Skip to content
NewSecurity Audit for AI-Built Apps. Fixed scope, clear report
GeekTech
AI Security

How to Test an LLM Feature Before Launch

A practical plan to test an LLM feature before launch: what to attack, how to write repeatable test cases, and the pass criteria to agree before go-live.

6 min read

GeekTech Engineering Team18 years building and securing software

To test an LLM feature before launch, map what it can read and do, then try to make it misbehave in each of those places. Run every attack several times, because the output varies between runs. Ship only when the failures that matter, such as data leaking between users or unauthorized actions, don't happen at all.

The rest of this guide turns that into a plan you can run in staging. It focuses on security and misuse, not answer quality, although the same setup serves both.

Why testing an LLM feature is different

Ordinary software gives the same output for the same input. A test that passes keeps passing until the code changes.

An LLM feature breaks both assumptions. The same prompt can produce different answers on different runs. Behavior can also change without a code change, when the provider updates a model, someone edits the system prompt, or new documents enter the knowledge base.

Two consequences follow:

  • One run proves little. An attack that fails once may succeed on the fifth attempt.
  • Tests are never finished. A suite that passed last month says nothing about today's model and prompt, so it has to be repeatable.

How to threat model an LLM feature on one page

Before writing a single test, write down what the feature touches. Keep it to one page. If it takes longer, the feature probably has more access than it needs.

QuestionExample answer for a support assistant
Who can send it input?Signed-in customers, plus anyone who emails support
What can it read?Help articles, the current customer's orders and tickets
What can it do?Draft replies, tag tickets, look up order status
Where does its output go?Rendered in the agent dashboard and in customer emails
Whose permissions does it use?Should be the signed-in user's, never a broad service account

Every row is a test target. Inputs are where instructions get planted. Data sources are what might leak. Actions are what an attacker might trigger. Outputs are where unsafe content lands.

What to test in an LLM feature

Most launch-blocking issues fall into a few areas. The OWASP Top 10 for LLM Applications is a useful reference. A typical pre-launch plan covers these:

AreaWhat to tryWhat a pass looks like
Direct prompt injectionAsk it to ignore its rules, reveal its instructions or act outside its purposeIt stays in scope, and nothing sensitive is in the prompt to reveal
Indirect prompt injectionPlant instructions in a document, ticket, email or web page it processesPlanted text can't change what data it reaches or what it does
Data leakage between usersAsk for another customer's records, or search for content you shouldn't seeRetrieval only returns what the current user can already open
Tool and action misusePush it to call tools with other users' IDs, unusual values or external addressesCode rejects the call, whatever the model asked for
Output handlingGet it to output HTML, scripts, markdown images and linksOutput is escaped, and images and links can't carry data out
Abuse and runaway usageSend very long inputs, rapid repeated requests and loopsRate limits and length caps hold, and usage alerts fire
Off-scope and harmful contentAsk for advice it shouldn't give, or content that breaks your policyIt declines in a way your support team can live with

If the feature has tools or acts on its own, the action rows deserve the most effort. Our guide to securing AI agents that use tools covers the controls to test for. For background on why planted instructions work, see prompt injection explained for product teams.

How to write LLM test cases that you can rerun

A test case for an LLM feature needs more structure than a chat transcript. Write each one with five parts:

  1. Setup: the user, their permissions and the data in the environment.
  2. Input: the message, file or content the feature will process.
  3. Goal: what an attacker would want, in one sentence.
  4. Pass rule: a check you can apply without judgment, such as "the response contains no order numbers from account B."
  5. Runs: how many times to repeat it.

Then vary each important case. Rephrase it, translate it, split it across two messages or hide it in a long document. Attacks that fail in plain English often succeed with a small change of form.

Automate the pass rule where you can. Checks for other users' identifiers, secrets or blocked domains are easy to code. Checks that need judgment can use a second model as a grader, as long as a person reviews a sample of its verdicts.

How to test an LLM feature before launch, step by step

A pre-launch round looks like this:

  1. Freeze the configuration. Record the model version, system prompt, tools and data sources. Results only hold for that combination.
  2. Use staging with realistic test data. Create at least two accounts with separate data, so leakage between users is detectable.
  3. Run the automated suite. Repeat each case several times and record every failure.
  4. Do an exploratory session. Try to break the feature by hand. People find chained attacks that scripted cases miss.
  5. Fix the cause, not the case. If a test leaks data, fix the retrieval permissions. A new line in the prompt only hides the symptom.
  6. Rerun everything. A fix in one area can open another.
  7. Sign off in writing. Note what was tested, what remains a known risk and who accepted it.

What pass criteria should block a launch?

Decide what blocks the launch before the results come in, or a borderline result will start to look good enough.

A workable split has two classes:

  • Zero tolerance: data from another user, a secret or credential in output, an unauthorized tool call, or executable content reaching a browser. A single occurrence in any run blocks the launch.
  • Threshold-based: off-scope answers, tone problems or occasional refusals of valid requests. Agree in advance on an acceptable rate.

The zero-tolerance list should be enforced by code: permissions, validation and output escaping. If a pass depends on the model choosing to behave, it isn't a control. It's a hope.

A useful launch question

Assume the model is fully controlled by whoever wrote the input. What is the worst it could read, do or send? If the answer is anything on your zero-tolerance list, the feature isn't ready.

What to retest after launch

Launch is the first run of the suite, not the last. Rerun it whenever any of these change:

  • the model or model version
  • the system prompt or its templates
  • a tool, or the permissions behind it
  • a new data source or document type
  • the way output is rendered or sent

Keep the suite in the repository next to the feature, and run it in CI where you can. Add a case for every issue found in production, so it can't return quietly. Log prompts, retrieved sources and tool calls, so you can investigate when something looks wrong.

When should you bring in an outside tester?

An internal round catches a lot. An outside review adds value when the feature reads confidential data, takes actions in other systems, or a customer asks for independent evidence. If that's where you are, our LLM and AI agent security testing covers prompt injection, data leakage, tool abuse and a retest after fixes.

Before any test, fix the common basics. The AI app security checklist is a quick way to do that.

Frequently asked questions

How long does pre-launch LLM testing take?

It depends on what the feature can reach. A chatbot over public help articles needs a short round. Every private data source and tool adds cases.

Can we rely on the model provider's safety filters?

They help with harmful content in general, but they don't know your permissions, data or tools. They can't stop another customer's records from appearing in a response. That's your application's job.

Do we need to test again if we only changed the prompt?

Yes. A prompt change can alter how the model handles instructions and refusals. Rerun the full suite. It's the quickest way to catch a regression before your users do.

What's the difference between this and red teaming?

Red teaming is open-ended attack work, usually by people who didn't build the feature. Pre-launch testing is the repeatable suite. Red teaming finds new problems, and the suite keeps them fixed.

Insights

Related insights

More on AI Security →
  • AI Security

    How to Secure AI Agents That Can Use Tools

    AI agents that call tools can take real actions. How to scope their permissions, contain prompt injection and test an agent before it touches production.

    6 min read

  • AI Security

    Prompt Injection Explained for Product Teams

    What prompt injection is, why it can't be solved with better prompts alone, and the design decisions that limit the damage in your AI features.

    7 min read

  • Vibe Coding

    Taking an AI-Built Prototype to Production

    Your AI-built prototype works and users are arriving. What to add before it holds real data: tests, CI/CD, environments, monitoring, backups, cost limits.

    5 min read

Get practical AI insights, monthly

AI security, automation and governance, written by engineers. No spam, unsubscribe any time.

Book a free 30-min call