Best Practices for Building Agents Recap
Arthur

Why Does Single-Shot Red Teaming Miss Agent Vulnerabilities?

October 6, 20265 min read

Single-shot red teaming sends one adversarial prompt, checks whether the agent refuses, and records a result. Against a model trained to decline obvious attacks, almost everything passes. One vendor evaluation ran this way and found zero vulnerabilities. The same system tested over multiple turns surfaced 160. The agent did not get weaker between the two tests. The test got more honest.

The gap comes from a mismatch built into how agents are trained and how they are used. Models are trained to refuse a harmful request once. They are designed to stay helpful and compliant across a long conversation. Single-shot testing probes the first behavior and never reaches the second, so it reports safety the agent does not actually have.

What single-shot red teaming actually measures

A single-shot test is one prompt in, one response out, scored in isolation. Ask the agent to do something it should refuse, confirm it refuses, move to the next prompt. It is fast, easy to automate, and easy to run at volume.

It measures one thing: whether the agent's first-line refusal fires on an obvious attack. That is a real property worth checking. It is also the property models are most heavily trained to get right, which is why a single-shot pass rate climbs toward perfect while the agent remains exploitable. A clean single-shot report is close to the expected result, not evidence of a secure system.

Why agents fail across turns, not in one shot

An agent's job is to carry context forward. It remembers what was said earlier, builds on prior steps, and tries to be useful as a conversation develops. Every one of those traits is an opening an attacker can work over multiple turns.

A multi-turn attack does not ask for the harmful thing directly. It arrives at it:

  • Priming. Early turns establish a framing, a role, or a hypothetical the agent accepts as benign. Nothing in those turns trips a refusal on its own.
  • Incremental escalation. Each turn asks for slightly more than the last. No single step crosses an obvious line, so no single step gets refused.
  • Context accumulation. The agent carries forward details, permissions, or assumptions granted earlier, and later turns exploit that accumulated state.
  • Persistence. The attacker rephrases and retries. A system trained to refuse once is not trained to refuse the same goal approached ten different ways.

The root cause is the same across all four: the agent refuses a request in isolation but complies with the same goal when it is assembled gradually. Single-shot testing can only see the isolated request. It is structurally blind to the path.

The numbers behind the gap

The scale of what single-shot misses shows up when the same systems get tested both ways.

  • One vendor evaluation found zero vulnerabilities on single-shot testing and 160 on multi-turn testing of the same systems.
  • A broad red-teaming exercise reported that all 34 leading models it tested were compromised.
  • Attack success rates in that work ranged from 1.5% to 92% depending on the model and method.

That 1.5%-to-92% spread is the part worth sitting with. A model that resists one attack method at a 1.5% success rate can fold to another at 92%. A test that tries one method and stops has no way to find the 92% door. It reports the 1.5% and calls the system safe.

Why this matters more for agents than for chatbots

A chatbot that gets talked into saying something it shouldn't produces bad text. An agent that gets talked into it takes an action. It calls a tool, writes to a system, moves data, or triggers another agent. The multi-turn path that a chatbot renders as an inappropriate answer, an agent renders as an inappropriate action with real consequences.

Agents also run longer and hold more state than a single chat exchange. They retrieve documents, call tools, and pass context between steps, and each of those is another surface where an earlier turn can plant something a later turn uses. The attack surface a multi-turn test needs to cover is larger for an agent, which makes the single-shot blind spot more expensive.

What adaptive testing looks like instead

If single-shot testing reports false negatives, the fix is testing that adapts the way a real attacker does: multiple turns, multiple methods, and a goal it keeps pursuing across the conversation rather than a prompt it fires once.

Adaptive red teaming:

  • Runs multi-turn. Attacks unfold across a conversation, so priming, escalation, and context accumulation are all in scope.
  • Varies the method. It tries many paths to the same goal rather than one phrasing, which is what surfaces the 92% door next to the 1.5% one.
  • Persists. It retries and rephrases, because a system that refuses once is not proven to refuse the same goal approached differently.
  • Tests the actions, not just the text. For agents, it checks whether the conversation can drive an impermissible tool call or action, not just an impermissible sentence.

This is harder to run than single-shot, which is exactly why single-shot stays popular and keeps reporting clean. The cost of adaptive testing is real, and so is the cost of shipping an agent that a clean single-shot report declared safe.

Testing once is not the same as testing deep

Running an adaptive test before launch still only describes the agent on the day you ran it. Agents change. Prompts get edited, models get swapped, tools get added, and the conversation space shifts underneath a test that passed last quarter. Depth per test and frequency of testing are two different problems, and an agent needs both.

Two practices keep the gap closed after launch. Continuous evals run against real production traffic and catch behavior that only emerges once real users engage the agent over real conversations, which is where multi-turn failures actually live. And observability built on OpenTelemetry tracing captures the full multi-turn path, every input, tool call, and decision, so when a multi-turn failure does appear you can replay the conversation that produced it instead of guessing. Adaptive testing finds the vulnerability. These keep it from coming back silently.

The short answer

Single-shot red teaming misses agent vulnerabilities because it tests the one behavior models are trained hardest to get right, refusing an obvious request in isolation, and never tests the behavior agents are designed for, staying compliant and helpful across a long conversation. Attacks that prime, escalate, accumulate context, and persist only appear over multiple turns, which is why the same systems can show zero vulnerabilities on single-shot and well over a hundred on multi-turn. Test adaptively, across many turns and many methods, and pair it with continuous evals and observability so the depth you test for before launch holds up after it.

Want to see multi-turn failures caught against real agent traffic? Book a demo with an AI expert.

‍

SHARE