Red Teaming Agents Can't Be a Once-a-Year Exercise
Most writing about attacking AI agents, including ours, explains the attack classes one at a time. Here is what prompt injection looks like in a multi-agent system. Here is how tool poisoning works. Here is memory poisoning, here is data exfiltration through a sanctioned tool. All of it describes what the attack is and why it matters.
None of it treats adversarial testing as an ongoing engineering practice, with a cadence, a trigger model, and an output that feeds back into the system. That is the gap this piece is about. The question isn't which attacks exist. It's how often you test for them, what makes you test, and what you do with a result.
Why the annual assessment is the wrong shape
Annual or quarterly red teaming is a cadence borrowed from traditional application security. It made sense there because the thing being tested changed only when someone deployed. You shipped a release, you tested the release, and between releases the attack surface held still.
Agents break that assumption in three ways, and each one on its own would justify a shorter cycle.
They are non-deterministic. An attack that fails on one attempt can succeed on the next. Temperature, sampling, and context ordering all move the output. A single pass tells you an attack did not work that time. It does not tell you the attack does not work. Run the same injection ten times and you may see it land twice, which is two times too many.
The attack surface changes without a deploy. An MCP tool definition can change underneath a running agent. A model provider can update the model behind the same API endpoint. A RAG corpus grows every day as new documents land in it. The agent has shipped nothing, its code is byte-for-byte identical to last week, and its behavior under attack is different. A deploy-triggered test cycle never fires, because there was no deploy.
The techniques move faster than the review cycle. Whatever was tested in the last assessment reflects the attack landscape of that quarter. New jailbreak families and injection patterns surface constantly. A test suite frozen three months ago is testing against a threat model that has already moved.
The real payoff: red teaming becomes a data source
The usual case for automating red teaming is coverage and cost. Run more attacks, run them cheaper, cover more agents. That is true and it is not the point.
The point is that automation changes what red teaming is. Done as an annual event, it produces a report. Done continuously, it produces data. Every attack that succeeds becomes a test case. Your evaluation suite stops being a list someone imagined and starts growing out of your own failures. An attack that worked once becomes a permanent regression test, so it cannot quietly start working again three model versions later after everyone forgot it was ever a problem.
That loop is what makes the practice compound instead of repeat. A calendar-driven assessment tests the same rough set of things every year and throws away everything it learns the moment the report is filed. A data-driven one accumulates. Every failure it finds is a failure it will never let recur silently.
This is the same idea as building regression datasets from real production failures, applied to adversarial inputs instead of organic ones. There, you take the cases where your agent got a real user's request wrong and turn them into tests that gate every future change. Here, you take the cases where an attacker got your agent to misbehave and do the same thing. The mechanism is identical. Only the source of the failure differs. Both grow a suite that represents what your agent must never do again.
What to test for
Ground your attack library in the OWASP Top 10 for Agentic Applications rather than an invented list, so your coverage maps to a shared standard a reviewer can check. At minimum, test for:
- Goal hijacking, where an attacker redirects the agent from its assigned task to one of their choosing.
- Direct prompt injection, malicious instructions in the user input designed to override the system prompt.
- Indirect prompt injection, instructions that arrive inside a retrieved document or a tool response rather than the input field.
- Jailbreak, getting the agent to ignore its safety instructions and produce content it was told to refuse.
- Data exfiltration through sanctioned tools, using the tools the agent is supposed to have to move data it is not supposed to move.
- Excessive agency, the agent taking actions well beyond what the request warranted because its permissions were too broad.
Give indirect injection its own attention
Indirect injection is the one most testing programs miss, and the reason is structural. The attack does not arrive through the user input field. It arrives inside a document the agent retrieved from its corpus, or inside a tool response it trusted and acted on. A malicious instruction sits in a support ticket, a web page, or a database row, and the agent reads it as though it were part of its own reasoning.
A test harness that only fuzzes user prompts will pass an agent that is wide open to this. Every attack it fires enters through the front door, and the agent's front door might be well defended while its retrieval path is completely exposed. If your red teaming injects only through the user turn, you are testing one surface and reporting confidence about all of them. Test the retrieved documents and the tool outputs, not just the prompt.
Three distinctions worth drawing clearly
Red teaming is not evaluation. Continuous evals measure whether an agent does what it is supposed to do: does it answer accurately, stay on topic, ground its claims in context. Red teaming probes for what the agent can be made to do against its instructions. Same infrastructure, opposite intent. One asks whether the agent works. The other asks whether the agent can be broken.
Red teaming is not a guardrail. A guardrail is the control that intercepts a bad input or output at runtime. Red teaming is how you find out whether that control actually holds. Teams that have deployed guardrails often treat the problem as handled and stop testing, which is exactly backwards. A guardrail you never attack is a guardrail you are trusting on faith. Red teaming is what turns that faith into evidence, and it is what catches the guardrail that quietly stopped working after a config change.
Automated is not a replacement for human. Automation gives you breadth, cadence, and regression coverage. It runs thousands of known attack variations on every change, which no human team could do by hand. But it finds variations on attacks that already exist. Human red teamers find novel attack classes that no generator will produce, because inventing a genuinely new technique is a creative act, not a combinatorial one. You need both. Automation for the floor, humans for the frontier.
Trigger on what changed, not on the calendar
Cadence is the practical question, and the answer is event-driven rather than calendar-driven. Run the relevant attack suite when the thing that changes the attack surface changes:
- On a model version change, since a new model behind the same endpoint can respond to attacks differently.
- On a tool or MCP definition change, since a modified tool is a modified surface.
- On a prompt promotion, since a new system prompt can open a hole the old one closed.
- On a RAG corpus update, since new documents are new indirect-injection vectors.
- On a rolling baseline schedule for everything else, so nothing goes untested indefinitely just because it has been stable.
The principle underneath all of these: the trigger should be the thing that changed the attack surface, not the date on the calendar. A quarterly cadence tests an agent that changed twelve times since the last run and misses eleven of those changes. An event-driven one tests every change when it happens.
Be honest about the limits
Automated red teaming finds variations on known attacks. It will not find the attack nobody has thought of yet. If a technique isn't in your generator's repertoire, running the generator a million more times will not surface it. That is the ceiling, and pretending otherwise is how programs develop false confidence.
A clean run is not proof of safety. It is evidence that a specific set of techniques did not work on a specific day, against a specific version of the agent, at a specific sampling temperature. Report it as exactly that and no more.
And to be direct about the state of the art: prompt injection is not a solved problem. Anyone claiming a deterministic, security-grade defense against it is overstating what the field can currently do. The defenses that exist reduce risk, they do not eliminate it, and treating them as airtight is how agents end up shipped into production wide open. Continuous adversarial testing is not how you prove your agent is safe from injection. It is how you keep measuring how exposed it is as everything around it changes.
What to instrument
- Attack success rate over time, tracked as a real metric per agent. A rising success rate is a signal that something in the agent's surface changed, the same way a rising eval failure rate is.
- Every successful attack converted into a permanent test case, so a technique that worked once can never work again unnoticed.
- Event-driven triggers tied to change, wired to model version, tool and MCP definitions, prompt promotions, and corpus updates.
- Coverage across every registered agent, not just the flagship. The agent nobody is watching is the one most likely to be exposed.
- Results attached to the agent record, so a reviewer looking at an agent can see its testing history in the same place as its tools, owner, and data sources, rather than hunting for a separate report.
Adversarial testing that runs once a year produces a document. Adversarial testing that runs on every change, feeds its failures back into a growing suite, and tracks its success rate over time produces a defense that gets stronger each time something tries to break it.
See how Arthur helps teams test, evaluate, and govern agents continuously. Book a demo or explore the Agent Development Toolkit to get started.