A Runaway Agent Is a Control Failure, Not a Budget Problem
Most writing on agent spend, including ours, frames token limits as savings: cheaper runs, lower bills, better unit economics. That framing is fine as far as it goes, and it misses what the limit actually does in the moment it fires.
When an agent enters a retry loop or a recursive delegation spiral, the spend ceiling is frequently the only thing that stops it. The bill you avoided is real, but it is a side effect. The control is the point. Loop caps and spend ceilings are containment controls that happen to save money, and once you see them that way, you instrument them differently.
A cost symptom, not a cost problem
A runaway agent burning tokens in a loop is doing three things at once. It is failing at its task. It is consuming resources. And it is often holding a connection, a lock, or a database session while it does. That is a control failure with a cost symptom, not a cost problem with a technical cause.
The distinction changes who owns it and how fast it gets caught. Treat runaway spend as a budget line and it surfaces in a monthly review, weeks after the damage. Treat it as a control and it gets a real-time ceiling and an alert the moment it trips. Same number on the invoice, entirely different response time.
What actually runs away
Four failure modes account for most runaway execution, and each spirals for a different reason.
Retry loops. A tool call fails, the agent retries, the retry fails the same way, and nothing bounds the number of attempts. Because agent behavior is non-deterministic, the loop can persist without ever making progress. The agent keeps trying because, from its perspective, the next attempt might work.
Recursive delegation. One agent calls another, which calls back, or a planner spawns subagents that spawn subagents of their own. Without a depth limit, the tree keeps growing until something external stops it. Each level looks locally reasonable, which is what makes the total invisible until the spend spikes.
Unbounded tool-calling. An agent that can call tools in a loop, with no cap on calls per run, can hammer an API or a database far past anything useful. A single invocation issuing hundreds of queries is rarely doing productive work, but nothing inside the agent's own logic tells it to stop.
Pathological context growth. Each step appends to the context window. Cost per call climbs as the window grows, so a long-running agent gets more expensive per step the longer it runs. This one compounds the others: a retry loop with growing context is more expensive on its hundredth iteration than its first.
Traces are what make these visible. Instrumenting an agent to emit spans for its LLM calls, tool invocations, and token counts turns a runaway from a mysterious bill into a legible sequence you can watch spiral.
The limits, as controls
Frame each limit by what it contains, not what it saves.
Per-run step and iteration caps. Bound how many reasoning-action cycles a single invocation can execute. This contains loops that make no progress, since a loop that never advances will hit the cap and stop rather than run forever.
Token and spend ceilings, per invocation and per window. A hard stop on consumption for a single run, plus a rolling ceiling across a time window. The per-invocation ceiling catches one runaway request. The windowed ceiling catches a fleet of small runaway runs that individually stay under the limit but collectively don't.
Delegation depth limits. Cap how deep the agent-calling-agent tree can go. This contains recursive delegation directly, putting a floor under how far a spawning spiral can descend before it is forced to stop.
Wall-clock timeouts. Bound total execution time regardless of token count. This catches an agent stuck waiting on a slow dependency or looping slowly, cases where token-based ceilings alone would take too long to trip.
A ceiling is not a guardrail
This is the distinction worth being sharp about, because the two get conflated.
A guardrail inspects content. It asks whether an input is a prompt injection, whether an output is a hallucination, whether a response contains PII. It reads what the agent is saying and makes a judgment about it.
A spend ceiling inspects nothing. It counts resource consumption and trips at a threshold. It has no opinion about whether the agent's calls are good or bad, only about how many there have been.
Same containment goal, different mechanism. A ceiling is a circuit breaker on consumption, and it fires on behavior a content guardrail cannot see: an agent looping over perfectly valid, individually reasonable calls, forever. Every request that guardrail would inspect passes cleanly. The problem is the volume, not the content, and only a consumption limit catches it.
This is why you need both. The guardrail catches bad content. The ceiling catches bad resource behavior. Neither substitutes for the other, and a system that has guardrails intercepting bad inputs and outputs but no consumption ceiling is still wide open to the runaway loop.
The tradeoff
A ceiling that stops a runaway can also stop real work. A research agent calling tools dozens of times and building up context as it goes looks a lot like a loop for a while. Set the ceiling too low and you cut off legitimate tasks before they finish. Set it too high and the runaway does its damage before the limit trips.
There's no universal number, so the ceiling has to be per-agent. Pick it against the failure you'd rather live with: cutting off the occasional long task, or letting the occasional loop run longer than you'd like. Decide that when you build the agent, not while you're watching it burn through an incident.
What to instrument
- A defined step, token, spend, and depth limit per agent, rather than a single global default that fits nothing well.
- Limit breaches emitted as telemetry, so a trip is a visible event and not a silent kill.
- An alert when an agent approaches or hits a ceiling, since repeated trips signal a real defect rather than an unlucky run.
- Boundary behavior defined per agent: hard stop versus pause-for-approval.
- The resolved limits attached to the agent record, so a reviewer can see what bounds each agent operates under.
A spend ceiling that only shows up on the invoice is doing half its job. Wire it as a control, alert on it, and the runaway loop becomes a caught event instead of a line item you explain after the fact.
Arthur helps teams discover, govern, and monitor AI agents across their environment, including the execution limits that keep a runaway from becoming an incident. Book a demo to see it on your own agents, or explore the Agent Development Toolkit to start building.