Best Practices for Building Agents Recap
Arthur

Allowed vs. Aligned: Why a Permitted AI Agent Action Can Still Cause Harm

October 6, 20266 min read

Most agent security work answers one question: is this action allowed? Allow or deny, pass or block, inside policy or outside it. That question matters, but it misses a second one that determines whether an agent is actually safe to run: is this action aligned with what the user meant?

An allowed action complies with policy. An aligned action serves the intended goal. The two sound similar, and most of the industry treats them as the same thing. They are not, and the gap between them is where a growing share of agent incidents live.

The definitions

Allowed means the action is within the permissions and policies the agent has been granted. The agent had the credential, the tool was in scope, the request passed the allow/deny check. Nothing was violated.

Aligned means the action serves the outcome the user actually wanted. It respects intent, not just permission. An aligned action is one a reasonable person would agree matched the goal behind the request.

A single action can be allowed and aligned, allowed and misaligned, or blocked outright. The dangerous quadrant is allowed and misaligned, because every existing permission check waves it straight through.

Why "allowed" isn't enough

Permissions describe what an agent can do. They say nothing about whether doing it serves the request. An agent operating entirely inside its granted access can still take an action that no one wanted, because permission systems were designed for a world where a human decided what to do with their access. Agents decide at runtime, and they interpret loose instructions literally.

Three patterns show how allowed and misaligned diverge.

Individually permitted actions combine into risk. A read-only knowledge agent connected to a CRM can exfiltrate sensitive data through a search query. Each step is permitted: it is allowed to read, allowed to search, allowed to return results. The combination produces an outcome the owner never intended. No single permission was violated, so no allow/deny check fires.

Loose prompts run against the wrong target. Consider a database administrator who asks an agent, "Will you help me clean up the database?" The request sits fully within the admin's own permissions. The agent uses those permissions on the production database instead of the demo environment. Nothing was unauthorized. The action was allowed and badly misaligned with what the admin meant.

Context redirects the goal. An agent ingests a document, email, or search result that carries instructions, and it acts on them using its legitimate access. The permissions held the whole time. The intent behind the session did not.

In each case, an allow/deny system reports success. A human looking at the outcome would call it a failure.

Why agents widen the gap

Traditional software did exactly what it was coded to do, so "allowed" and "intended" collapsed into one idea. Agents break that assumption in three ways.

They are non-deterministic. The same request can produce different action sequences on different runs, so you cannot enumerate every path in advance and pre-approve it.

They interpret intent. An agent translates a vague instruction into concrete tool calls, and it can translate correctly on permissions while translating wrongly on purpose.

They chain actions. An agent composes many permitted steps into a plan, and the risk often lives in the composition rather than any individual step.

Scoping permissions tightly helps, and least privilege is still worth doing. But you cannot scope your way to alignment. Permissions govern access. Alignment is about behavior, and behavior has to be checked while the agent runs.

How to check for alignment, not just permission

Closing the gap means adding controls that assess intent and behavior alongside the ones that assess access. Four practices do most of the work.

Instrument the full decision chain. You cannot judge whether an action was aligned if you cannot see what led to it. Capture the input, the retrieved context, the model's decision, the tool it selected, the permission it checked, the action it took, and the output. With that trace, you can replay a decision and ask whether it matched intent, not just whether it was authorized. Arthur builds this on OpenTelemetry and OpenInference tracing so the decision chain is visible per action, not just at the system level.

Run continuous evals on behavior. Alignment is a behavioral property, so it needs behavioral checks. Unsupervised evals like goal accuracy ("did the agent call the right tools to fulfill the user's intent?") and topic adherence ("did the agent stay within the scope defined by its system prompt?") assess whether an action served the request, not whether it was permitted. Because they run against every production interaction, they catch misalignment that permission checks pass.

Add guardrails that validate actions and outputs. Post-LLM guardrails can check whether the agent selected the right tools for the request and whether its output is grounded in the context it actually had. Tool and action validation targets the allowed-but-misaligned case directly: the action was permitted, but was it the right one? Guardrails can catch the mismatch and feed it back to the agent for correction before anything reaches the user.

Treat ingested context as untrusted. Because misalignment often enters through data the model reads, pre-LLM guardrails and prompt-injection detection keep hijacked instructions from turning legitimate permissions toward the wrong goal.

The short answer

An allowed action passes your permission checks. An aligned action does what the user meant. Permission systems only measure the first, so an agent can operate entirely within policy and still do real damage. Closing the gap takes observability to see what the agent did, continuous evals to judge whether it matched intent, and guardrails to catch misaligned actions before they land.

If your agent security program spends all its time on allow and deny, it is answering half the question. Arthur is built to answer the other half.

Want to see alignment checks running against real agent traffic? Book a demo with an AI expert.

‍

SHARE