Best Practices for Building Agents Recap
Arthur

A Count Is Not an Inventory: Grading Agent Discovery Findings by Evidence Strength

September 9, 20264 min read

Every agent discovery tool reports a number. We found 271 agents. The number goes on a slide, into a board deck, in front of an auditor. It looks like progress, and in one narrow sense it is. You now know that 271 is bigger than zero.

But those 271 findings are not the same kind of object. One of them emits live telemetry with named tools, a known model, and identified data sources. Another is a resource sitting in a cloud console that has never once been invoked. A third is a single process name spotted on an employee laptop. Counting all three as "an agent we found" produces an inventory that looks complete and can't be acted on.

The fix is to treat detection confidence as a governance primitive. Until you grade the evidence behind each finding, you can't triage, you can't prioritize, and you can't tell an auditor what your inventory actually proves. A count answers how many. Governance needs to know what you know about each one, and how you know it.

Why a flat list fails

A flat count hides the only question that matters once discovery is running: what is the evidence behind each finding, and what does that evidence let you do?

The failure mode is concrete. A dashboard reads 271. A security lead is asked which of those 271 needs attention this week and has no way to answer, because the number carries no information about certainty, behavior, or ownership. Everything is flattened into a single undifferentiated pile. The high-volume production agent that's already well behaved sits in the same bucket as the unnamed process making outbound calls to a model provider from an unmanaged machine.

Grading turns that pile into a set of decisions. Each finding gets sorted by the strength of the evidence behind it, and each tier of evidence carries a different governance consequence.

A grading model: four tiers of evidence

The point of tiering is to be honest about what each finding proves. Define each tier by what the evidence supports, what it does not, and what you're supposed to do about it.

Traced. Live telemetry, with spans, tool calls, the model in use, and the data sources the agent touches all visible. You know what this agent does because you can watch it do it. This is the strongest evidence a discovery system can produce, and it comes from agents that emit traces with spans and tool calls the way instrumented systems do. Governance consequence: a traced agent can be evaluated and policy-enforced immediately. You have everything you need to assess its risk surface and apply controls.

Partial. The resource exists and is deployed, but no invocations have been observed. You have the declared definition, not the observed behavior. It might be a draft someone abandoned, it might be a staging artifact, it might go live tomorrow with production traffic. The evidence tells you the agent could run, not that it has. Governance consequence: a partial finding needs a watch state. It will either activate, at which point it should move up to traced, or it should be cleaned up. Leaving it unresolved means carrying a resource you can neither assess nor retire.

Thin. A single weak signal. A process name on an endpoint. One outbound call to a model provider API. You know something is running, but you may not know its name, its owner, or what data it touches. The evidence establishes existence and almost nothing else. Governance consequence: a thin finding needs investigation before it means anything. It's a lead, not a conclusion. The work is to pull more signal until it resolves into a real object or turns out to be noise.

Unattributed. Telemetry is arriving, sometimes at high volume, with no service name and no way to map it back to a team or an application. You have behavior without identity. You can see what it's doing and can't say whose it is. Governance consequence: this is the hardest case, because the more spans arrive, the more urgent it becomes and the less you can act on it. Volume signals that something real is running in production, but with no owner to route it to, urgency has nowhere to go.

What grading changes operationally

Sorting findings by evidence strength changes three things about how a governance program runs.

Triage order. A thin finding of an outbound model-provider call from an unmanaged laptop may deserve attention before a traced agent that's already behaving well. The traced agent is known, scoped, and governable today. The thin finding is unknown and sitting outside managed infrastructure. Sorting by volume or by discovery timestamp would put these in exactly the wrong order, burying the finding that carries the most uncertainty under the ones you already understand. Grading sorts by what you don't know, which is where the risk lives.

What you tell an auditor. "We discovered 271 agents" invites the obvious follow-up: how do you know, and what does that number represent? "We have 180 traced, 40 partial, 38 thin, and 13 unattributed, with a defined resolution path for each" is a defensible position. It says you understand the limits of your own detection and you have a plan for the findings that aren't yet resolved. The first answer describes a spreadsheet. The second describes a program.

Coverage honesty. Grading exposes where your sensors are weak. If most of your endpoint findings never rise above thin, that's a statement about your endpoint sensor, not about the agents behind them. The grade separates what's true about the agent from what's true about your ability to see it. Without that separation, a weak sensor looks like a quiet environment, which is the most dangerous misread a security team can make.

The resolution path

Grading isn't the end state. It's the input to a workflow that moves findings up the tiers toward certainty and ownership.

Thin findings get investigated until they resolve into partial or traced, or get dismissed as noise. Unattributed telemetry gets fingerprinted and correlated until it maps back to a named application. Partial agents get watched until they activate or get retired. The direction of travel is always the same: toward evidence strong enough to govern on, and toward an owner accountable for the result.

The goal is that every finding eventually lands in an owned, governed application, which is where a discovered inventory becomes something you can actually govern, with a named owner, a risk classification, and guardrails and policy enforcement applied against it. A count tells you something exists. A graded inventory tells you what you can prove, what you still need to run down, and what to do about each one.

Govern every agent, not just the ones you can count

Arthur discovers agents across your environment and grades what it finds by the strength of the evidence behind each detection, so your inventory reflects what you actually know. From there, every finding moves toward an owned, governed application with the right controls applied.

Book a demo to see how Arthur turns discovery findings into a governable inventory, or explore the Agent Development Toolkit to go deeper on building and governing production agents.

‍

SHARE