From 1,200 Agents to a Governance Queue: How to Triage and Score AI Agents After Discovery
Discovery used to be the hard part. Finding the agents running across an enterprise, the ones in vendor software, the ones a business user spun up, the ones a framework deployed under the hood, took real work. That problem is mostly solved. Telemetry, MCP monitoring, network analysis, and cloud platform APIs now surface agents faster than any team can act on them.
The new bottleneck sits one step later. Discovery hands you a list of 1,200 agents and no instruction on what to do with it. You can't review them all at once, and most teams freeze at exactly this point. The inventory is complete and useless at the same time. What stands between a list and a governed environment is triage: scoring each agent and turning the pile into an ordered queue.
Why discovery output isn't a to-do list
A raw inventory of 1,200 agents is noise. It tells you agents exist. It doesn't tell you which one can write to a production database, which one reads customer PII, and which one summarizes meeting notes for a single team.
Treat the list as uniform and you lose either way. Review it top to bottom and the high-risk agent sits unreviewed for weeks behind forty harmless ones. Spread attention evenly and a low-stakes internal helper gets the same scrutiny as an unowned agent with write access to financial systems. Neither is governance. Both are just motion.
Triage fixes this by ranking the list. It decides what gets human attention first, what can wait, and what can be monitored passively without a review at all. A queue is actionable in a way a list never is.
The two questions triage answers
Good triage separates two questions that are easy to blur together.
The first is a confidence question. How sure are we this is actually an agent, and that we understand what it is? An agent confirmed across telemetry, MCP, and network signals is well understood. One inferred from a single weak signal might be an agent, might be a stray API call, might be something misclassified.
The second is a risk question. If this is an agent, how much damage could it do? An agent with write access to customer records and no owner is a different problem from a read-only agent summarizing public documents.
These feed the same priority ranking but mean different things, and keeping them distinct is what makes the score defensible. A high-risk agent you're only 40% sure exists needs investigation to resolve the uncertainty. A low-risk agent you're 100% sure about can wait regardless of how confident you are. Collapse the two questions into one number and you lose the ability to tell "investigate this" from "govern this."
Scoring an agent's risk exposure
Risk exposure is the heavier half of the score, and it should come from concrete signals rather than an abstract severity label. The factors that matter most:
Data sensitivity. What can the agent reach? An agent with a path to PII, health records, financial data, or proprietary IP scores higher than one touching aggregate or public data.
Access type. Read versus write is one of the sharpest dividing lines. An agent that can only read is bounded. An agent that can write, update records, trigger transactions, or modify systems can cause damage that outlives a single response.
Tool and system access. The more tools and systems an agent can call, the larger its blast radius. An agent wired to a database, an email API, and an internal service has more ways to go wrong than one with a single tool.
Autonomy. An agent that acts without a human in the loop carries more risk than one that proposes actions for approval. Continuous evals on an agent's production behavior can sharpen this signal over time.
External exposure. Agents that are customer-facing or reachable from outside the network invite prompt injection and misuse that purely internal agents don't.
Ownership. An unowned agent scores higher almost by definition. No owner means no one accountable for its behavior, no one to answer a compliance question, and no one to call when it misbehaves.
The point of scoring on signals rather than gut feel is that the score holds up to scrutiny. When a security lead asks why a given agent jumped the queue, "it can write to customer records, it has no owner, and it's reachable externally" is an answer. "It felt risky" is not.
Factoring in discovery confidence
Confidence changes what a risk score means in practice. Two agents can carry the same risk exposure and warrant completely different actions depending on how well you understand them.
An agent confirmed across multiple discovery signals, emitting telemetry, exposing an MCP server, visible in network traffic, is well characterized. You know its tools, its access, and its behavior, so a high risk score points straight to governance.
An agent inferred from one weak signal is a different situation. The high risk score might be real, or it might be an artifact of incomplete information. Here the first move is investigation, not governance. You resolve the uncertainty before you decide what controls to apply.
This is why confidence and risk combine rather than average. A high-confidence, high-risk agent goes to the front of the governance queue. A low-confidence, high-risk agent goes to the front of the investigation queue, which is a different line. A high-confidence, low-risk agent can wait. Separating the two keeps you from either ignoring a real threat you haven't confirmed or burning review cycles on a false positive.
From score to priority tiers
A score is only useful if it maps to action. In practice that means grouping scored agents into a handful of tiers, each with a clear operational meaning.
Govern now. High-risk, high-confidence agents, especially unowned ones with sensitive access. These need an owner, a risk classification, and enforced policy before anything else moves.
Review soon. Agents with meaningful risk that aren't immediate emergencies, or high-risk agents still awaiting confidence confirmation. They enter the queue behind the first tier and get worked through in order.
Monitor passively. Low-risk, well-understood agents. These don't need an individual review to operate. They stay in the inventory under continuous monitoring, and they only resurface if something about them changes.
Unowned high-risk agents deserve a specific call-out because they tend to jump the queue regardless of other factors. An agent with broad access and nobody accountable for it is the exact profile behind most agent-related incidents. Assigning an owner is often the single highest-leverage triage action available.
Keeping scores current
Triage isn't a one-time sort. An agent's score reflects its risk surface, and that surface moves. An agent that scored low last quarter gains a new tool, points at a new data source, or has its access scope widened, and its risk jumps. An agent gets an owner, and its score drops because someone is now accountable for it.
A static score decays into a wrong answer. The queue has to re-rank as agents change, which means re-scoring on change rather than on a calendar. The same discovery signals that catch a new agent also catch when an existing one's capabilities shift, and that shift should feed straight back into the score. An agent that crosses a risk threshold because of a new permission should move up the queue the moment the change is detected, not at the next quarterly review.
Triage at scale without reading every agent
The number that makes manual triage impossible is the whole reason triage matters. Nobody is going to read 1,200 agent profiles and assign each a thoughtful risk rating. By the time they finished, discovery would have added another two hundred.
Scoring has to be automated. The signals that drive the score, data access, write capability, tool count, autonomy, external exposure, ownership status, come from the same telemetry and discovery data that found the agent. That data can be scored programmatically the moment an agent appears, which means prioritization keeps pace with discovery instead of falling behind it.
Automation handles the ranking. Humans handle the top of the queue. Reviewers spend their attention on the agents that scored into "govern now," not on the long tail of low-risk helpers that a passive monitor can watch. That division is what makes governance tractable at thousands of agents instead of dozens.
How Arthur helps
Arthur's discovery feeds directly into scoring and tiering, so a newly discovered agent arrives with the signals needed to rank it: what it can access, what it can do, and whether anyone owns it. Agents are classified by risk, unowned high-risk agents are flagged for attention, and scores update as agents change. The result is a governance queue rather than a flat inventory, so teams work the agents that matter first.
A complete list of agents is where governance starts, not where it ends. Triage is what turns that list into a plan you can act on. Book a demo to see how Arthur scores and prioritizes agents at scale.