Best Practices for Building Agents Recap
Arthur

Sensor Health Is a Governance Control, Not IT Hygiene

September 9, 20266 min read

Most agent discovery writing, including ours, treats detection as a solved input. You point sensors at your environment and an inventory comes back. The dashboard shows a number, the number is large, and the number is reassuring.

Here is the case nobody writes about. A sensor stops working and nothing tells you. The inventory doesn't shrink. No alert fires. The count on the dashboard holds steady at the figure it reached the last time the sensor ran. You keep governing that number, and you keep trusting it, while it quietly stops meaning anything. You are no longer looking at your environment. You are looking at a snapshot that stopped updating and still looks alive.

A silent sensor is worse than no sensor at all

A gap you know about is a gap you can plan around. If a source was never connected, you know that region, that platform, or that account is dark, and you can weight your review accordingly. You account for the blind spot.

A gap you believe is covered is different. It is a false assurance, and you act on it. You tell an auditor the environment is monitored. You sign off on a coverage claim. You deprioritize a region because the sensor "already has it." Every decision downstream of a dead sensor inherits its failure, and none of them carry a warning.

That is why sensor health cannot sit in the IT-hygiene pile with certificate renewals and disk-space alerts. It is a governance control. It should be monitored, thresholded, and reported like one, because the integrity of your entire agent inventory depends on it.

How sensors fail quietly

The dangerous failures are the ones that leave the count looking normal. Here is where they come from and why nothing catches them.

Credentials expire. An API secret rotates or lapses. The sensor returns a 401 on every poll from that point forward. Findings freeze at the last successful scan, but the total count stays flat, and a flat count reads as stability rather than failure. Nothing about the number tells you it stopped moving because the sensor went blind, not because the environment went quiet.

Permissions narrow. Someone tightens a role in a cloud account for reasons that have nothing to do with you. The sensor still authenticates. It still returns results, just fewer of them. This is the hardest case to catch, because partial success looks like success. There is no error to alert on. The sensor is working exactly as designed against a scope that silently shrank.

Scope drifts. A new region, subscription, or project gets stood up and never added to the sweep. The sensor is healthy and blind at the same time. Every existing source reports cleanly, the dashboard looks green, and an entire slice of the environment is invisible because nobody wired it in.

A source was never configured. A connector exists in the product but was never finished during onboarding. It sits at zero forever. Zero is indistinguishable from "nothing there," so the source that was supposed to cover your endpoint fleet or your gateway traffic reports the same value it would report if that surface were genuinely empty. Nobody questions a zero.

Continuous sources get conflated with scanned ones. Telemetry-based sources report as data arrives, not on a fixed scan cycle. So a quiet period is ambiguous. It might mean nothing new is happening, or it might mean ingestion broke upstream. Judging a continuous source by the logic you use for a scanned one produces exactly the wrong read, and the two failure modes look identical from the dashboard.

What follows operationally

Once you accept that sensors fail silently, three things change about how you run discovery.

Staleness needs a threshold, not a timestamp. "Last scan 14 minutes ago" is useless on its own, because nobody defined what too old means for that source. And the answer differs by source class. Endpoint inventory that shifts slowly tolerates a very different staleness than a cloud API sweep that should turn over quickly. A timestamp only becomes a signal once it is measured against a defined threshold for that specific source type. Without the threshold, you are reading a clock with no idea what time is late.

Coverage should be reported as a ratio, not a count. Sources reporting over sources configured. "Six of eight sources contributing" tells a reviewer far more than "271 agents discovered," because it exposes the shape of what you don't know. The headline agent count can only go up or stay flat, which is exactly the behavior a dead sensor produces. The ratio moves the moment a source drops, and it moves in a direction someone will notice.

Sensor failures need an owner, and it usually isn't the governance team. An expired credential in an endpoint security platform is fixed by the team that owns that platform. A narrowed role in a cloud account is fixed by whoever administers that account. Governance detects the failure and routes it. Someone else resolves it. This is the part most programs get wrong, because they assume the team that watches the inventory is also the team that can repair the pipes feeding it. It usually isn't, and leaving that handoff undefined is how a flagged sensor sits broken for a quarter. Assigning an accountable owner is the same discipline that turns a discovered inventory into a governable application, applied one layer down to the sensors themselves.

What to instrument

Sensor health is not abstract. It reduces to a short list of things you can measure and alert on.

  • Per-source last successful run. Not "last attempt." Last time the source actually returned data, timestamped and visible.
  • Per-source finding delta over time. A sudden drop to zero from a source that was returning findings yesterday is a signal, even when the total count barely moves.
  • A defined staleness threshold per source class. Set what too old means for endpoint, for cloud API, for telemetry ingestion. Alert when a source crosses it.
  • An alert when a source returns zero where it previously returned findings. This catches the narrowed-permissions and expired-credential cases that a flat total hides. Telemetry-based sources are worth watching most closely here, since a quiet telemetry stream is ambiguous by default and needs a defined expectation to interpret.
  • A coverage ratio surfaced wherever the inventory count is surfaced. If the dashboard shows the number of discovered agents, it shows the fraction of configured sources reporting right next to it.

The rule worth stating directly: never show a discovery count without showing the coverage behind it. A number on its own invites trust it hasn't earned. A number paired with "six of eight sources reporting, two stale past threshold" tells the reviewer precisely how much the count is worth and where to look before they rely on it.

Govern the sensors, not just the agents

Discovery is only as trustworthy as the sensors feeding it, and sensors fail in ways that never announce themselves. Treating their health as a first-class governance control, with per-source freshness, coverage ratios, defined staleness thresholds, and a clear owner for every failure, is what keeps your inventory honest. Everything you build on top of discovery, from triage to policy to what you tell an auditor, inherits the reliability of the layer underneath it.

See how Arthur reports sensor health and coverage alongside every discovery finding. Book a demo or explore the Agent Development Toolkit to get started.

‍

SHARE