Best Practices for Building Agents Recap
Arthur

Agent security & governance

Govern what you
cannot see.

Agents are proliferating faster than any team can track, creating security blind spots, compliance risk, and unmanaged sprawl. Arthur discovers every agent in your environment and brings it under one governance framework.

Policies Governance controls enforced across your applications 6 POLICIES ACTIVE
Status Owner Application + POLICY
Only approved apps use restricted tools

Tools that touch sensitive data — PII, customer records, code repositories — are limited to applications that have been approved for them.

1 0
2 applications4 months ago
Agents run within fiscal budget limits

Consumption-billed applications have completed budget review, with monitoring to keep spend inside the approved envelope.

1 1
3 applications4 months ago
Agents touching sensitive data check for PII

Applications reading datasets that contain PII enforce the PII guardrail, and that guardrail is passing consistently.

1 0
2 applications5 months ago
Public-facing apps are protected from prompt injection

Applications exposed to the public internet have the prompt injection guardrail enabled and actively passing.

1 0
3 applications4 months ago
Running workspace compliance checks… 6 of 6 policies · 2 need attention

Discover

Comprehensive visibility into every agent. Anywhere.

Manual self-reporting cannot keep pace with agent proliferation. Arthur runs five sensors in parallel, each one tuned to a different way agents show up.

01 OTel stream detection

Listens to OpenTelemetry streams to catch new agents, new tools, and configuration changes as they happen.

02 MCP server monitoring

Watches for new MCP servers to catch agents as they come online, and flags capability changes in real time.

03 Network-layer analysis

Inspects network traffic for LLM API signatures to catch agents that bypass instrumentation entirely, including non-standard frameworks.

04 Cloud API discovery

Queries AWS Bedrock, Google Vertex AI, and Azure AI Foundry, which advertise their running agents through APIs.

05 Endpoint detection

Scans employee machines for agents running locally, including personal assistants and local models.

The result

A live library of every agent, sanctioned or not.

Discovery runs continuously across every environment: cloud, endpoints, network, and SIEM. The library updates itself as agents appear and change capability, so official and unofficial agents surface in the same place.

Learn more →

Four surfaces agents arrive through

The registered agents you know about are the tip of the iceberg. Agents arrive through four different surfaces, and each one needs a different sensor to find it. That is why there are five sensors and four surfaces: one surface can take more than one.

Arthur · agent discovery
271 agents discovered
needs triage
Endpoint · Jamf Pro 184
Cloud · Bedrock, Vertex 17 + 11
SIEM · idx_ai_egress 7
Evidence Traced Partial Thin
marketing-image-generation-agent GCP · us-central1
reconciliation-agent GCP · europe-west4
OpenClawpersonal agent macOS 15.3 laptop · 8m ago Register
ollama servelocal model macOS 15.3 laptop · 3m ago Register
2 blind spots Gateway and CI pipeline classes have no sensor reporting.
Surface 01 Internally developed agents
Surface 02 Third-party agent builders
Surface 03 SaaS vendors turning on AI
Surface 04 Personal AI assistants

Observe

See inside every agent, not just its output.

Your inventory says an agent exists. Arthur also shows what it is doing: every step it takes, every tool it calls, what each one costs, and how its behavior moves over time.

Trace support-agent-v4 · demo-run-99-q9 4 sub-agents 1 tool call 21.9s end to end
0s5s10s15s20s
AGENTagent run: supportAgent 21.9s
CHAINtemplate prompt: support-websearch 69ms
AGENTagent run: websearchAgent 5.68s
LLMllm: gpt-4o-mini 828 tok · $0.000378 5.67s
TOOLtool: websearchTool retrieval · 3 docs 2.32s
AGENTagent run: draftAgent 3.02s
LLMllm: gpt-4o-mini 2.97s
AGENTagent run: reviewAgent 12.46s
LLMllm: gpt-5.1 12.41s
Eval on this trace groundedness 0.94 tool selection 0.91 PII leakage 0.00 PASSED

For the teams building agents

Inner workings, step by step Reasoning, tool selection, sub-agent hand-offs, and retrieved context, span by span.
Runtime monitoring on live traffic Continuous evals for accuracy, groundedness, and tool selection, attached to real requests.
Token cost and latency Spend and time per span, model, and tool call, so a slow or expensive path is visible.
Data and model drift Input and output distributions tracked against baseline to catch decay before users feel it.

For the teams accountable for them

Cost and usage per application Spend attributed per application, model, and team, so AI cost is a managed line item.
Reporting for board, risk, and compliance Committee-ready reporting on coverage, incidents, and policy performance, drawn from live data.
Configurable alerting into your systems Findings stream into the SIEM, ticketing, and paging tools your teams already run.
Audit trail behind every decision Every eval result, policy check, and enforcement action retained as evidence for examinations.

Govern

Policy that runs, not policy that sits in a document.

Your governance framework becomes enforced rules on live traffic, with an audit trail behind every decision and a way to stop an agent immediately when it goes wrong.

01 Operationalize the framework

Translate your written governance framework into machine-checkable policies, scoped per application, model, and environment.

02 Enforce in real time

Guardrails evaluate every request and response for prompt injection, PII exposure, restricted models, and off-policy tool use, and return a verdict your application enforces inline.

How a policy decision travels

Arthur evaluates every request and returns a verdict. Arthur returns the verdict; your application enforces it. Wire it in via the SDK, or through a gateway or chat interface you already control.

Stage Trigger What happens Where
01 Evaluate Every request and response Scored against your policies for prompt injection, PII, restricted models, and off-policy tool use Validation API
02 Verdict A check fails Returns which claim, entity, or tool call failed and why, with full reasoning Arthur
03 Enforce Your application receives the verdict Blocks, redacts, or re-prompts in your application via the SDK Your app
04 Detect drift Enforcement stops being applied Arthur flags agents where policy is no longer running Continuous
05 Escalate Anomalous behavior or a rogue agent Send an alert notification to stakeholders, owners, or automation with full evidence for fast action Your systems

Deploy securely

Built for the enterprise.

Three deployment options. Fully hosted SaaS, on-prem in your own data center, or a hybrid (also called federated) architecture. RBAC, SSO, and deployed engineering support come with every one. The diagram below shows the hybrid or federated architecture: the evals engine runs next to your workloads, so sensitive data never leaves your environment.

AI applications / data plane

Gen AI applicationsAI modelsAI agents
Evals engine (OSS)

Lightweight, open-source, drop-in — deployed in minutes, next to your AI workloads. Inference data stays local.

Unidirectional access — only anonymized metrics cross

NO SENSITIVE DATA LEAVES

Centralized control plane

DashboardsAlertsManagementAPIsRBAC & SSO

Centralized visibility and governance, hosted by Arthur or by you, with real-time aggregate metrics across every business unit.

For complex enterprises, a single data plane per line of business or cloud account scales across business units without compromising security.

Procurement

Procure through your cloud marketplace.

Arthur is listed on both the AWS and Google Cloud marketplaces, so you can buy through billing and procurement you already have in place.

Amazon Web Services

AWS Marketplace

Deploy the evals engine inside your own AWS account, so inference data stays where it already lives, and govern agentic, generative, and predictive systems from one control layer.

Google Cloud

Google Cloud Marketplace

Add discovery, policy evaluation, and observability to the stack you are already building on Google Cloud, billed through the account your teams use today.