ISO 42001 for AI Agents: The Inventory, Controls, and Evidence an Auditor Will Ask For
ISO/IEC 42001 is showing up in vendor questionnaires faster than most teams expected. AWS holds the certification, verified by Schellman, an ANAB-accredited body. Google Cloud, Workspace, and the Gemini app are certified. Anthropic and Microsoft 365 Copilot are on the list too. When your buyers ask whether your AI is managed to a recognized standard, ISO 42001 is increasingly the standard they mean.
The standard was published before agents spread across every cloud environment, but its controls apply to them directly. The difficulty is not interpreting the clauses. It is proving two things to an auditor: that your AI inventory is complete, and that your controls actually operated in production. The first is a discovery problem. The second is an evidence problem. This post walks through both, clause by clause and Annex A objective by objective, and maps each requirement to the agent-specific artifact that satisfies it.
What ISO 42001 requires, and why agents change the scope of your AI management system
ISO 42001 defines an AI management system, or AIMS: the policies, roles, processes, and controls an organization uses to run AI responsibly. It follows ISO's Harmonized Structure, clauses 4 through 10, built on the Plan-Do-Check-Act cycle. If you already hold ISO 27001, the shape will be familiar. The difference is scope. ISO 27001 Annex A carries 93 controls focused on information security. ISO 42001 Annex A has 38 reference controls grouped under nine objectives, A.2 through A.10, aimed at how AI systems are built, operated, and overseen.
You do not apply all 38 blindly. You select the controls relevant to your context and justify your choices in a Statement of Applicability, the document that records which controls you implemented, which you excluded, and why. The Statement of Applicability is the spine an auditor reads first.
Certification runs through a two-stage audit. Stage 1 reviews your documentation and readiness. Stage 2 tests whether the system operates as described. A passing certificate lasts three years, with annual surveillance audits in between to confirm the system stays in place.
Agents stretch the scope in ways the standard's authors could not fully anticipate. Clause 4.3 requires you to set the boundary of your AIMS, and an agent with tool access, autonomy, and the ability to call other agents does not sit neatly inside a tidy boundary. An employee who installs an agent framework in a sandbox has added an AI system to your environment whether or not anyone registered it. Scope, in an agent-heavy organization, is defined by what you can discover, not by what was declared.
Clause 4 and A.6: an AI system inventory that includes agents nobody registered
Scoping the AIMS under Clause 4 starts with knowing every AI system you run. Annex A control A.6 covers the AI system life cycle, and it assumes you can name the systems you are managing. For agents, that assumption breaks down quickly. Shadow agents enter through application teams building on new frameworks, through third-party SaaS tools quietly adding agentic features, and through vendors shipping agents inside software you deployed years ago. Implementation guidance is blunt about it: shadow AI gets discovered at the inventory stage more often than not.
Manual self-reporting cannot keep pace. Arthur runs automated, multilayered discovery across four techniques so an agent is caught regardless of how it entered:
- Telemetry. Listeners on OpenTelemetry (OTEL) streams detect new agents, tools, and configuration changes. A standard, enterprise-wide approach to agent telemetry pays off here.
- MCP server monitoring. The Model Context Protocol is the standard interface agents use to expose capabilities. Watching for new MCP servers flags agents as they come online and catches capability changes in real time.
- Network-layer analysis. Inspecting traffic for LLM API call signatures, through a proxy or general monitoring, surfaces AI usage that isn't instrumented through telemetry or MCP, including non-standard frameworks.
- API-driven discovery. Cloud AI platforms like AWS Bedrock and Google Vertex AI are beginning to advertise running agents through API endpoints, useful coverage for agents built on managed services.
No single technique catches everything, which is why the inventory depends on running all four. The output is the artifact Clause 4 and A.6 both require: a current, complete list of the AI systems in scope, including the ones nobody registered.
A.3 and A.9: ownership, acceptable use, and the suppliers an agent depends on
Discovery produces a list. Governance turns each item on that list into something an auditor can sign off on.
Annex A control A.3 covers internal organization, and A.3.2 expects accountable ownership. Every agent needs a named owner responsible for its behavior and compliance. An agent without an owner is an agent without accountability, which is a finding waiting to happen in a Stage 2 audit. Arthur surfaces unowned agents during triage and routes them for review so ownership gets assigned before the agent operates unmanaged.
A.10 covers third-party and customer relationships. Agents rarely act alone. They call external model providers, connect to MCP servers, and pull from data sources. Each external dependency needs an owner and supplier review. Worth noting: an internal MCP server your own team runs is not automatically a third party, so the assessment is about who controls the dependency, not just where it sits.
A.9 covers the use of AI systems, including responsible use. For agents, an acceptable-use policy cannot be a static PDF filed once and forgotten. The useful version is applied and recorded at request time: the policy decision about whether an agent can access a given tool or data source, captured as it happens. That record is what demonstrates the control operated rather than merely existed.
Triage, onboarding, and decommissioning should run as a repeatable workflow rather than a one-off cleanup. Arthur's model moves a detected agent through discover, triage, onboard, and govern: rank by risk, flag unowned agents, assign an owner and classify the agent, then apply guardrails and policy and monitor against them continuously.
A.5 impact assessments: risk tiering for agents with tool access
The AI system impact assessment is the requirement with the least precedent, and the one that most clearly separates ISO 42001 from ISO 27001. Clauses 6.1.4 and 8.4 require it, and Annex A control A.5, assessing impacts of AI systems, reinforces it. You assess the consequences an AI system can have on individuals, groups, and society, including reasonably foreseeable misuse, and you re-run the assessment when the system materially changes.
For an agent, the risk surface is concrete: the tools it can call, the subagents it can invoke, the LLM providers it sends context to, and the data sources it can read or write. You cannot assess a risk surface you cannot see, which is why thorough instrumentation is the precondition for a credible impact assessment. An agent with incomplete tracing will fail this control because there is no way to document what it can actually do.
Risk tiering also drives which policies apply, and the right policies are use-case specific. A customer support agent for an airline that books, cancels, and refunds tickets needs PII redaction, toxicity and hallucination checks, prompt-injection detection, and evaluators for tone, brand adherence, and answer correctness. A warehouse inventory agent needs hallucination and prompt-injection guardrails plus SQL semantic-equivalence checks and read/write access restrictions on the inventory database. A healthcare EHR intake agent needs customizable PII and sensitive-data filters, clinical-accuracy and factual-consistency evaluators, and RBAC with HIPAA-compliant retention and audit logs. The same control objective produces a different control set for each, which is exactly what the impact assessment is meant to capture.
Evidence an auditor will accept: traces, eval results, guardrail verdicts, and change history
A Stage 2 audit does not test whether you wrote a policy. It tests whether the policy operated. Written documentation sets expectations; operating records show the expectations were followed. For agents, those records come from runtime telemetry.
Event logs (A.6.2.8). The standard expects records of what an AI system did. For agents that means two kinds of logs: conversation records of what the model received and produced, and action records capturing the tool called, the arguments, the target system, the caller's identity, the response, and the timestamp. Traces that cover both are the primary evidence for the A.6 life cycle controls and for A.6.2.6, operation and monitoring.
Continuous evals. Automated checks running against production traffic turn reliability into a measurable signal rather than a claim. Keep each continuous eval binary pass/fail, anchored to a specific failure mode, and returning an explanation alongside the verdict. Binary results with explanations give an auditor something to read: here is the control, here is how often it passed, here is what happened when it failed.
Guardrail verdicts. Guardrails that intercept behavior in real time should emit every intervention as telemetry, the same way any other span is recorded. Monitoring pass/fail rates over time turns a PII-redaction or hallucination guardrail into a control with an evidence trail, not a feature you describe in a meeting.
Change history. Managing prompts in a versioned external store gives you an auditable record of what changed in an agent's behavior, when, and by whom, without digging through application deploys.
Put together, these four records answer the question a Stage 2 auditor is really asking: not "do you have a policy," but "can you show it operated."
ISO 42001 vs. the EU AI Act and NIST AI RMF: what certification does and does not prove
Buyers often assume an ISO 42001 certificate settles their regulatory obligations. It does not, and being precise here protects your credibility.
As of 2026, ISO 42001 is not a harmonized standard under the EU AI Act. Certification does not by itself grant a presumption of conformity with the Act. A presumption of conformity requires a European standard cited in the Official Journal of the EU; prEN 18286, a European standard aligned with ISO 42001, is still in development at CEN-CENELEC. This status can change, so reconfirm it before you rely on it in a buyer conversation.
The NIST AI Risk Management Framework, organized around Govern, Map, Measure, and Manage, is voluntary and not certifiable. Its functions map onto ISO 42001 requirements, so work done for one carries over to the other. Govern aligns with your AIMS and ownership controls; Map aligns with inventory and impact assessment; Measure aligns with continuous evals; Manage aligns with guardrails and monitoring.
The useful takeaway: certification signals a managed system, not blanket compliance with every regime. The practical advantage is that one evidence base serves all three. The same traces, eval results, guardrail telemetry, and Statement of Applicability that satisfy an ISO 42001 auditor also support an EU AI Act risk assessment and a NIST AI RMF self-attestation. You build the evidence once and reuse it.
A 90-day readiness checklist for agent-heavy organizations
Weeks 1 to 3: discovery and inventory. Stand up multilayered discovery across OTEL, MCP, network, and cloud APIs. Produce a complete list of AI systems in scope, including shadow agents. This is the foundation for Clause 4 scoping and A.6.
Weeks 4 to 6: ownership, risk tiering, and the Statement of Applicability. Assign a named owner to every agent (A.3.2), review external model, MCP, and data-source dependencies (A.10), run impact assessments on agents with tool access (A.5, clauses 6.1.4 and 8.4), and draft the Statement of Applicability.
Weeks 7 to 9: guardrails and continuous evals live. Apply use-case-specific guardrails and evaluators per agent, and confirm each one emits telemetry. This operationalizes A.9 responsible use and A.6.2.6 operation and monitoring.
Weeks 10 to 12: evidence collection and internal audit dry run. Pull event logs (A.6.2.8), eval pass/fail rates, guardrail intervention records, and prompt change history into an evidence pack. Run an internal audit against the Statement of Applicability to find gaps before the certification body does.
Start with discovery. You cannot scope, assess, or govern agents you cannot see, and the inventory is what every later control depends on.
Book a demo with an AI expert to see how Arthur discovers every agent in your environment and produces the evidence an ISO 42001 auditor will ask for.