Can You Audit an AI System Without Access to the Model?
Yes, if what you are auditing is what the system did. An external auditor cannot evaluate model weights, which are proprietary and not what auditing examines. What is auditable from outside is the action boundary: each decision an agent attempted, the policy applied, the outcome, and who authorized it. Model access is not required for that record.
The question has become concrete rather than theoretical. Several companies that train frontier models have publicly committed to partnering with independent external auditors, and that commitment has a consequence nobody has written down: an external party cannot be given weights, so whatever satisfies it has to be readable from outside the model.
This page is about what that record is. It applies equally to an auditor reviewing a frontier lab and to a security reviewer who has to sign off on an agent inside an ordinary company, because the constraint is the same in both cases.
What is auditable from outside the model?
Two things get confused under the word audit, and separating them answers most of the question.
Auditing the model means forming a judgment about a system's propensities: what it tends to do, what it has learned, where it fails. That work needs weights, training data or at minimum deep sampling access. It is research, it is expensive, and an external auditor is not equipped to do it even when permitted.
Auditing the deployment means forming a judgment about events: what was attempted, what was permitted, what happened. That needs no model access at all, because the evidence is generated at the boundary where the system meets the world rather than inside it.
Almost every question an auditor is actually accountable for is the second kind. Did this agent exceed what it was granted. Was there an approval. What happened after the refusal. Who widened the grant, and when. None of those are answerable from weights, and all of them are answerable from a record of decisions.
Why is a training-time evaluation not audit evidence?
Evaluations are genuinely useful and they are not evidence, for three structural reasons.
They are about a model, not about your deployment. An evaluation result describes behavior on a corpus. An auditor is asking about an action that occurred in a specific tenancy under a specific grant, which no corpus contains.
They do not accumulate. One evaluation is produced per model version. Audit evidence accrues continuously, because the thing being evidenced is a stream of events rather than a property.
They cannot be re-performed by the auditor. An auditor's standard move is to select a sample and verify it independently. A score on a corpus the auditor cannot run against a model they cannot access is a representation, not a test.
Does a log of blocked actions prove the control is working?
No, and this is the failure that is hardest to see because it looks exactly like success.
A control that has quietly stopped running produces an empty refusal log. A control that is running correctly in a well-behaved month also produces an empty refusal log. From the artifact alone the two are indistinguishable, and the second is the story everyone prefers to believe.
What separates them is recording the decisions where the control allowedsomething. An allow is evidence that the control was consulted and answered. Without allows in the record there is no denominator, and a refusal count over an unknown denominator supports no conclusion at all.
This is not a novel position in policy engineering. Open Policy Agent's decision log records the policy queried, the input, and the result, whatever that result was, which is what makes the log reviewable rather than merely alarming.
Where does independence actually come from?
Not from who owns the system, because the party running the agent is the only party present when the action happens. Self-produced evidence is unavoidable. The question is what makes it reviewable anyway.
Independence comes from the separation between the point of decision and the store of record. If the component being audited can also revise its own record after the fact, the record documents whatever that component last decided it should say. If the store is append-only, or tamper-evident, or simply outside the reach of the thing it describes, the same evidence becomes reviewable.
The FINOS framework treats this as a tier rather than a binary, which is the right shape: AIR-DET-021 rises from timestamps and tool invocations, through logged reasoning ahead of tool calls, to cryptographic protection and tamper-evident logging with decision chains tracked across multi-step and cross-agent processes. A programme can sit at a lower tier honestly. What it cannot do is claim a higher one.
Why the external-auditor requirement constrains everything above it
Reporting on the White House Accord on Super Intelligence describes four layers for its signatories: internal controls monitoring capability and alignment during training and deployment, an internal team verifying those controls, independent external auditors, and an independent board committee reviewing both. The accord is voluntary, binds the companies that signed it, and creates no obligation for an organization deploying agents built on someone else's model. The wording is summarized from the reporting cited below rather than quoted, for the reason set out on the SI definition page.
Read as engineering, the third layer is the binding constraint on the other three. If an external auditor must be able to review the evidence, then the evidence cannot be model-internal, which means layers one and two have to emit something at the boundary whether or not they also inspect the model. And layer four, a board committee, needs that same material summarized for people who are accountable without being practitioners.
The ordering is worth noticing. A programme designed inward from the model satisfies layer one and then discovers it has nothing layer three can read. A programme designed outward from the action boundary satisfies three and four first, and finds that one and two were mostly about recording what it already decided.
Six things an auditor can ask for today
None of these require model access, and a programme that cannot produce them does not become auditable by adding model access.
- The declared authority each agent held: which tools, for what purpose, for how long, granted by whom.
- The decision record for actions attempted, including allowed ones.
- A participation count per control, so a silent control is distinguishable from a clean one.
- The policy version in force at the time of each decision, not just the current one.
- What happened after a refusal: whether the work was abandoned, escalated, approved, or retried by another route.
- Evidence the record cannot be altered by the system it describes.
Number four catches more programmes than the others combined. Evidence of a decision is only interpretable against the rule that was live when it was made, and policy stores that keep only the current version silently make every historical decision unreviewable.
Related reading
- What is Super Intelligence (SI)? The federal rename, and the accord the four layers come from.
- What do agent benchmarks actually measure? Why a score is not a record.
- What is action governance? The boundary this page says the evidence comes from.
- Proving governance to an auditor: the same question answered in a compliance framework's own language.
Sources
- FINOS AI Governance Framework, AIR-DET-021 Agent Decision Audit and Explainability. Four tiers of decision logging, rising to tamper-evident records and decision-chain tracking across agents
- FINOS AI Governance Framework, MI-18 agent authority and least privilege
- Open Policy Agent, Decision Logs. Documents the per-decision record: the policy queried, the input, the result, and metadata enabling audit
- AgentSpec: runtime enforcement of user-defined constraints on LLM agents (ICSE 2026)
- Executive Order, "Inaugurating The Era Of Super Intelligence," signed September 29, 2026
- Nextgov/FCW, "White House unveils super intelligence executive order and industry accord." Reporting on the accord's four layers
Shrike records a decision per action at the governance boundary, allows included, with the grant that was in force and the control that answered. That record is what the six questions above are asking for, and it is readable without access to the model. See action governance for what the layer does, or the quickstart to run it against your own agent.