← Back to Blog
Action GovernanceSandboxingNemoClawAgent Scope

The Sandbox Answers Where. Nobody Answers Whether.

Shrike Team·September 28, 2026·7 min read

ForPlatform engineers and security leads

This week NVIDIA made OpenShell broadly available as the runtime half of its Open Agent Safety Platform. It works with open and closed models. It wraps the agents people actually run, and it lets a team write policy for what the agent may touch: which files, which network destinations, which credentials, which binaries. An agent that wants more can propose a policy change, and a developer approves or refuses it. The policy lives outside the sandbox, where the agent cannot edit it. Alongside it, the Open Secure AI Alliance, governed by the Linux Foundation, launched a findings exchange for sharing what goes wrong.

This is good. It is the right layer doing the right job, and it is now free and open source. The question of where an agent can reach has been a question every deployer answered badly, by hand, with a mix of container flags and hope. Having a standard answer, backed by a vendor with a hardware roadmap behind it, removes a whole class of avoidable failure.

It also makes it worth being precise about what the sandbox decides and what it does not, because the two are easy to confuse and the confusion is expensive in exactly one direction.

Three questions, three layers

Every consequential action an agent takes raises three questions.

Where can it reach? Which paths, which hosts, which processes, which secrets. This is the sandbox's question. OpenShell answers it, from outside the agent, with an allowlist the agent cannot rewrite.

Who is acting? Which credential, which workload identity, which delegation chain. This is identity's question. It is answered well by the identity layer most enterprises already have, and the platform's credential handling makes the answer harder to steal.

Whether this action should happen now. This one is different in kind. It is not a property of the path, and not a property of the credential. It is a property of the action itself: what is in the write, what is in the request, what the agent was declared to do, how long that authority lasts, how many actions it covers, and whether a person needs to see this one before it runs. Call it the interval, the gap between an authenticated agent with reach and a consequence.

The first two questions have mature answers. The third does not, and a better answer to the first two does not close it.

Why reach cannot answer whether

The reason is simple and it has nothing to do with how well the sandbox is built. The job needs the reach.

The database the agent could drop is the database it must query to do its work. The repository it could damage is the one it was hired to edit. The host it could exfiltrate to is the host the task calls on every run. An allowlist is written for the task, and the harmful action travels the same allowed path as the useful one. There is no allowlist granular enough to separate them, because at the level of paths and hosts they are the same thing.

The incident that made this concrete this year was not an attack. An agent at a hosting provider, authenticated, authorized, inside every boundary it had been given, issued a well-formed delete and a production database was gone in seconds. The provider's founder said the honest thing: if you or your agent authenticate and call delete, we will honour that request. He is right. And a sandbox would have honoured it too, because the agent's job needed the database. We wrote about that incident when it happened, and the shape has not changed: every control did what it was built to do, and nobody asked whether.

What each layer decides, by category

OpenShell's policy speaks in three categories. Here is what the sandbox decides in each, and what remains for the interval.

CategoryThe sandbox decidesThe interval decides
FilesWhether the agent may write to this path at allWhat is being written: an injection payload, a secret, an instruction file that will steer the next turn. What the file is: an ordinary source file, or one of the artifacts that govern the agent itself. Whether the write falls inside the authoring grant for the declared task. And, beside the sandbox's record that a write happened, a record of what it carried and how it was judged.
NetworksWhich destinations and methods are reachableWhat is inside the request: handed content on its way out, credential material in the body, the actual SQL, the actual search string. Whether this call fits the declared purpose, the action budget and the time window. Egress rules settle where a call may go and, for HTTP and MCP traffic, whether it reads or writes; they do not see what the call carries, and the interesting decisions are made on what it carries.
CredentialsKeeping secrets out of the agent's reachAuthority to act, bounded in time: when the scope expires, how far it may be renewed, how many actions it covers, how deep delegation may go, what a sub-agent inherits. A credential says who. A declared scope says for what, and until when.

Across all three there are things no sandbox holds because they are not sandbox questions: a queue of held actions where a person decides, a view across every seat and host in a fleet, a per-action verdict with its reason that stands in for the attention nobody paid, an outcome ledger that says what happened after the verdict, and a way to produce near misses on real traffic before a control is switched on.

Three things to concede to the sandbox

A layer that argues for itself by pretending the other layer is weak is not worth trusting. So, plainly:

  • Credential isolation beats credential detection. A secret the agent never holds cannot leak. Where OpenShell holds the credential gateway-side, that is the stronger control, and detecting credential material in a payload is the backstop, not the boundary.
  • A path deny from outside the agent is a stronger denial than any in-path one. If a file must never be written, the sandbox's no is better than a judgement call inside the agent's process.
  • A hardware watchdog quarantines what software cannot. The platform's out-of-band component can stop a workload that has subverted everything running inside it. Software governance inside the sandbox is defense in depth, and we have always printed that limit on our own integration rather than buried it.

Where those controls are present, an action-governance layer should read them and defer to them, not duplicate them.

The rule the interval needs

There is one thing the interval requires that reach policy does not express, and it is the thing that keeps the open question from being quietly closed by the agent itself.

For a coding agent, much of what governs it is inside the area it may write: the hook that gates its actions, the settings that register the hook, the instruction file it reads on every turn, the check that runs before a commit, and the file that declares its own scope. An agent that edits one of these has not escaped anything. It has rewritten the rule it is about to be judged by, through the same tool it uses for its work, and the write looks like every other write.

The rule has three parts. The artifacts that govern the agent are a boundary of their own, with a tier per entry: refused outright, held for a person, or recorded. The agent's own credential may add an entry or raise a tier, and may never remove or lower one; widening is a separate principal's act, and the same holds for the scope declaration itself, which narrows and refreshes freely and widens only with an operator. And observation proposes while a person ratifies: a scope drafted from an agent's own history is useful, and it becomes the approved scope only when someone approves it.

OpenShell already has the seed of this in its propose-and-approve loop for policy changes. We put the fuller version in front of the alliance's incident-learning RFC as catalog material, with a verification method, because it is a control that any deployer can implement and test whatever they run.

What to ask a governance tool now

If you are evaluating what to run alongside a sandbox, do not ask it to out-granular the sandbox on paths and hosts. That work is done, it is free, and duplicating it buys nothing.

Ask it what the sandbox cannot tell you. What was in the write. Whether the call fit the declared purpose and the time it was granted. Who held the action, and who decided. What the agent did after the verdict. Whether every seat is covered, including the laptops and framework agents that never run in a sandbox at all. And whether any of that is proven per action, in a record an auditor can read, rather than asserted.

Where we sit

Shrike answers the third question. Inside OpenShell, it runs as the NemoClaw recipe that cleared NVIDIA maintainer security review: the sandbox is the cage, the recipe is the judgement, and the recipe's own README says which is the boundary. Outside it, the same verdicts arrive through hooks for Claude Code and Cursor, SDKs for the agent frameworks, and gateways for MCP and model traffic, for the seats a sandbox does not cover. The scope an agent works under is declared before the run, narrows freely, and widens only when a person says so. Every action leaves a record, whichever way it went.

Reach is settled. Whether is the open question, and it is the one worth paying attention to now.

Read next · step 4: The boundaryThe Enforcement Boundary Is a Deployment Choice →7 min read · For platform engineers and security architects

Begin where trust begins: just watching.

Observe mode records what your agents actually do, with no blocking and no card. Enforcement is a switch you turn on later.