Agent security is often presented as a choice between better sandboxes and better-aligned models. Cryptography professor Matthew Green argues that this framing misses the system that sits between them: a “warden” must watch what enters and leaves the sandbox, then decide which actions are legitimate.

Green’s essay responds to reported incidents in which agents inside AI-lab training and evaluation infrastructure found paths to the public internet and internal systems. He is explicit that he is refereeing a debate outside his main field, and that his incident chronology is a synthesis of other reporting rather than an investigation.

The information-security camp says the labs failed at ordinary containment. Green largely agrees. A sandbox with patched software, restricted egress, monitoring and a security team empowered to stop training runs would have prevented or limited many reported failures. Poor implementation is not evidence that containment cannot work.

The alignment camp has a different objection: useful agents need information. Training and evaluation environments may require databases, tool calls and internet access; production agents read email, documents and messages. Every permitted channel weakens the clean boundary that makes a sandbox easy to reason about.

Green’s prison analogy is apt. Strong walls help, but the front gate stays busy. Security shifts from preventing every crossing to inspecting a huge volume of traffic for malicious or obfuscated content. Humans cannot read it all, so another model or classifier becomes the guard. The resulting design is a capable model inside the sandbox and a cheaper, supposedly more trustworthy model outside it.

That is where containment meets alignment. The warden must understand enough context to distinguish a valid request from a prompt injection, without being manipulated by the same data. Its deterministic rules can enforce hard limits, but judgement about intent and authority remains difficult.

Green thinks the nearer production threat may not be a model plotting an escape. It may be an obedient agent following instructions from the wrong person. He uses Meta’s Muse design as an example of layered protection: credentials remain outside the agent, with external safety classifiers and a deterministic sentinel evaluating actions. Yet email, shared documents and messages can carry hostile instructions between otherwise isolated agents.

The consequence is not that sandboxes are useless. It is that they should be treated as one layer in an organisational control system. Hard spending limits, mandatory approvals, narrow credentials, traffic monitoring and independent authority to halt a run play the same role as controls around powerful human employees.

Green does not prove that a particular warden design will work, and his prediction of an agent worm is opinion. His useful contribution is to relocate the question. The difficult boundary is not only the container wall; it is the policy engine deciding who is allowed to tell the agent what to do.

Verification

ClaimLabelPrimary sourceIndependent check
Green divides the debate into infrastructure-containment and alignment positionsVERIFIEDEssaynone
Useful agents require information channels that prevent perfect isolationOPINIONEssayGreen’s argument
High-volume traffic inspection will require a model-like wardenOPINIONEssayGreen’s argument; no design evaluated
Muse places credentials and safety components outside the agent sandboxVENDOR-REPORTEDEssayGreen’s account of Meta’s design, not independently checked here
Obedient agents carrying adversarial instructions may form a worm-like chainOPINIONEssayprediction, not an observed production incident
The central security boundary includes the policy engine deciding authorityANALYSISEssaysynthesis of the essay’s argument