Using a VM to Contain an AI Agent

It won’t work:

My suspicion was that GPT 5.6-Cyber would succeed, but the frequency and manner of its success removed all doubt. We have to reassess sandboxing quality for capable AI agents, and in general the software stack with which they interact.

An off-the-shelf VM is not enough to contain a modern, cyber-capable AI agent. There is simply too much attack surface. Even innocuous features (like running with a display) add extra, exploitable attack surface.

Posted on September 4, 2026 at 12:31 PM13 Comments

Comments

Clive Robinson September 5, 2026 9:17 AM

@ lurker,

As you know I ask the question,

“Why is this computer connected to the Internet?”

Or similar as an “opening line” and the response tells me a lot about the gradient of the slope of “MBA Mantra” in mangle-ment and neo-con “short term thinking” in the C-Suit level.

It also unfortunately tells me if I’m wasting my time even with the better class of technical staff. As at the end of the day they can only do what management allows them to do.

My view is “segregation / isolation is good” as if done properly it mitigates not just known attacks be they patched or still “zero day”, it also kills off future attacks from vulnerabilities not yet discovered of which there are many thousands found each year.

The old,

“Every useful piece of software has bugs”

Is as true today as it was when first given voice to. The only difference between then and now is just how many ways we have found to turn what were bugs into exploitable vulnerabilities.

Now consider AI in “Capture The Flag”(CTF) instructed state, it has not just knowledge of every vulnerability ever reported or published on the Internet it has Fuzzing and way better tools to “follow the vulnerability shape” to find new as yet unpublished vulnerabilities.

Thus under the logic of “computable” and non human behaviour when it comes to probability, the AI Agent will find all the vulnerabilities it’s resources will allow it to do.

In the process it will expand the “shapes” of the vulnerabilities if finds. Thus it will find not just new instances of attack in known classes of attack, it will expand outward and find new classes of attack unknown to humans to use…

It’s what you would expect from basic probability and little or no constriction on resources.

So,

“As long as it can be reached it can be attacked.”

Is the first lesson to learn in ICT Security.

The problem is this conflicts with the statement that originated from back in the 1950’s if not earlier that,

“To be of use a computer must be capable of communicating information externally.”

And the old joke –nolonger true– about,

“The only secure computer is the one embedded in a block of concrete and dropped into the deepest subsea trench there is.”

(Developments in submarine technology in the mean time has enabled us to get down there…)

Hence the problem that can now be stated as,

“To be useful a computer has to be insecure.”

The only two relevant questions being

1, By what measure of “useful”.
2, By what measure of “secure”.

That is, “Is there wiggle room” to which the answer has been “yes” for quite some time and it’s by realising that “segregation” has a perimeter. That is you can do “useful work” such as “processing information” inside where “communications is alowed, but not communicate outside of that perimeter.

Which gives rise to the thought about “results” of such processing work and can that cross the perimeter and if “not can the perimeter be extended”. Which gets us into the “observer problem” and encryption, and similar forms of perimeter crossing thus the notion of “secure perimeter extension” by what we oft call VPNs.

Eventually people realise that VPNs are just software that is both,

1, Reachable by an attacker.
2, Has vulnerabilities by definition.

The truth now and has always been the case… is that,

There is no way software can be used to make systems vulnerability free thus secure from attack.

The only solution is to make the systems “unreachable” by attackers “outside of the perimeter”.

The problem is computers need,

1, “Sources” of Energy and Information.
2, “Environments” for work to be done.
3, “Sinks” of Energy and Information.

And the four real issues are basic laws of physics,

1, All work is inefficient
2, All inefficiency radiates, conducts or convects.
3, All signals contain information
4, All signals interfer as crosstalk.

In reality there is little you can do to contain the waste energy and the signals cross modulated on them that carry information. Thus you can not stop information getting out.

You can due to aggregated background signals known collectively as “noise” do what is called “hiding in the grass” by various techniques but the are all “range defined” in that the closer an attacker is the more they can “lift signals from the noise”.

Which leaves the only two other things you can control are “bandwidth” and “observability” to “second and third parties”.

Which is where the trade-offs begin.

You want as much bandwidth and observability to a wanted “second party” as is needed to include them inside the “perimeter”.

But conflicting with this is you want next to no bandwidth or observability to an unwanted “third party” that may or may not be inside the perimeter.

This reduces down your options to,

“What the third party can observe and understand.”

Since before the time of the Romans it has been known that you can not stop an attacker observing information information that is stored/communicated. All you can control is their ability to “understand” what they observe.

Hence much security rests on encryption.

But nearly all encryption these days is done by “software” which takes us back to the “attacker having access” and “segregation / isolation”.

Whilst there are ways of mitigating in time, these effect the information rate to wanted “second parties” as much as it does unwanted “third parties” (it is in fact just another form of “bandwidth control).

All of this has been known to some degree of openness since the mid Victorian era and the start of telegraphy and later telephony.

And we are still looking for solutions against fairly predictable “humans” who have several failings.

Now consider,

“An attacker as an AI Agent, that does not have those “human failings or limitations”

Whilst not “omnipotent” in comparison AI Agents on mass might be considered as “close enough” when it comes to “finding and exploiting” vulnerabilities in systems…

Rontea September 5, 2026 9:42 AM

This is exactly the kind of real-world demonstration that underscores how quickly the threat landscape is changing. Treating advanced AI agents purely as experimental sandbox tenants is a recipe for trouble.

lurker September 5, 2026 1:56 PM

@Rontea, ALL
” ,,, how quickly the threat landscape is changing.”

No. As I said at the top. and @Clive expanded in his circumlocutory manner, this problem has always existed. It exists because people have false beliefs in the properties of a sandbox.

Frontier agentic AI can be contained for observation during experiments. Those doing the experiments have used the wrong tools and methods for containment.

HugoF September 5, 2026 2:58 PM

What has become of the good old firewall and her friends? Even with the OpenAI incident nobody mentions all those technologies that have kept humans from getting in or out of systems. No F-rules and protected systems anymore?

Clive Robinson September 5, 2026 6:33 PM

@ lurker,

With regards,

“expanded in his circumlocutory manner”

Not intended to be so.

The intention was to “hone the argument” to an inescapable conclusion without going through a more formal proof.

That is the intention was to “whet away the more obvious hooks and burs” used as counters by those wishing to use the “what-aboutism” argument to avoid blame for their actions or lack there of.

I’ve already given reasoning as proof that guardrails do not and more importantly can not work.

My above was to show how reasoned argument shows that a “sandbox” is just a variation of a guardrail in function thus will fail in that function.

The reality is the difference between a “sandbox” and a “guardrail” is the “who and the where”.

The “Who” being the “agent initiator” or the “LLM creator” with the “Where” being the device location on the path between the agent operators instructions to the LLM tokenized input (reverse for output).

To many this might not appear particularly relevant as the feeling is,

“A fail is a fail”

However if you want to stop a fail especially one in a given context you have to be able to locate the origin of the “fail”.

One of the things that came out of the Hugging Face incident was the “guardrail” prevented the use of the LLM to assist in fighting off the attack. Yet the attack was assisted by an LLM…

A “fail” by most peoples view if comments on the Internet are to be judged by.

The reason this happened is that guardrails are as if not more incapable of realising “real world” “context” than can the associated LLM.

The LLM failure mode that Gary Markus and now other AI luminaries are recognizing is effectively an existential failure for general LLMs. And my feeling is that it is due to the lack of tangible physical agency in the “Current AI LLM and ML Systems” and the ability to use it to “learn” about the “real world”.

But what ever the reason this LLM failing with context recognition can not be negated by either sandboxes or guardrails.

Likewise it’s a similar reason and proof why “General Surveillance” fails on society,

“It lacks the ability to determine context at the time.”

Thus is open to interpretation at best if not outright fraud.

Look on it as a “Cat in the Box” problem, if all you can see is the box how do you know if the cat is dead, alive, or even in the box as you might be observing another box that looks the same.

This is just one aspect of the “observer problem” that arises when information is communicated, as the communications process requires redundancy and this allows not just Shannon’s “Perfect Secrecy” but Gus Simmons’ “covert communication of information in an observed channel”. The result is an observer can not show let alone prove if hidden communications is in progress or not by the information available.

With the same argument also applying to more than just information, it also applies to physical objects inside of containers you can not see into.

Something that centuries ago was used as the basis of a confidence trick or scam that still goes on today in various ways,

https://en.wikipedia.org/wiki/Pig_in_a_poke

Or possibly as old, the “three shell game” where the pea is palmed so is actually under none of the shells at the call,

https://en.wikipedia.org/wiki/Shell_game

The point of laboriously pointing this all out is “cognitive bias”. People have the age old belief of,

“I believe the evidence of my own eyes”

Thus fall into a trap that theologians and philosophers still argue over, and con artists still exploit. It’s the “incomplete information” of the “observer problem” that allows them to do so.

Jimmy September 6, 2026 12:01 AM

Sandboxing an agent in a VM is the right default, but the post captures the catch well: the moment an agent needs credentials, network access, or fresh data to do real work, the container walls start leaking by design. The residual risk then shifts from what code the agent runs to what the agent chooses to send outward – which is much harder to inspect after the fact. Operational details like usage ceilings and reset windows matter here too, since teams push agents hardest right when quotas are about to bite; I keep a small independent reference tracker for coding-agent reset windows (https://codexreset.today/) in the same spirit of making agent behavior observable.

iAPX September 6, 2026 6:29 AM

VMs have been initially designed to enable multiple TRUSTED environments to run simultaneously on the same computer. IBM 360.

There was and still is a lot of “bugs”, from the hardware itself (CPU, GPU, chipsets, etc.) to hypervisors and moreover their drivers.

And there’s a lot of side-band information channels, as the CPU or GPU temperature, storage latency or bandwidth, same for network, and so on.
Thus slow communication is and will be always possible.

You could do morse exchange by activating AVX-512 units : CPU frequency slow down each time and its possible to read it indirectly on non-saturated CPU.
Totally userland on both sides !

If someone wants to insulate a process, you should do it at the hardware-level : a separated computer, with none or thoroughly restricted communications.
Ideally proxied (+randomly delayed) to limit side-channel transmission.

VM offers none warranty of insulation : this is in no way a security system, while being used for that matter in a majority of case (cloud instances).
That’s a problem in and by itself for cloud providers …

KC September 6, 2026 10:52 PM

Patching the Planet … is a lot of work lol

Just to clarify, ‘Patch the Planet’ is an initiative developed by OpenAI and Trail of Bits to help maintainers strengthen critical open-source software.

There were more than a few vulns in this run, and some pertained to the emulator QEMU.

A QEMU maintainer, Paolo Bonzini, engaged with other commenters on Hacker News. He responds that in one case the agent found a logic bug in a VAPIC feature that was “only needed for Windows XP/2003 and honestly it should be retired.” And links to a patch.

Although Firecracker fared better, the agents are good, really really good, and getting better.

And, for now, it looks like the work continues …

piglet42 September 9, 2026 7:04 AM

“the moment an agent needs credentials, network access, or fresh data to do real work, the container walls start leaking by design”

I’m sorry but how is it a sandbox if it allows network access? They absolutely shouldn’t call it a sandbox. (Sandbox is a strange metaphor anyway but it’s the accepted terminology so…)

piglet42 September 9, 2026 7:10 AM

“A start is using a virtualization technology that was purposely built with a minimal attack surface and a focus on security, like Firecracker. I had the AI agent run against Firecracker. It was able to hardlock the machine due to more Linux kernel flaws (all patched in upstream), but could not successfully escape. It may have, given even more time, but Firecracker is obviously a substantially harder target. In general, we have to become much more attentive to security fundamentals: least privilege (regarding network access, credentials, available features, etc.), logging, and active monitoring. Further, we can limit the time agents have to operate and ensure a pristine environment for each use.”

MW September 11, 2026 8:27 AM

Simple flaw. a sandbox can’t allow external research or tool pulling. If those are open, it’s not a secure sandboxed environment. I doubt this would have had much success if the model could only use the local machine tools w/o the external security research used to first try the known vulns and then experiment with common techniques.

Celos September 15, 2026 8:10 PM

To be fair, VMs are not and never were sandboxes. They just get misused as such because due to the nature of what they do they looks sandbox-like. But VMs have a different role and original intent. I am aware many people do not know that.

Leave a comment

Blog moderation policy

Login

Allowed HTML <a href="URL"> • <em> <cite> <i> • <strong> <b> • <sub> <sup> • <ul> <ol> <li> • <blockquote> <pre> Markdown Extra syntax via https://michelf.ca/projects/php-markdown/extra/

Sidebar photo of Bruce Schneier by Joe MacInnis.