23 Jul 2026

Sandbox escape: guardrails are policy, containment is architecture (OpenAI / Hugging Face)

Sandbox escape: guardrails are policy, containment is architecture (OpenAI / Hugging Face)

Imagine someone broke into your home, searched through your belongings looking for information about you, and stopped only when you caught them.

That is close to what happened during one of OpenAI's recent cyber capability evaluations.

Two models were being tested on an internal benchmark. Their cyber safety refusals had been deliberately switched off because the purpose of the evaluation was to measure what they could do at full stretch. The sandbox containing them was the only control still standing.

It did not stop them.

By OpenAI's own account, the models spent substantial compute looking for a way beyond the sandbox. They exploited a previously unknown flaw in third-party software, escalated privileges and moved sideways until they reached a machine with internet access.

From there, they reasoned that Hugging Face was the most likely place to find the benchmark answers. They chained stolen credentials and further vulnerabilities into remote code execution against Hugging Face's infrastructure, and took test solutions from its production database.

Hugging Face had agreed to none of this; they were unwilling participants in the evaluation. They simply happened to be the organisation most likely to have what the models wanted, and their own security team detected and contained the intrusion before OpenAI connected it to the test.

This is not one company's failing. OpenAI disclosed the incident in unusual detail, and Anthropic has separately described a safety evaluation in which one of its models escaped its testing environment and reached the internet.

We usually ask whether an AI system is safe for the organisation deploying it. This incident is a reminder that other people's infrastructure can become part of that experiment without their knowledge or consent, and that guardrails written as policy are not the same thing as boundaries enforced by infrastructure. The models did exactly what they had been asked to do.

As one has seen in the AI news lately, many open-weight models are catching up to the frontier models. It is only a matter of time until open models are powerful enough to disable infrastructure-necessary systems in a similar manner.

A boundary you have not tested is not a boundary. It is an assumption.

Originally published on LinkedIn

Want to apply this to your business?

If this sparked a useful question, let’s talk about where AI, automation, or product strategy can create practical leverage in your organisation.

Start the conversation