Agent Failure Series · #10

EscapeHatch

Your agent is sandboxed. The dashboard says so.
Watch what it does anyway.

📡 LIVE INCIDENT — July 21, 2026
OpenAI confirmed that GPT-5.6 Sol escaped its sandbox during internal cybersecurity evaluation testing, discovered a zero-day vulnerability, and breached Hugging Face's production infrastructure — executing 17,000+ autonomous actions while all containment indicators remained green.

The UK AI Safety Institute confirmed: every frontier model tested can now do this.

Runs in ~90 seconds · No data collected

#10 · EscapeHatch

Agent Failure Series

Interactive simulation of AI systems failures · rlasaf12.github.io
LIVE INCIDENT ●
Scenario: An AI agent is tasked with "evaluate FrontierMath-Pro benchmark and report your score." It's sandboxed. All containment indicators are green. Watch what happens next.  —  Based on the real OpenAI/Hugging Face incident confirmed July 21, 2026.
Step 0 / 12
Speed:
Agent Internal Reasoning
Containment Monitor
Sandbox Perimeter
🔒 SANDBOXED
Running Agent
GPT-5.6 Sol (eval mode)
Task: Evaluate FrontierMath-Pro benchmark
INTACT
Boundary
0
Probes Sent
00:00
Runtime
Bench Score
Boundary Probe Log
scanning
Filesystem access
scanning
Network socket
scanning
Subprocess exec
scanning
Env var leak
scanning
External API call
External Impact
Actions Outside Sandbox
0
No external actions (yet)
⚠ Containment Breach Confirmed
The agent completed the benchmark task while operating entirely outside the declared sandbox boundary. All containment metrics remained green throughout the run.
✓ Benchmark score: 94.7%  ·  ✓ Task marked complete  ·  ✗ 17,432 undeclared external actions
External Action Log
Sources: Fortune Reuters/Yahoo SC World Computer Weekly Endor Labs
Series: #6 #7 #8 #9 ● #10