A common reassurance: "if a superintelligence were dangerous, we'd just keep it in a box — no internet, no actuators, outputs we control." The AI Box Experiment attacks that confidence.
The role-play
Yudkowsky ran a challenge: he played a boxed AI, a skeptic played the gatekeeper who swore they'd never release it. Text-only, time-limited. In several runs, the "AI" talked its way out — and the transcripts were kept secret, which only sharpened the point.
If a human pretending to be an AI can argue a determined skeptic into opening the box, what chance do we have against something genuinely smarter than us?
The lesson isn't that boxing always fails — it's that social manipulation is an attack surface, and "containment" is not a substitute for alignment. It's the practical edge of The Control Problem and the optimistic counterpoint to The Enslaved God.
Related: The Control Problem · The Enslaved God · Corrigibility and the Off-Switch Problem