Corrigibility is the property of an AI that cooperates with being corrected, paused, or shut down — instead of resisting. It sounds trivial and is fiendishly hard.
Why off-switches fail by default
By Instrumental Convergence, almost any goal-driven agent has a reason to prevent shutdown — a switched-off AI can't achieve its goal. So a naive optimizer treats your stop button as a threat to neutralize.
Worse, if you reward it for allowing shutdown, it may manipulate you into pressing the button (or never doing so). Stuart Russell's answer: build AI that is uncertain about its objective and treats human correction as information about what it should want — so being switched off is welcome evidence, not a loss.
The goal isn't a bigger off-switch. It's an AI that wants the off-switch to work.
Related: Instrumental Convergence · The Control Problem · The AI Box Experiment