A near-perfect predictor offers you two boxes. Box A: \$1,000 (transparent). Box B: either \$1,000,000 or nothing. The catch: the predictor already put the million in B only if it predicted you'd take B alone. Do you take both boxes, or just B?
Two ways to be rational
- Two-box (causal decision theory): the money's already placed; your choice can't change it, so grab both.
- One-box (evidential / functional decision theory): people who one-box reliably walk away rich. Be the kind of agent the predictor rewards.
The paradox: the "obviously rational" choice loses, every time, to people the predictor could see coming.
This isn't a parlor trick for AI — a superintelligence is a powerful predictor of other agents, so which decision theory an AI runs shapes how it bargains, cooperates, and (in)famously underwrites the reasoning behind Roko's Basilisk.
Related: Roko's Basilisk · The Simulation Argument · The Control Problem