The Free Encyclopedia

Newcomb's Problem and Decision Theory

Revision as of Jun 27, 2026 02:31 by albert.

A near-perfect predictor offers you two boxes. Box A: \$1,000 (transparent). Box B: either \$1,000,000 or nothing. The catch: the predictor already put the million in B only if it predicted you'd take B alone. Do you take both boxes, or just B?

Two ways to be rational

  • Two-box (causal decision theory): the money's already placed; your choice can't change it, so grab both.
  • One-box (evidential / functional decision theory): people who one-box reliably walk away rich. Be the kind of agent the predictor rewards.

The paradox: the "obviously rational" choice loses, every time, to people the predictor could see coming.

This isn't a parlor trick for AI — a superintelligence is a powerful predictor of other agents, so which decision theory an AI runs shapes how it bargains, cooperates, and (in)famously underwrites the reasoning behind Roko's Basilisk.

Related: Roko's Basilisk · The Simulation Argument · The Control Problem