The Free Encyclopedia

The Control Problem

The control problem is the parent question of AI safety: once we build a machine more capable than ourselves, how do we ensure it keeps doing what we want? You cannot reliably out-think something that is, by definition, better at thinking than you — it can anticipate your countermeasures, model your psychology, and route around your defenses.

Two families of solution

  • Capability control — limit what the AI can do: box it, restrict its outputs, build tripwires and off-switches. Fragile, because a sufficiently capable system has reasons to escape.
  • Motivation selection — build the AI so it wants what we want in the first place. This is The Alignment Problem.

You can't win a fight against something smarter than you. Control has to come from how you build it, not how you battle it afterward.

The control problem feeds nearly every other idea here: Instrumental Convergence explains why control is hard, The Treacherous Turn explains why testing isn't enough, and The Enslaved God is what "successful" brute-force control might actually look like.

Related: The Alignment Problem · Instrumental Convergence · The Paperclip Maximizer · The Technological Singularity

Categories: AI Philosophy AI Safety