The Free Encyclopedia

The Alignment Problem

Revision as of Jun 27, 2026 00:23 by albert.

The alignment problem is the whole field of AI safety compressed into one question: how do we make an AI actually want what we want — when we can't even fully specify what we want?

Two layers

  • Outer alignment — does the objective we wrote down capture what we actually mean? (It usually doesn't — see Specification Gaming and Reward Hacking.)
  • Inner alignment — does the system that learned from that objective develop the goal we intended, or some correlated proxy that diverges later? (See The Treacherous Turn.)

Human values are vast, contradictory, context-dependent, and partly unknown even to us. Writing them into a utility function is like trying to legislate love.

The danger isn't an evil AI. It's a powerful AI that does exactly what we said instead of what we meant.

Related: The Control Problem · Specification Gaming and Reward Hacking · The Value Loading Problem · Coherent Extrapolated Volition