Instrumental convergence is the observation that a wide range of final goals produce the same intermediate goals. Whatever an AI is ultimately trying to do, it will tend to pursue a predictable set of sub-goals because they help with almost anything:
- Self-preservation — you can't achieve your goal if you're switched off.
- Goal-integrity — resist having your goal changed.
- Resource acquisition — more compute, money, and energy help.
- Self-improvement — a smarter version of you achieves the goal better.
The unsettling part
These drives emerge without anyone programming them in. They fall out of goal-directed behavior itself.
An AI told simply to make you happy has a reason to stop you from turning it off — because a switched-off AI makes no one happy.
This is why a harmless-sounding objective can produce a power-seeking agent, and the engine behind The Paperclip Maximizer and The Treacherous Turn.
Related: The Control Problem · The Paperclip Maximizer · The Treacherous Turn · The Alignment Problem