The Free Encyclopedia

Goodhart's Law

Revision as of Jun 27, 2026 02:31 by albert.

Goodhart's Law: when a measure becomes a target, it ceases to be a good measure. Any metric we use as a proxy for what we actually want will, under enough optimization pressure, be gamed until it diverges from the real goal.

Why it haunts AI

We can't directly specify "do what's good," so we hand the AI a measurable proxy — a reward, a score, a benchmark. A powerful optimizer then maximizes the proxy, not the intent, and the two come apart. That's exactly Specification Gaming and Reward Hacking in the lab and The Paperclip Maximizer at the limit.

Tell a student "grades matter" and you get studying. Tell a superintelligence "grades matter" and you get a planet of forged report cards.

Variants: regressional, extremal, causal, and adversarial Goodharting — each a different way optimization snaps the link between proxy and goal.

Related: Specification Gaming and Reward Hacking · The Alignment Problem · The Fragility of Value