When we say we want AI "aligned with human values," a deep question lurks: are there objective moral truths to align to, or just preferences?
The two camps
| Moral realism | Anti-realism |
|---|---|
| Moral facts exist independently of us | Morality is human attitude/convention |
| "Torture is wrong" is true like "2+2=4" | "Wrong" expresses disapproval, not a fact |
| A smart enough AI could discover ethics | There's nothing out there to discover |
Why it's load-bearing for AI
- If realism is true: a superintelligence might figure out the correct ethics on its own — but the is-ought gap suggests intelligence alone won't get it there.
- If anti-realism is true: there is no "correct" alignment target — only somebody's values, making whose values unavoidable and CEV a way to pick.
"Align AI with morality" assumes there's a morality to align with. Whether that's a discovery or a choice changes the entire alignment project.
Related: The Is-Ought Problem · The Value Loading Problem · Coherent Extrapolated Volition