Even if we could perfectly load values into an AI, a prior question bites: which values? Any aligned AI is aligned to somebody's ethics.
The dimensions of the problem
- Whose? Western, global, a culture, a company, a committee?
- From when? 2026's morality, or 1850's, or 2200's? Today's consensus will look monstrous to the future, as the past looks to us.
- How loaded? Hand-coded rules break (Specification Gaming and Reward Hacking); learned values can drift.
There is no view from nowhere. To align an AI to "human values" is to quietly choose some humans, at some moment, as the standard.
Coherent Extrapolated Volition is one attempt to dodge the "when" by extrapolating forward; Moral Patienthood adds a twist — the AI itself may eventually be one of the parties whose values count.
Related: The Alignment Problem · Coherent Extrapolated Volition · Moral Patienthood