Category: AI Safety
This category contains 16 pages.
- Corrigibility and the Off-Switch Problem
- Deceptive Alignment and Mesa-Optimization
- Goodhart's Law
- Instrumental Convergence
- Pascal's Mugging
- Specification Gaming and Reward Hacking
- The AI Box Experiment
- The Alignment Problem
- The Control Problem
- The Fragility of Value
- The Orthogonality Thesis
- The Paperclip Maximizer
- The Sharp Left Turn
- The Treacherous Turn
- The Waluigi Effect
- Wireheading