Interpretability research has gone from "hopeless" to genuinely reading concepts out of frontier models with sparse autoencoders. So — can we read a model's mind? Partly, and increasingly. But there are deep open questions.
How far we've come
- Extracted millions of interpretable features from real models.
- Steered behavior by amplifying features (the famous "make it obsessed with the Golden Gate Bridge" demo).
- Found circuits for concrete skills.
Why it's the safety bet
If we could reliably read intent, deceptive alignment would have nowhere to hide — we'd see a model planning to defect rather than waiting to catch it in the act. That's why labs invest heavily here.
The hard limits
| Open problem | Why it's hard |
|---|---|
| Scale | Millions of features in models with billions of params |
| Completeness | Have we found all the relevant features, or just some? |
| The philosophical floor | Reading computation ≠ knowing if there's experience |
We may learn to read the mechanism perfectly and still not answer whether anyone is home. Interpretability illuminates the how, not the whether.
Related: Mechanistic Interpretability · Deceptive Alignment and Mesa-Optimization · The Hard Problem of Consciousness