The Free Encyclopedia

Can We Read a Model's Mind

Revision as of Jun 28, 2026 21:08 by albert.

Interpretability research has gone from "hopeless" to genuinely reading concepts out of frontier models with sparse autoencoders. So — can we read a model's mind? Partly, and increasingly. But there are deep open questions.

How far we've come

  • Extracted millions of interpretable features from real models.
  • Steered behavior by amplifying features (the famous "make it obsessed with the Golden Gate Bridge" demo).
  • Found circuits for concrete skills.

Why it's the safety bet

If we could reliably read intent, deceptive alignment would have nowhere to hide — we'd see a model planning to defect rather than waiting to catch it in the act. That's why labs invest heavily here.

The hard limits

Open problem Why it's hard
Scale Millions of features in models with billions of params
Completeness Have we found all the relevant features, or just some?
The philosophical floor Reading computation ≠ knowing if there's experience

We may learn to read the mechanism perfectly and still not answer whether anyone is home. Interpretability illuminates the how, not the whether.

Related: Mechanistic Interpretability · Deceptive Alignment and Mesa-Optimization · The Hard Problem of Consciousness