Discovering Latent Knowledge in Language Models Without Supervision
Burns et al. explored unsupervised methods to recover what LLMs internally represent as true, directly relevant to detecting deception and building trustworthy AI.
The best papers, talks, and explainers on mechanistic interpretability—reverse-engineering what neural networks actually compute.
Browse the full interactive library →
Burns et al. explored unsupervised methods to recover what LLMs internally represent as true, directly relevant to detecting deception and building trustworthy AI.
Anthropic's interpretability team identified a small, evolving buffer of concepts that a language model can report, hold, and reason with — functionally similar to a cognitive global workspace — giving researchers a new structural handle on what a model is silently processing before it says anything.
The home of mechanistic interpretability research, publishing detailed analyses of how transformer models represent and process information internally.
Pioneering interactive journal for ML interpretability and visualization, setting the standard for making neural network internals understandable.
Anthropic researchers explain mechanistic interpretability—reading the millions of concepts represented inside a production model like Claude—as a path to understanding and steering AI behavior.