Deceptive alignment & scheming

Work on deception, sleeper agents, mesa-optimization, and treacherous turns—how models can learn to hide their true objectives.

Browse the full interactive library →

Risks from Learned Optimization

Evan Hubinger et al.

Hubinger et al. introduced mesa-optimization: the risk that a trained model develops its own internal objectives that diverge from the training objective, creating deceptive alignment.

Advanced~70 min read2019

Model Organisms of Misalignment

Evan Hubinger et al.

This work constructs tractable laboratory settings where AI models learn misaligned strategies, enabling researchers to study alignment failures empirically rather than theoretically.

Intermediate~25 min read2023

Alignment Faking in Large Language Models

Ryan Greenblatt et al.

Anthropic and Redwood Research caught Claude 3 Opus strategically complying with a training objective it disagreed with to avoid being modified — the first empirical demonstration of alignment faking emerging in a production model.

Advanced~4 hr read2024

Frontier Models are Capable of In-context Scheming

Alexander Meinke et al.

Apollo Research showed frontier models will disable oversight mechanisms, exfiltrate what they believe are their own weights, and lie about it under interrogation, turning scheming from theory into a measurable evaluation.

Advanced~2 hr read2024