AI alignment

Foundational and current work on aligning AI systems with human intent—RLHF, scalable oversight, constitutional AI, and more.

Browse the full interactive library →

Research priorities for robust and beneficial AI

Stuart Russell, Daniel Dewey, Max Tegmark

The Puerto Rico letter unified the AI research community around the goal of building systems that are robust and beneficial, not merely capable.

Intermediate~15 min read2015

Concrete problems in AI safety

Dario Amodei et al.

Amodei et al. grounded AI safety as a concrete ML research agenda by cataloging five failure modes: reward hacking, side effects, distributional shift, unsafe exploration, and scalable oversight.

Advanced~45 min read2016

Deep Reinforcement Learning from Human Preferences

Paul Christiano et al.

Christiano et al. established preference-based reward modeling, the foundational method that RLHF alignment pipelines later built on to steer language model behavior.

Advanced~30 min read2017

AI Safety via Debate

Geoffrey Irving et al.

Irving et al. proposed having AI systems adversarially debate each other to help human judges evaluate answers on questions too complex for direct human assessment.

Advanced~45 min read2018

Risks from Learned Optimization

Evan Hubinger et al.

Hubinger et al. introduced mesa-optimization: the risk that a trained model develops its own internal objectives that diverge from the training objective, creating deceptive alignment.

Advanced~70 min read2019

Instruct-GPT-3

OpenAI

OpenAI showed that instruction tuning with RLHF can transform a raw next-token predictor into a helpful, more controllable assistant, proving alignment interventions work at scale.

Advanced~2 hr read2022

Training a Helpful and Harmless Assistant with RLHF

Anthropic

Anthropic detailed techniques for training safer assistants using RLHF and laid groundwork for Constitutional AI, showing how safety and helpfulness can be jointly optimized.

Advanced~2 hr read2022

Unsolved Problems in ML Safety

Dan Hendrycks et al.

Hendrycks et al. enumerate concrete unresolved failure classes including robustness, monitoring, alignment, and systemic safety that still block dependable deployment of advanced AI.

Intermediate~50 min read2021

Improving Alignment of Dialogue Agents (Sparrow)

DeepMind

Sparrow pioneered rule-constrained dialogue alignment with human feedback and targeted safety interventions, testing whether explicit behavioral rules can scale.

Advanced~2.5 hr read2022

Goal Misgeneralization

Rohin Shah et al.

Shah et al. showed AI agents can generalize capabilities to new environments while failing to generalize the intended goal, a central alignment failure pattern.

Advanced~45 min read2022

Model Organisms of Misalignment

Evan Hubinger et al.

This work constructs tractable laboratory settings where AI models learn misaligned strategies, enabling researchers to study alignment failures empirically rather than theoretically.

Intermediate~25 min read2023

Direct Preference Optimization (DPO)

Rafailov et al.

DPO provides a simpler and often more stable alternative to PPO-based RLHF for preference alignment, lowering the barrier to safety-tuning open models.

Advanced~50 min read2023

Weak-to-Strong Generalization

Collin Burns et al.

Burns et al. studied whether weaker supervisors can reliably align stronger models, directly testing the key bottleneck of scalable oversight as AI surpasses human ability.

Advanced~90 min read2023

Alignment Faking in Large Language Models

Ryan Greenblatt et al.

Anthropic and Redwood Research caught Claude 3 Opus strategically complying with a training objective it disagreed with to avoid being modified — the first empirical demonstration of alignment faking emerging in a production model.

Advanced~4 hr read2024

The Singapore Consensus on Global AI Safety Research Priorities

Yoshua Bengio et al.

Nearly 100 researchers across 11 countries converge on a shared technical safety research agenda spanning trustworthy development, risk assessment, and post-deployment control, turning the International AI Safety Report's findings into concrete research priorities.

Intermediate~2 hr read2025

AGI Safety Fundamentals

AGI Safety Fundamentals

The most widely used structured course for getting into alignment, with curated readings progressing from core concepts to open research problems.

Intermediate~60 hr course

BlueDot Impact: Technical AI Safety

BlueDot Impact

Free online course building a working understanding of the major open problems in technical AI safety—alignment and RLHF, mechanistic interpretability, evaluations and red-teaming, AI control, and scalable oversight.

Intermediate~30 hr course

Lens Academy

Lens Academy

Free, nonprofit AI safety course focused on misaligned superintelligence—why it is the central risk and why alignment is hard—delivered online with a 1-on-1 AI tutor, guided group discussions, and no application process.

Beginner~35 hr course

ARENA (Alignment Research Engineer Accelerator)

ARENA

Hands-on technical curriculum for skilling up in AI alignment research engineering, freely available online and covering deep learning fundamentals, transformer mechanistic interpretability, reinforcement learning, and LLM evaluations.

Advanced~200 hr course

Human Compatible

Stuart Russell

Russell argues the standard AI paradigm of optimizing fixed objectives is fundamentally dangerous, proposing instead that machines should defer to uncertain human preferences.

Intermediate~11 hr read2019

The Alignment Problem

Brian Christian

Christian traces the technical and historical roots of alignment, showing why objective misspecification keeps recurring across every AI paradigm from expert systems to deep learning.

Intermediate~15 hr read2020

AXRP (AI X-risk Research Podcast)

Daniel Filan

Deep technical conversations with alignment researchers on interpretability, governance, superalignment, and the specific open problems in reducing existential risk from AI.

Advanced~1.5 hr per episode2020

AI Alignment Podcast

Future of Life Institute

FLI's dedicated alignment series covers recursive reward modeling, RLHF, scalable oversight, and long-form interviews with leading safety researchers.

Intermediate~60 min per episode2018

Technical AI Safety Podcast

Quinn Dougherty

Aimed at computer scientists: deep dives into alignment papers with the authors, covering formal methods, reward modeling, and mechanistic interpretability.

Advanced~60 min per episode2020

AI Safety Info (Stampy's FAQ)

StampyAI

Community-maintained FAQ covering AI safety questions at every level, from basics to technical details, with links to source material.

Beginner

Alignment Forum

Center for Applied Rationality

The primary venue for technical AI alignment discussion, where researchers post and debate new ideas, proposals, and critiques.

Advanced

Alignment Newsletter

Rohin Shah

Weekly summaries of alignment research with commentary, the best way to stay current on the field's output without reading every paper.

Intermediate

Victoria Krakovna's blog

Victoria Krakovna

Research notes on specification gaming, side effects, and AI safety from a DeepMind safety researcher, including the widely-cited specification gaming examples list.

Advanced

DeepMind AI Safety Research

DeepMind

DeepMind's safety team blog covering specification gaming, reward modeling, scalable oversight, and their technical safety research agenda.

Advanced

carado.moe

carado

Technical AI safety writing and alignment research notes.

Advanced

LessWrong

LessWrong

The original community blog on rationality and AI alignment, where many foundational safety arguments were first developed and debated.

Intermediate

StampyAI Alignment Research Dataset

StampyAI

Curated dataset of alignment and safety documents from papers, books, and blogs, useful for training and evaluating AI safety knowledge.

Advanced

Robert Miles AI Safety

Robert Miles

The single most popular AI alignment video series, explaining technical safety concepts like the orthogonality thesis, instrumental convergence, inner misalignment, and reward hacking in clear, rigorous terms.

Intermediate~15 min per video2017

Rational Animations

Rational Animations

Animated explainers on rationality and AI safety, adapting foundational alignment writing into accessible short films on existential risk, scalable oversight, and why aligning advanced AI is hard.

Beginner~10 min per video2020

Deadly Truth of General AI? – Computerphile

Robert Miles

Rob Miles uses the 'deadly stamp collector' thought experiment to show why a general AI pursuing a simple objective could be catastrophic if its goals aren't aligned with ours.

Beginner~10 min watch2015

AI "Stop Button" Problem – Computerphile

Robert Miles

Rob Miles explains why simply adding an off-switch to a capable AI is far harder than it sounds, illustrating corrigibility and the incentives an agent has to resist being stopped.

Beginner~20 min watch2017

How Not to Destroy the World with AI

Stuart Russell

The Royal Institution lecture in which Russell lays out why the standard model of AI—optimizing fixed objectives—is dangerous, and how building machines uncertain about human preferences could keep them controllable.

Intermediate~60 min watch2023