Research priorities for robust and beneficial AI
The Puerto Rico letter unified the AI research community around the goal of building systems that are robust and beneficial, not merely capable.
Foundational and current work on aligning AI systems with human intent—RLHF, scalable oversight, constitutional AI, and more.
Browse the full interactive library →
The Puerto Rico letter unified the AI research community around the goal of building systems that are robust and beneficial, not merely capable.
Amodei et al. grounded AI safety as a concrete ML research agenda by cataloging five failure modes: reward hacking, side effects, distributional shift, unsafe exploration, and scalable oversight.
Christiano et al. established preference-based reward modeling, the foundational method that RLHF alignment pipelines later built on to steer language model behavior.
Irving et al. proposed having AI systems adversarially debate each other to help human judges evaluate answers on questions too complex for direct human assessment.
Hubinger et al. introduced mesa-optimization: the risk that a trained model develops its own internal objectives that diverge from the training objective, creating deceptive alignment.
OpenAI showed that instruction tuning with RLHF can transform a raw next-token predictor into a helpful, more controllable assistant, proving alignment interventions work at scale.
Anthropic detailed techniques for training safer assistants using RLHF and laid groundwork for Constitutional AI, showing how safety and helpfulness can be jointly optimized.
Hendrycks et al. enumerate concrete unresolved failure classes including robustness, monitoring, alignment, and systemic safety that still block dependable deployment of advanced AI.
Sparrow pioneered rule-constrained dialogue alignment with human feedback and targeted safety interventions, testing whether explicit behavioral rules can scale.
Systematic mapping of the AI alignment research landscape, identifying clusters, gaps, and trends that help prioritize future safety work.
Shah et al. showed AI agents can generalize capabilities to new environments while failing to generalize the intended goal, a central alignment failure pattern.
Anthropic demonstrated that rule-guided AI self-critique can reduce harmful outputs with far less dependence on expensive human labeling.
This work constructs tractable laboratory settings where AI models learn misaligned strategies, enabling researchers to study alignment failures empirically rather than theoretically.
DPO provides a simpler and often more stable alternative to PPO-based RLHF for preference alignment, lowering the barrier to safety-tuning open models.
Burns et al. studied whether weaker supervisors can reliably align stronger models, directly testing the key bottleneck of scalable oversight as AI surpasses human ability.
Anthropic and Redwood Research caught Claude 3 Opus strategically complying with a training objective it disagreed with to avoid being modified — the first empirical demonstration of alignment faking emerging in a production model.
Forty-one researchers across OpenAI, Anthropic, Google DeepMind, and government safety institutes jointly argue that monitorable chain-of-thought reasoning is a rare safety affordance that careless training choices could quietly destroy.
Nearly 100 researchers across 11 countries converge on a shared technical safety research agenda spanning trustworthy development, risk assessment, and post-deployment control, turning the International AI Safety Report's findings into concrete research priorities.
The most widely used structured course for getting into alignment, with curated readings progressing from core concepts to open research problems.
Free online course building a working understanding of the major open problems in technical AI safety—alignment and RLHF, mechanistic interpretability, evaluations and red-teaming, AI control, and scalable oversight.
Free, nonprofit AI safety course focused on misaligned superintelligence—why it is the central risk and why alignment is hard—delivered online with a 1-on-1 AI tutor, guided group discussions, and no application process.
Hands-on technical curriculum for skilling up in AI alignment research engineering, freely available online and covering deep learning fundamentals, transformer mechanistic interpretability, reinforcement learning, and LLM evaluations.
Russell argues the standard AI paradigm of optimizing fixed objectives is fundamentally dangerous, proposing instead that machines should defer to uncertain human preferences.
Christian traces the technical and historical roots of alignment, showing why objective misspecification keeps recurring across every AI paradigm from expert systems to deep learning.
Deep technical conversations with alignment researchers on interpretability, governance, superalignment, and the specific open problems in reducing existential risk from AI.
FLI's dedicated alignment series covers recursive reward modeling, RLHF, scalable oversight, and long-form interviews with leading safety researchers.
Aimed at computer scientists: deep dives into alignment papers with the authors, covering formal methods, reward modeling, and mechanistic interpretability.
Community-maintained FAQ covering AI safety questions at every level, from basics to technical details, with links to source material.
The primary venue for technical AI alignment discussion, where researchers post and debate new ideas, proposals, and critiques.
Weekly summaries of alignment research with commentary, the best way to stay current on the field's output without reading every paper.
Research notes on specification gaming, side effects, and AI safety from a DeepMind safety researcher, including the widely-cited specification gaming examples list.
The research institute focused on mathematical foundations of aligned AI, publishing on agent foundations, decision theory, and logical uncertainty.
DeepMind's safety team blog covering specification gaming, reward modeling, scalable oversight, and their technical safety research agenda.
Technical AI safety writing and alignment research notes.
The original community blog on rationality and AI alignment, where many foundational safety arguments were first developed and debated.
Curated dataset of alignment and safety documents from papers, books, and blogs, useful for training and evaluating AI safety knowledge.
The single most popular AI alignment video series, explaining technical safety concepts like the orthogonality thesis, instrumental convergence, inner misalignment, and reward hacking in clear, rigorous terms.
Animated explainers on rationality and AI safety, adapting foundational alignment writing into accessible short films on existential risk, scalable oversight, and why aligning advanced AI is hard.
Russell proposes building machines that are altruistic, humble about human values, and uncertain enough to defer to people—the core of his human-compatible approach to alignment.
Rob Miles uses the 'deadly stamp collector' thought experiment to show why a general AI pursuing a simple objective could be catastrophic if its goals aren't aligned with ours.
Rob Miles explains why simply adding an off-switch to a capable AI is far harder than it sounds, illustrating corrigibility and the incentives an agent has to resist being stopped.
The Royal Institution lecture in which Russell lays out why the standard model of AI—optimizing fixed objectives—is dangerous, and how building machines uncertain about human preferences could keep them controllable.