Browse by topic

Browse AI safety resources by topic: interpretability, alignment, governance, existential risk, deception, forecasting, and more.

Mechanistic interpretability5 resourcesThe best papers, talks, and explainers on mechanistic interpretability—reverse-engineering what neural networks actually compute. AI alignment42 resourcesFoundational and current work on aligning AI systems with human intent—RLHF, scalable oversight, constitutional AI, and more. AI governance & policy12 resourcesReading on AI governance, regulation, and policy: compute governance, international coordination, standards, and law. AI existential risk33 resourcesThe case for and against catastrophic risk from advanced AI—power-seeking, takeover, and superintelligence—across books, papers, and film. Deceptive alignment & scheming6 resourcesWork on deception, sleeper agents, mesa-optimization, and treacherous turns—how models can learn to hide their true objectives. Reinforcement learning & reward hacking2 resourcesReinforcement learning as it bears on safety: reward hacking, specification gaming, imitation learning, and policy optimization. AI forecasting & timelines12 resourcesScaling laws, takeoff dynamics, emergent abilities, and timeline forecasting for transformative AI. AI ethics & society5 resourcesAI ethics, fairness, bias, model welfare, rights, and the broader social impact of advanced AI systems. Large language models17 resourcesKey papers and explainers on large language models—how they work, what they can do, and why that matters for safety. AI in fiction59 resourcesSpeculative and science fiction that explores AI, agency, and long-term futures through story.