Proximal Policy Optimization (PPO)
PPO stabilized policy gradient training and became the optimization backbone behind RLHF pipelines including early ChatGPT, making it foundational infrastructure for alignment work.
Reinforcement learning as it bears on safety: reward hacking, specification gaming, imitation learning, and policy optimization.
Browse the full interactive library →
PPO stabilized policy gradient training and became the optimization backbone behind RLHF pipelines including early ChatGPT, making it foundational infrastructure for alignment work.
Christiano et al. established preference-based reward modeling, the foundational method that RLHF alignment pipelines later built on to steer language model behavior.