Skip to content

training

PPO

Proximal Policy Optimization, a reinforcement learning algorithm widely used in RLHF training of language models. PPO updates the policy in small, controlled steps to improve stability, preventing the model from changing too dramatically in any single update.

In practice

PPO was used to train ChatGPT and InstructGPT, balancing response quality improvement with maintaining coherent behavior.