Skip to content

training

Reinforcement Learning from Human Feedback

A training paradigm that uses human preferences to train AI systems to produce outputs humans find more helpful, accurate, and safe. RLHF involves collecting comparison data from human raters, training a reward model, then optimizing the AI policy using RL.

In practice

InstructGPT was one of the first models to demonstrate that RLHF dramatically improves helpfulness and safety.