training
Reinforcement Learning from Human Feedback
A training paradigm that uses human preferences to train AI systems to produce outputs humans find more helpful, accurate, and safe. RLHF involves collecting comparison data from human raters, training a reward model, then optimizing the AI policy using RL.
In practice
InstructGPT was one of the first models to demonstrate that RLHF dramatically improves helpfulness and safety.