training
RLHF
Reinforcement Learning from Human Feedback, a training technique where human evaluators rank model outputs to train a reward model, which then guides the language model to generate more helpful, harmless, and honest responses through reinforcement learning.
In practice
ChatGPT was trained with RLHF where human raters compared pairs of responses and indicated which was better.
In the index
Tools that mention RLHF
- Scale AITrust 72
Matched on each tool’s own description and feature list, highest trust score first.