How It Works
RLHF (Reinforcement Learning from Human Feedback)
A training technique that fine-tunes a model using human rankings of its outputs, making its behavior better match what people actually want.
Origin
Foundational method from Paul Christiano and colleagues (OpenAI/DeepMind, 2017); popularized for language models via OpenAI’s 2022 InstructGPT paper.
