How It Works
RLVR (Reinforcement Learning with Verifiable Rewards)
A post-training method where a model is fine-tuned using reinforcement learning with rewards from an automatic verifier (for example, checking whether a math answer or code output is correct), instead of rewards from human preference ratings.
Origin · no single documented coiner
The term crystallized as a category label around early 2025, after DeepSeek-R1 showed strong reasoning from reinforcement learning with automatically checkable rewards; it was soon formalized and analyzed in academic work such as Wen et al. at Microsoft Research Asia.
