How It Works
RLCD (Reinforcement Learning for Calibrated Decisions)
TypeSafe describes it as training for “calibrated decisions: answers with epistemically honest probabilities” — where RLHF optimizes for responses human raters prefer and RLVR for outputs a program can verify, RLCD optimizes for typed decisions whose stated confidence matches how often they turn out right. The resulting model returns structured values software can use directly instead of generated text.
Origin
Named by TypeSafe AI founder Diogo Almeida in the company’s 15 September 2026 launch post for Jev, its first “System One” model, which returns typed probabilistic decisions instead of generated text. The term comes from the company itself, and its performance claims are not yet independently verified.
