Paul Christiano
Alignment researcher behind reinforcement learning from human feedback
An alignment researcher who led the work introducing reinforcement learning from human feedback, the training method behind today’s chat assistants, and who shaped how the field talks about alignment and the alignment tax.
In the library
Nothing of theirs is filed in the hub yet. Browse the full library.
Recommended next
Hand-picked from the hub based on what Paul Christiano covers.
What Is RLCD? Reinforcement Learning from Contrast Distillation
A plain-language explainer on RLCD, a way of aligning language models by learning from contrasting outputs rather than human ratings alone.
Why this: Covers RLHF and AI alignment too
Wang Yangming, Claude, and AI Alignment: How Philosophy Entered Anthropic's Safety Work
We0 article (July 2026) on philosopher Harvey Lederman's research on the Confucian thinker Wang Yangming and its connection to Anthropic's alignment work on Claude. Published on the blog of an AI website-builder company.
Why this: Also about AI alignment
Reinforcement Learning ebook (Weights & Biases)
A free ebook walking through reinforcement learning from the basics to RLHF, written for practitioners rather than researchers.
Why this: Also about RLHF
