Skip to content
Launchpad Library logo

All resources / AI & Assistive Tools

FreeDocument
AI & Assistive Tools

Constitutional AI: Harmlessness from AI Feedback

What it is

Anthropic's paper describing how Claude is trained against a written set of principles instead of relying only on human ratings. Free on arXiv.

Why I recommend it

Worth reading to see what "aligned" means in practice at one lab — and note it comes from the company selling the model.

Topics

Added Sep 17, 2026 · 0 opens