Back to Anthropic (Claude)All frontier AI labs
Ethics and safety policies
Anthropic (Claude): what they say, and what they leave out
Every claim on this page comes from documents Anthropic (Claude) published itself, linked below so you can check me. Checked September 19, 2026.
A policy is not an audit
Almost every document linked on this page was written by the company it governs, and graded by that same company. A long policy list does not make a lab more ethical than a short one. This page keeps three things apart: what the lab has committed to in writing, what its documents do not cover or where they have been credibly criticized, and — clearly labeled at the end — my own read.
Their stated position
Anthropic was founded explicitly around AI safety research and publishes a Responsible Scaling Policy that ties capability levels to required safeguards, with the stated commitment to pause deployment if safeguards aren't ready.
It publishes a written constitution — the principles its models are trained against — and a good deal of its interpretability research.
The documents themselves
Read these rather than this summary if a decision depends on it.
- Responsible Scaling Policy
AI Safety Level tiers and the safeguards each one requires before deployment.
- Claude's Constitution
The written principles used to train model behavior.
- Usage policy
Prohibited uses, including restrictions on high-risk decisions about people.
- Transparency hub
Its published reporting on enforcement, safeguards and model evaluations.
What they have committed to
Stated in their own published policies, in writing.
- — Publishes system cards with pre-deployment safety evaluations, including red-team findings.
- — Restricts its highest-capability models to vetted organizations rather than general release.
- — Publishes interpretability research that makes its own models easier to criticize.
What those documents don't cover
Plain absences, and criticism that has been documented publicly — not speculation.
- — The safety levels are self-assessed. Anthropic decides which tier a model is in and whether the safeguards suffice.
- — It has revised the Responsible Scaling Policy as models advanced, which critics read as moving the goalposts it set itself.
- — Training-data sources are undisclosed, and it settled a major authors' copyright case rather than litigating the question.
- — The company argues frontier development is dangerous and is also racing to do it — a tension it acknowledges but has not resolved.
My read
This section is my opinion, clearly separated from everything above. Disagree with it freely.
Of the big labs, Anthropic publishes the most material that could be used against it, which I count in its favor. That is a disclosure practice, not an external check, and the pause commitment has never been tested in public.
