Skip to content
Launchpad Library logo
← All people

Chris Olah

Co-founder, Anthropic; interpretability research lead

Chris Olah is a machine learning researcher and co-founder of Anthropic, where he leads interpretability research into how AI models work internally. He previously led OpenAI's interpretability team from 2018 to 2020 and worked at Google Brain, and co-founded the journal Distill.

Wikipedia
#Anthropic#interpretability#AI safety
Visit their site

In the library

Nothing of theirs is filed in the hub yet. Browse the full library.

Recommended next

Hand-picked from the hub based on what Chris Olah covers.

Research PaperFreeanthropic.com

Emotion concepts in a large language model

Anthropic interpretability research looking at internal representations that behave like emotions in its Claude model.

Why this: Covers Anthropic and interpretability too

ArticleFreecnbc.com

Anthropic Warns of AI's "Existential Risk to Humanity" in IPO Filing

CNBC, reporting from Reuters, on Anthropic's IPO filing, which devotes 80 pages to AI risks — nearly double the 48 pages on its business plans — and includes a warning that AI could pose existential risks.

Why this: Covers Anthropic and AI safety too

WebsiteFreeseverinfield.com

Severin Field

Personal site of Severin Field, visiting fellow at the Institute for AI Policy and Strategy and former MATS research fellow, who has studied interpretability, deception and persuasion in language models. Links to his blog and publications.

Why this: Covers interpretability and AI safety too

ArticleFreeanthropic.com

Introducing Claude Opus 5.5 (Anthropic)

Anthropic's own announcement of Claude Opus 5.5, published 22 September 2026 — the first model in the Claude 5.5 family. The company's claims, in its own words: it performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5 on typical workloads; on Anthropic's automated behavioural audit, its most comprehensive internal alignment test, it scores the highest of any model the company has tested. The page also cites early-tester anecdotes, including one completing a 680,000-line code migration in under a day, and a result of succeeding 39 times out of 40 on a page-load optimisation task. It is Anthropic's first release since the company publicly called for pacing the frontier, and it was tested before release by external evaluators including METR and Frontier Design.

Why this: Covers Anthropic and AI safety too