John Wentworth
Independent AI alignment researcher working on natural abstractions
John Wentworth is an independent AI alignment researcher who focuses on formalizing abstraction and understanding agency. He is known for the natural abstraction hypothesis, the idea that a wide variety of minds will learn roughly the same high-level concepts humans use, and writes as johnswentworth on LessWrong and the Alignment Forum.
In the library
Nothing of theirs is filed in the hub yet. Browse the full library.
Recommended next
Hand-picked from the hub based on what John Wentworth covers.
A Three-Facet Framework for AI Alignment
Essay proposing three distinct facets of AI alignment: control (the system cannot cause unacceptably bad outcomes even if trying), intent alignment (the system tries to do what the user wants), and values alignment (the system refuses widely unacceptable actions). Uses a stock-trading agent scenario to show how the facets can conflict.
Why this: Also about ai alignment
Wang Yangming, Claude, and AI Alignment: How Philosophy Entered Anthropic's Safety Work
We0 article (July 2026) on philosopher Harvey Lederman's research on the Confucian thinker Wang Yangming and its connection to Anthropic's alignment work on Claude. Published on the blog of an AI website-builder company.
Why this: Also about ai alignment
