
Melanie Mitchell
Santa Fe Institute professor; author of a leading plain-language AI guide
Complexity researcher who writes about what current AI systems can and cannot do, arguing that benchmark performance is routinely mistaken for understanding.
Key arguments & positions
- Argues benchmark scores are regularly mistaken for understanding, and that models fail at the abstraction and analogy humans use easily.
- Warns against both dismissing and over-crediting current systems, and pushes for evaluation designed to resist memorization.
- Skeptical of near-term AGI timelines while treating present-day harms as the more tractable problem.
Accomplishments
- Davis Professor of Complexity at the Santa Fe Institute; PhD under Douglas Hofstadter.
- Author of Artificial Intelligence: A Guide for Thinking Humans (2019) and Complexity: A Guided Tour.
- Writes the AI: A Guide for Thinking Humans newsletter and co-leads work on abstraction-and-reasoning evaluation.
Papers & key writings
Links
In the library
Nothing of theirs is filed in the hub yet. Browse the full library.
Recommended next
Hand-picked from the hub based on what Melanie Mitchell covers.
How to talk about "AI" without adding to the anthropomorphization
Emily M. Bender and Alex Hanna's Mystery AI Hype Theater 3000 newsletter on word choices that stop us describing software as if it thinks, feels or understands.
Why this: Also about AI hype
Arena Blog (LLM and Agent Evaluation Research)
Research and product writing from the team behind the Arena model leaderboards: how coding-agent harnesses change cost and success rates, how the agent leaderboards are built, and their academic partnership calls.
Why this: Also about AI evaluation
A Call for Control of Frontier AI Models
A short joint statement issued in September 2026 by heads of state and government — launched by President Alexander Stubb of Finland and Prime Minister Jonas Gahr Støre of Norway, with 22 leaders from 20 countries signed on at launch. It asks for three things: mandatory pre-deployment testing and independent evaluation by qualified evaluators with real access; coordinated common standards and shared reporting of serious safety incidents, with scientific capacity available to countries in every region; and for UN member states to explore an international institution that could set standards, verify compliance and convene states when capability thresholds are crossed.
Why this: Also about AI evaluation
