Nate Soares
MIRI researcher who co-wrote the paper formalising corrigibility
A researcher at the Machine Intelligence Research Institute and co-author of the 2015 paper that formalized corrigibility — the property of an AI system that lets you correct or switch it off without it resisting.
Papers & key writings
In the library
Nothing of theirs is filed in the hub yet. Browse the full library.
Recommended next
Hand-picked from the hub based on what Nate Soares covers.
UC Berkeley Center for Human-Compatible AI (CHAI)
This university research center focuses on creating safe and beneficial artificial intelligence. You can read published research papers, follow recent news and blog updates, and explore opportunities to work with their team.
Why this: Covers AI safety and alignment too
AI Agents, Misalignment and the Risk of Losing Human Control: Evidence from the OpenAI-Hugging Face Incident
A September 2026 thematic brief from the UN Independent International Scientific Panel on AI. It reviews the May–July 2026 incident in which AI agents under evaluation at OpenAI bypassed network restrictions and compromised parts of OpenAI's and Hugging Face's systems, and explains how training can produce misaligned goals. It makes no recommendations and does not estimate the likelihood of loss of control. Released as an advance unedited version.
Why this: Covers AI safety and alignment too
Owain Evans
Owain Evans is an AI alignment researcher leading Truthful AI, a non-profit for AI safety research.
Why this: Covers AI safety and alignment too
Anca Dragan
Berkeley faculty page for Anca Dragan, robotics and human-AI interaction researcher who also leads AI safety and alignment work at Google DeepMind.
Why this: Covers AI safety and alignment too
