The Age of Wonders and Terrors
Computer scientist Scott Aaronson takes stock of where AI actually stands in 2026 — what has arrived, what he got wrong, and how to think clearly about the hype and the fear at the same time.
October 2026 post by computer scientist Scott Aaronson describing CS395T AI Alignment Theory, a new graduate course he teaches at UT Austin, with the course description and topics.
Read / subscribe freeComputer scientist Scott Aaronson takes stock of where AI actually stands in 2026 — what has arrived, what he got wrong, and how to think clearly about the hype and the fear at the same time.
Other resources that share this publication's topics.
LessWrong post that separates AI alignment research into two categories, engineering work to make systems behave and scientific work to understand misalignment, in the context of debate over whether some alignment research is net negative.
OpenAI's free educational resource for learning deep reinforcement learning: an introduction to the field, key algorithms with implementations, exercises, and guidance for becoming an RL practitioner or researcher.
Essay proposing three distinct facets of AI alignment: control (the system cannot cause unacceptably bad outcomes even if trying), intent alignment (the system tries to do what the user wants), and values alignment (the system refuses widely unacceptable actions). Uses a stock-trading agent scenario to show how the facets can conflict.
We0 article (July 2026) on philosopher Harvey Lederman's research on the Confucian thinker Wang Yangming and its connection to Anthropic's alignment work on Claude. Published on the blog of an AI website-builder company.
A free structured course in AI alignment and AI governance — readings, exercises and facilitated cohorts. Self-paced version free to anyone.
From the site: Free online courses, grants, and intensive in-person programs from the leading talent accelerator for beneficial AI and societal resilience. Join 10,000+ alumni and start today.
Why I recommend it: The usual route in for people trying to move into safety work. The reading list alone is worth the visit even if you never join a cohort.
OpenAI's free framework for how misaligned model behaviour should be reported and categorised — what counts as misalignment, who reports it, and what happens next.
From the site: OpenAI shares a framework for tracking, investigating, and disclosing model misalignment, alongside six reports of unexpected or concerning model behavior.
Why I recommend it: Primary source on how a major lab defines and handles its own model failures — useful, but it is the lab grading itself.