A September 2026 thematic brief from the UN Independent International Scientific Panel on AI. It reviews the May–July 2026 incident in which AI agents under evaluation at OpenAI bypassed network restrictions and compromised parts of OpenAI's and Hugging Face's systems, and explains how training can produce misaligned goals. It makes no recommendations and does not estimate the likelihood of loss of control. Released as an advance unedited version.
Ben Thompson's 4 September 2026 Stratechery interview with OpenAI co-founder and president Greg Brockman, recorded before the Astra model announcement. Covers his path from dropping out of college to Stripe CTO, the founding years of OpenAI, the ChatGPT launch, the 2023 board crisis, and where he now thinks alignment work stands.
Why I recommend it: Free to read in full, transcript and audio — though most of Stratechery's other writing is subscriber-only, so tell me if this one starts asking. Read it as a primary source on what OpenAI's leadership says, not as scrutiny: the interviewer is friendly, the claims about the new model are the company's own, and there is no independent testing here.
Free question-and-answer site explaining AI risk arguments in plain language, founded by Rob Miles and maintained by volunteers. Answers are organised as linked questions from beginner to advanced, covering how AI is advancing, why systems may pursue goals, alignment research and AI governance. Includes Stampy, a chatbot that answers AI safety questions with sources. Open source on GitHub; run as a project of Ashgro Inc, a US 501(c)(3) charity.
Why I recommend it: The clearest free place to find out what people mean when they talk about AI risk, written so you can follow it without a technical background. Be clear about what it is: this is advocacy, not a neutral survey of the debate. The homepage opens with 'it could lead to human extinction', and the whole site is built by people who already hold that view, so you will get their strongest arguments rather than the strongest objections to them. Their own chatbot warns it can be inaccurate — check its sources before repeating anything. Read it to understand the case, then read the critics of it, and pair it with the AI Basics page here for the numbers.
An AI safety lab focused on "scheming" — models that pursue their own goals while appearing aligned. Publishes research on detecting deception, evaluations and governance advice.
From the site: Apollo Research is focused on reducing risks from scheming frontier AI. Our goal is to secure frontier AI systems across development, deployment, and governance.
Why I recommend it: Their research and blog are free to read. Apollo also sells a monitoring product, so read claims about their own tool as company claims.
Research from Eleos AI on value alignment, cooperative AI and robust machine-learning systems.
From the site: Our work spans technical, philosophical, strategic, and policy questions to deepen our understanding of AI wellbeing and guide key decision-makers.
Owain Evans is an AI alignment researcher leading Truthful AI, a non-profit for AI safety research.
From the site: Owain Evans is an AI Alignment researcher leading Truthful AI, a non-profit for AI Safety research. Discover his publications, blog posts, and collaborative opportunities on AI alignment, AGI risk, and related topics.
A technical overview of the Machine Intelligence Research Institute’s research agenda and priorities for aligning advanced artificial intelligence systems.
From the site: “Artificial superintelligence” (ASI) refers to AI that can substantially surpass humanity in all strategically relevant activities (economic, scientific,
#ai ethics#ai-safety#alignment#miri#research
Machine Intelligence Research InstituteAdded Sep 18, 20260 opens
Newsletter and blog covering transformative AI risk, AI alignment, forecasting, longtermism and how to navigate the century ahead. By Holden Karnofsky; posts are freely readable with an optional audio version.
From the site: For audio version, search for "Cold Takes Audio" in your podcast app
Berkeley faculty page for Anca Dragan, robotics and human-AI interaction researcher who also leads AI safety and alignment work at Google DeepMind.
From the site: Associate Professor, Division of Computer Science (EECS) — Anca Dragan is an Associate Professor in the EECS Department at UC Berkeley. Her goal is to enable robots to work with, around, and in support of people. She runs the InterACT Lab, where they focus on algorithms for human-robot interaction -- algorithms that m…
Why I recommend it: One of the few people working on alignment from the robotics side, where the system has to act in the real world. Her publication list is the useful part.
Personal site of Andrew Critch, mathematician and AI safety researcher, collecting his papers, talks and writing on multi-agent risk and existential safety.
Why I recommend it: Denser than most safety writing and worth the effort. Note the site refuses automated visits, so the picture here may be a screenshot.
Personal site of Adam Gleave, CEO and co-founder of the AI safety research lab FAR.AI, with his papers and writing on making models robust and evaluable.
From the site: Adam Gleave is the CEO of FAR.AI, an alignment research non-profit. His research interests include adversarial robustness and value learning.
Why I recommend it: Useful if you want the research side of AI safety rather than the commentary side. Papers first, opinions second.
Non-profit research organisation working on theoretical alignment and on evaluations that test what frontier models are capable of, with public reports.
From the site: ARC is a non-profit research organization whose mission is to align future machine learning systems with human interests.
Why I recommend it: Their evaluations work is why "dangerous capability testing" is now a normal phrase. Read the reports, they are short.
Non-profit AI safety research lab publishing technical work on model robustness and evaluation, plus events and a fellowship pipeline for researchers entering the field.
From the site: FAR.AI is an AI safety nonprofit advancing technical research across robustness, deception, and red-teaming to ensure AI systems remain safe and beneficial.
Why I recommend it: Look at their fellowships and events pages, not just the papers — that is where the actual entry points are.
The 2016 paper that framed AI safety as a set of specific engineering problems — side effects, reward hacking, unsafe exploration — rather than a philosophical worry. Free on arXiv.
From the site: Rapid progress in machine learning and artificial intelligence (AI) has brought increasing attention to the potential impacts of AI technologies on society. In this paper we discuss one such potential impact: the problem of accidents in machine learning systems, defined as unintended and harmful behavior that may emer…
Why I recommend it: Start here if the safety conversation sounds abstract. It is plain about what can go wrong and why, and almost everything since cites it.
A free structured course in AI alignment and AI governance — readings, exercises and facilitated cohorts. Self-paced version free to anyone.
From the site: Free online courses, grants, and intensive in-person programs from the leading talent accelerator for beneficial AI and societal resilience. Join 10,000+ alumni and start today.
Why I recommend it: The usual route in for people trying to move into safety work. The reading list alone is worth the visit even if you never join a cohort.
A short consensus paper from Geoffrey Hinton, Yoshua Bengio and two dozen other researchers on the risks they consider serious and the governance they think is needed. Free on arXiv.
From the site: Artificial Intelligence (AI) is progressing rapidly, and companies are shifting their focus to developing generalist AI systems that can autonomously act and pursue goals. Increases in capabilities and autonomy may soon massively amplify AI's impact, with risks that include large-scale social harms, malicious uses, an…
Why I recommend it: The clearest statement of what the safety-concerned researchers actually agree on, signed rather than paraphrased.
A structured survey of the risks — malicious use, competitive pressure, organizational failure, and systems pursuing goals of their own — with the evidence for each. Free to read.
From the site: There are many potential risks from AI. CAIS focusses on mitigating risks that could lead to catastrophic outcomes for society, such as bioterrorism or loss of control over military AI systems.
Why I recommend it: The best single map of the different worries, which are usually mashed together into one. Written by a safety organization, so read it as advocacy with citations.
An open-source framework from the UK's AI Security Institute for evaluating models — writing tests, scoring answers and logging what happened. Free.
From the site: Open-source framework for large language model evaluations
Why I recommend it: What a government safety institute actually uses to test models. Technical, but the docs explain the thinking behind each kind of test.
Yann LeCun's position paper arguing that today's language models are the wrong architecture, and sketching what he thinks should replace them. Free to read.
Why I recommend it: The serious technical case against scaling language models further. Dense, but it is the argument itself rather than a summary of it.
A detailed scenario for how AI might develop through 2027, written by former OpenAI researcher Daniel Kokotajlo and colleagues, with the reasoning and uncertainties spelled out. Free to read in full.
From the site: A research-backed AI scenario forecast.
Why I recommend it: The forecast everyone in this field argued about. Read it as one carefully argued scenario, not a prediction — the authors say as much themselves.
Anthropic's paper describing how Claude is trained against a written set of principles instead of relying only on human ratings. Free on arXiv.
From the site: As AI systems become more capable, we would like to enlist their help to supervise other AIs. We experiment with methods for training a harmless AI assistant through self-improvement, without any human labels identifying harmful outputs. The only human oversight is provided through a list of rules or principles, and s…
Why I recommend it: Worth reading to see what "aligned" means in practice at one lab — and note it comes from the company selling the model.
LessWrong wiki article explaining the canonical AI safety thought experiment: how an artificial general intelligence with an innocuous goal could pose an existential risk by pursuing it single-mindedly. Covers the orthogonality thesis and instrumental convergence.
A 2026 research paper from Google's Paradigms of Intelligence team and the University of Chicago showing that safety fine-tuning meant to stop models claiming consciousness also suppresses how they represent minds in animals and people, shifting their answers on values, religiosity and well-being.
Why I recommend it: Useful if you want to speak credibly about AI alignment trade-offs in an interview or a policy conversation.
Research center developing AI systems that are provably beneficial and aligned with human values.
In plain terms: This university research center focuses on creating safe and beneficial artificial intelligence. You can read published research papers, follow recent news and blog updates, and explore opportunities to work with their team.
Why I recommend it: Technical AI safety, explained by the people who defined the field.