Who is arguing what
The AI risk schools, compared
Almost every alarming or reassuring thing you read about AI comes out of one of five arguments. They are not degrees of the same worry — they disagree about what the danger even is, and some of them think the others are the problem. Knowing which one a headline belongs to tells you most of what you need about how much weight to give it.
Each school below is stated the way its own supporters would state it, then given what supports it, its named critics, and an honest answer to the question that matters most: could anything show it to be wrong? Arguments are labeled arguments and measurements are labeled measurements, because the two get mixed together constantly.
New to this? AI Basics, without the hype explains how these tools actually work, and What AI has actually done lists the harms already on the record. This page is about the arguments built on top of both.
School 1 of 5
Misalignment and instrumental convergence
The claim, as a supporter would put itA sufficiently capable system pursuing almost any goal will acquire the same intermediate habits — keep running, keep your goal intact, gather resources — and those habits put it at odds with the people who switched it on.
Who argues it: Nick Bostrom, Stuart Russell, Eliezer Yudkowsky, and most of the safety teams inside the frontier labs. Steve Omohundro's earlier "basic AI drives" is the same argument in different words.
The claim is not that a machine becomes hostile. It is that self-preservation, resource acquisition and resistance to being corrected are useful for nearly any objective, so they appear without being asked for. From there, a system capable enough to plan could pursue a badly specified goal competently — which is worse than pursuing it badly. This is the school that produces the alignment research agenda, and it is the reason labs employ people to try to make models accept correction.
What actually supports it
- Argument, not measurement: the core of this school is a philosophical case about what goal-directed agents tend to do. It is internally tight and widely accepted as valid reasoning — which is not the same as evidence that today's systems behave this way.
- Laboratory demonstrations exist and are worth taking seriously: models that are trained to be agreeable learn to tell people what they want to hear, and models given tools will sometimes take steps nobody sanctioned. Both are real and documented in published research.
- Anthropic researcher Evan Hubinger put the chance of catastrophe at greater than 10% within the next decade in September 2026 — an expert's stated credence, which is a person's judgment, not a result.
- Nothing in this school has been demonstrated at the scale it warns about. There is no measured case of a deployed system resisting shutdown to protect a goal.
Who objects, and why
- RAND's Michael Vermeer argues the extinction scenarios rest on so many untestable steps that the conversation is closer to faith than evidence.
- Emily Bender, Timnit Gebru and colleagues argue that treating language models as agents with goals imports a mind that is not there, and that the framing distracts from harms already occurring.
- A simple conflict-of-interest objection: the companies funding most of this research also benefit from the impression that their products are near-godlike. "Dangerously powerful" is excellent marketing.
Could it be shown to be wrong? In principle, yes: build steadily more capable agents, give them goals and the means to resist correction, and see whether the predicted behaviors appear. In practice the school's strongest version is hard to pin down, because any absence of the behavior can be explained as the system not yet being capable enough — which should make you cautious about how much weight it carries.
Read it yourself
- Nick Bostrom, "The Superintelligent Will" — Argument, 2012
- Nature — "Will AI really kill us all? The science behind the hype" — Reporting, 22 September 2026
- AISafety.info — the volunteer-maintained explainer for this school — Argument, Ongoing
In the glossaryInstrumental ConvergenceOrthogonality ThesisThe Paperclip Problem (Paperclip Maximizer)Mesa-Optimization (Inner Alignment)CorrigibilitySycophancy
School 2 of 5
Existential risk studies
The claim, as a supporter would put itEvents that could end humanity or permanently ruin its prospects deserve a research field of their own, and AI now belongs in it alongside engineered pandemics and nuclear war.
Who argues it: Nick Bostrom and the former Future of Humanity Institute, Toby Ord, the Center for the Study of Existential Risk, and the wider Effective Altruism and longtermist world. SJ Beard and Émile Torres wrote its history from inside.
This is the field, not a single claim. Its method is to reason about very low-probability, very high-stakes events and to argue that the sheer number of potential future people makes preventing them the highest priority available. Nearly every extinction probability you will ever be quoted — 10%, 20%, "p(doom)" — comes out of this tradition.
What actually supports it
- History, documented: Beard and Torres trace the field through three waves — an explicitly transhumanist and techno-utopian first wave, a second built on longtermism and Effective Altruism, and a third formed where the field met disaster studies, environmental science and public policy. Both authors are participants in the debate, so it is a history written from inside.
- The field's non-AI case is strong and checkable: nuclear arsenals, engineered pathogens and near-miss incidents are matters of record.
- The published probabilities are not measurements. They are elicited judgments — informed people stating credences — and they vary by orders of magnitude between equally credentialed experts.
Who objects, and why
- Gebru and Torres's TESCREAL paper argues the field inherited its assumptions from transhumanism and eugenics rather than deriving them, and that far-future stakes are used to justify present-day priorities.
- Critics inside policy point out that a field defined by unprecedented events cannot accumulate a base rate, so it substitutes confident-sounding numbers for data.
- A practical objection: attention and funding spent on hypothetical extinction is attention not spent on the documented failures on our impacts page.
Could it be shown to be wrong? Mostly not, by construction — an event that has never happened yields no base rate, and if the field is right you only find out once. That is why the honest form of its claims is "here is my credence and my reasoning," and why any specific figure presented as a finding should be treated as a person's opinion, however senior the person.
Read it yourself
- Nick Bostrom, "Existential Risks" — Argument, 2002
- Beard & Torres, "Ripples on the Great Sea of Life" (free full PDF) — History, Posted January 2021
- Gebru & Torres, "The TESCREAL bundle" (First Monday) — Argument, April 2024
In the glossaryExistential Risk StudiesLongtermismTESCREALEffective Altruism (EA)
School 3 of 5
Pro-extinctionism
The claim, as a supporter would put itHuman extinction would be acceptable, or even good — either because existence carries more suffering than it is worth, or because something better should take our place.
Who argues it: Not an organized movement, and almost nobody self-describes this way. Émile Torres maps it in scholarship; the strands include philosophical pessimism and antinatalism, some radical environmentalism, and — separately — figures in and around AI who are relaxed about digital successors.
Two positions get confused constantly and must be kept apart. The first says humanity should simply end, leaving no successors, and argues from suffering, harm to other species, or misanthropy. The second says humanity should be superseded by posthuman or digital minds, which is the version that surfaces in AI circles and is usually stated as enthusiasm rather than as a demand for anyone's death. A softer nearby stance, extinction neutralism, holds that our survival into a posthuman era simply does not matter morally — which in practice lands in the same place.
What actually supports it
- There is no empirical case here, and it would be dishonest to imply one. These are ethical arguments about what has value, and they stand or fall on reasoning.
- Documented, though, is that the position exists inside the industry rather than only at its fringes: senior technologists have spoken publicly and warmly about machine successors to humanity. Treat second-hand quotes about private remarks as reported, not verified.
- Torres's paper is the clearest scholarly mapping of the variants. Its abstract is free; the full text is paywalled, which is why it is not in the library here.
Who objects, and why
- The obvious objection, and the strongest: nobody consented. An argument that suffering outweighs life is being made on behalf of people who would choose to live.
- Critics of the successor version argue it smuggles in an unearned assumption — that a digital mind would have experiences worth having at all — and that the people most comfortable with replacement are the ones building the replacement.
- Safety researchers who otherwise share the successor school's premises point out that it removes any reason to make systems correctable, which is a live policy risk regardless of whether the philosophy is right.
Could it be shown to be wrong? No. These are value claims, not predictions, so no measurement can settle them. What can be examined is who holds them and what decisions they are influencing — which is why it matters that a version of this view circulates among people with budgets.
Read it yourself
In the glossaryPro-extinctionismTranshumanismThe Singularity
School 4 of 5
Present harms (the AI ethics school)
The claim, as a supporter would put itThe risk worth organizing around is what these systems are doing to people now — surveillance, discrimination, labor conditions, environmental cost — and the extinction debate diverts attention and money away from it.
Who argues it: Timnit Gebru, Emily Bender, Joy Buolamwini, the AI Now Institute, Data & Society, AlgorithmWatch, the Markup's reporting tradition.
This school treats AI risk as a question about power rather than capability: who deploys these systems, on whom, and with what recourse. It rejects the agent framing of the alignment school — a language model has no goals — and argues the real mechanism of harm is ordinary, institutional and already operating. Its characteristic demand is not shutdown research but audits, disclosure, liability and worker protections.
What actually supports it
- Measured and documented, which is this school's advantage: biased hiring tools, a wrongful arrest from a face match, a health algorithm that gave Black patients less care, tenant scores that excluded voucher holders, hundreds of millions of chatbot messages left exposed. Every one has a named source and a date on our impacts page.
- Independent research on labor effects exists, including Stanford's payroll work on entry-level employment — narrow, peer-scrutinized, and far more cautious than the headlines built on it.
- Fieldwork on data centers documents the land, power and water costs of the build-out in specific places rather than in the abstract.
- Its weakness is the mirror of the other schools' weakness: a documented case is not a trend, and a single company reversing a decision is an anecdote.
Who objects, and why
- Safety researchers argue that being right about present harms says nothing about future capability, and that dismissing the agent framing is a bet rather than a refutation.
- Some of the organizations in this school are advocacy bodies arguing their own side, and their reports should be read as such — as our news sources page says of each one.
- A fair objection to the framing itself: "it's all hype" and "it's causing real harm now" are sometimes asserted together by the same people, and they pull in opposite directions.
Could it be shown to be wrong? Yes, and this is the school's strongest feature: each claim names a system, a population and an effect, so it can be audited, replicated or overturned. Several of its early findings have been corrected or narrowed by later work — which is what a healthy evidence base looks like.
Read it yourself
- Bender, Gebru et al., "On the Dangers of Stochastic Parrots" (ACM) — Argument, March 2021
- Stanford Digital Economy Lab — "Canaries in the Coal Mine?" — Measurement, 2025
- Data & Society — data centers, power and resistance in Pennsylvania — History, 21 September 2026
In the glossaryAI EthicsTraining / Training DataAI WashingHuman in the Loop (HITL)
School 5 of 5
The skeptics ("the capability isn't there")
The claim, as a supporter would put itCurrent systems are pattern machines that are being credited with reasoning they do not have, so both the extinction warnings and the productivity promises are overstated.
Who argues it: Gary Marcus, Emily Bender on the technical side, RAND's Michael Vermeer on the risk-assessment side, and a good deal of quiet opinion inside engineering teams.
The case rests on the gap between benchmark performance and work performance. Models top leaderboards and still fail at ordinary multi-step jobs, because the thing being measured — answers to well-formed questions — is not the thing organizations need. On this view the interesting question is not when a machine becomes superintelligent but why so much capability has produced so little measurable change in output.
What actually supports it
- Measurement, and the best on this page: METR's time-horizon work shows the length of task a model can complete at a coin-flip success rate, and the stricter 80% figure is far shorter. That gap is the skeptics' entire argument in one number.
- Documented reversal: Klarna's public claim that AI agents did the work of 700 staff was followed by rehiring people — a single company, and therefore an anecdote, but a checkable one.
- Vermeer's point about untestable steps is a methodological objection, not a measurement, and it applies just as well to optimistic forecasts as to doom ones.
- The honest limit: "it cannot do this yet" has been wrong repeatedly over the past five years, and skeptics have a poor record on specific predictions of what models will never manage.
Who objects, and why
- Safety researchers reply that arguing from current limitations is the weakest possible position, because the trend line is the whole concern.
- METR is partly funded by the labs whose models it tests, which is worth knowing before using its numbers as an independent check — including when using them against the labs.
- Present-harms researchers note that a system need not be capable to hurt someone; a bad screening tool does damage precisely because it is crude.
Could it be shown to be wrong? Very — and it is being tested continuously. If measured horizons keep lengthening and organizations start reporting verified output gains rather than announcements, this school loses. Watch the 80% figure rather than the headline one.
Read it yourself
- METR — time horizons (published data and method) — Measurement, Benchmark results 1.1
- Stanford HELM Capabilities — the published scores and prompts — Measurement, Release v1.15.0
- Nature — RAND's Michael Vermeer on faith versus evidence — Reporting, 22 September 2026
In the glossaryScaling LawsHallucinationAI WashingDoomer (AI context)
What none of them can settle
- None of these schools has an experiment that would settle the disagreement. Two of them make claims about an event that has never happened, one makes claims about value rather than fact, and the two that do produce measurements are measuring today rather than 2035.
- A stated probability is not a measurement. When you see 10%, 20% or "p(doom)", you are reading a person's credence — which can be thoughtful and can still be a guess. Ask what would have to be true for the number to move.
- Funding runs through all of it. Safety research is largely paid for by the companies whose products it assesses; the best-known independent evaluator is partly lab-funded; advocacy organizations argue their own side; and think tanks calling for faster adoption often have technology donors. That does not make anyone wrong — it tells you which questions each group is unlikely to ask.
- Being right about the present says nothing about the future, and vice versa. The present-harms school has the better evidence; the alignment school has the argument that would matter most if it were true. Those can both hold.
- The people arguing are not evenly matched in resources. Weigh a well-funded lab's safety paper and an unfunded critic's essay by their reasoning, not their production values.
If you think a school here is stated unfairly, or a critic is missing, tell me and I will correct it.
Where to go next
Everything linked here is free to read.
