These Are the Most Urgent AI Risks, According to 272 Experts
MIT Sloan summary of research surveying 272 experts on which AI risks could cause the most harm in the next five years.
Resource Hub
2467 hand-picked resources, updated every week. Search it, filter it, or just browse a collection and see what catches your eye. Want today’s headlines instead? Read the free AI news feed.
Filtered by tag
Answers come only from resources in this hub, with the sources listed underneath.
20 resources
MIT Sloan summary of research surveying 272 experts on which AI risks could cause the most harm in the next five years.
Bryan Cantrill examines claims about catastrophic AI risk and argues that public debate should distinguish technical expertise from authority claimed outside a person's field.
Techdirt opinion piece (April 2024) arguing the effective altruism movement shifted its money and attention from global poverty to AI existential risk.
Why I recommend it: Openly argumentative and from 2024. Effective altruists dispute this framing; see Holden Karnofsky's and Luke Muehlhauser's profiles for their side.
An essay by SE Gyges examining whether the online rationalist community around AI risk behaves like a religious movement.
Why I recommend it: A critic's personal essay. Read alongside the rationalists' own writing on LessWrong so you hear both sides.
Microsoft's open-source toolkit for red-teaming AI systems: automated attack prompts, scoring of the responses, and repeatable runs. Free.
From the site: The Python Risk Identification Tool for generative AI (PyRIT) is an open source framework built to empower security professionals and engineers to proactively identify risks in generative AI system...
Why I recommend it: Built by the team that red-teams Microsoft's own AI products, and released as-is. Best paired with a written idea of what you are testing for.
The 2016 paper that framed AI safety as a set of specific engineering problems — side effects, reward hacking, unsafe exploration — rather than a philosophical worry. Free on arXiv.
From the site: Rapid progress in machine learning and artificial intelligence (AI) has brought increasing attention to the potential impacts of AI technologies on society. In this paper we discuss one such potential impact: the problem of accidents in machine learning systems, defined as unintended and harmful behavior that may emer…
Why I recommend it: Start here if the safety conversation sounds abstract. It is plain about what can go wrong and why, and almost everything since cites it.
A free structured course in AI alignment and AI governance — readings, exercises and facilitated cohorts. Self-paced version free to anyone.
From the site: Free online courses, grants, and intensive in-person programs from the leading talent accelerator for beneficial AI and societal resilience. Join 10,000+ alumni and start today.
Why I recommend it: The usual route in for people trying to move into safety work. The reading list alone is worth the visit even if you never join a cohort.
A short consensus paper from Geoffrey Hinton, Yoshua Bengio and two dozen other researchers on the risks they consider serious and the governance they think is needed. Free on arXiv.
From the site: Artificial Intelligence (AI) is progressing rapidly, and companies are shifting their focus to developing generalist AI systems that can autonomously act and pursue goals. Increases in capabilities and autonomy may soon massively amplify AI's impact, with risks that include large-scale social harms, malicious uses, an…
Why I recommend it: The clearest statement of what the safety-concerned researchers actually agree on, signed rather than paraphrased.
A structured survey of the risks — malicious use, competitive pressure, organizational failure, and systems pursuing goals of their own — with the evidence for each. Free to read.
From the site: There are many potential risks from AI. CAIS focusses on mitigating risks that could lead to catastrophic outcomes for society, such as bioterrorism or loss of control over military AI systems.
Why I recommend it: The best single map of the different worries, which are usually mashed together into one. Written by a safety organization, so read it as advocacy with citations.
An open-source framework from the UK's AI Security Institute for evaluating models — writing tests, scoring answers and logging what happened. Free.
From the site: Open-source framework for large language model evaluations
Why I recommend it: What a government safety institute actually uses to test models. Technical, but the docs explain the thinking behind each kind of test.
An open-source scanner that probes a language model for weaknesses — prompt injection, data leakage, jailbreaks, toxic output — and reports what it found. Free.
From the site: the LLM vulnerability scanner. Contribute to NVIDIA/garak development by creating an account on GitHub.
Why I recommend it: Point it at a model you are about to rely on and see how it fails before your users do.
A detailed scenario for how AI might develop through 2027, written by former OpenAI researcher Daniel Kokotajlo and colleagues, with the reasoning and uncertainties spelled out. Free to read in full.
From the site: A research-backed AI scenario forecast.
Why I recommend it: The forecast everyone in this field argued about. Read it as one carefully argued scenario, not a prediction — the authors say as much themselves.
An open-source tool for testing and red-teaming prompts and AI apps — run the same prompts across models, compare answers, and catch regressions. Free and self-hosted.
From the site: The AI Security Platform that catches vulnerabilities in development. Trusted by 156 of the Fortune 500 and 300,000+ developers worldwide.
Why I recommend it: The practical one: if you have built anything on top of a model, this is how you check a prompt change did not quietly make it worse.
Anthropic's paper describing how Claude is trained against a written set of principles instead of relying only on human ratings. Free on arXiv.
From the site: As AI systems become more capable, we would like to enlist their help to supervise other AIs. We experiment with methods for training a harmless AI assistant through self-improvement, without any human labels identifying harmful outputs. The only human oversight is provided through a list of rules or principles, and s…
Why I recommend it: Worth reading to see what "aligned" means in practice at one lab — and note it comes from the company selling the model.
A research nonprofit that independently evaluates frontier AI models to measure what they can actually do and what risks that creates. Reports are free.
From the site: METR is a research nonprofit that evaluates frontier AI models to inform the public about their risks and capabilities.
Why I recommend it: One of the few independent evaluators. Read their reports before you trust a lab's own capability claims.
A free arXiv preprint on recursive self-improvement in AI agents trained inside evolving simulated worlds.
Why I recommend it: Technical, and central to the safety debate about systems that improve themselves.
Stanford economist Charles I. Jones works out, in plain economic terms, how much money it would be worth spending to lower catastrophic risks from advanced AI — comparing it to the roughly 4 percent of GDP the U.S. effectively spent during Covid-19.
Computer scientist Scott Aaronson takes stock of where AI actually stands in 2026 — what has arrived, what he got wrong, and how to think clearly about the hype and the fear at the same time.
Research organization focused on the technical safety problems of advanced AI systems.
Why I recommend it: One perspective among several. Read it alongside the critics, not instead of them.
Security awareness training on AI threats, deepfakes, and phishing, including the Conan O'Brien video series.
Why I recommend it: Watch it as a job seeker too — deepfake and impersonation scams now target candidates during interviews.