Short answer: it's worth a few free hours before it's worth any money. Here is how to test whether ai safety is worth it for you, using the 100 free entries the library already holds.
A benchmark and tracker that documents reported instances of AI agents undertaking activity characterized as illegal, ranking major AI labs by aggregated incident counts.
Why I recommend it: This is exactly the kind of uncomfortable accountability tool our field needs. I include it because we cannot have thoughtful conversations about AI deployment without looking at real-world harm.
Partnership on AI's open library of guidance, frameworks, and case studies on responsible AI: synthetic media, labor and the economy, AI safety, fairness, and inclusive AI development.
Why I recommend it: When you need a credible source instead of a hot take, cite these. The labor and economy work is the most useful set for career conversations about automation.
MIT Technology Review report on recursive self-improvement — the idea that AI systems will soon rewrite and retrain themselves — and where the current evidence actually stands.
Why I recommend it: MIT Technology Review gives you a few free articles a month before a paywall. If you hit it, borrow via a library or your workplace subscription rather than paying at the door.
Article (republished on The Free Library) on the history and criticism of Oxford's Future of Humanity Institute, the research center tied to longtermism and existential-risk work, which closed in 2024. The page blocked an automated check, so this description is based on the title.
An Internet Archive copy of Eliezer Yudkowsky's autobiographical page from his sysopmind.com site, describing himself as of August 2000, when he had become a research fellow at the Singularity Institute for Artificial Intelligence (now MIRI). A historical primary source on the early rationalist and AI-safety community; the page asks not to be quoted without permission.
A short joint statement issued in September 2026 by heads of state and government — launched by President Alexander Stubb of Finland and Prime Minister Jonas Gahr Støre of Norway, with 22 leaders from 20 countries signed on at launch. It asks for three things: mandatory pre-deployment testing and independent evaluation by qualified evaluators with real access; coordinated common standards and shared reporting of serious safety incidents, with scientific capacity available to countries in every region; and for UN member states to explore an international institution that could set standards, verify compliance and convene states when capability thresholds are crossed.
Why I recommend it: Read this as the primary source rather than someone's summary of it — it is one page, in plain language, and you can read the whole thing in three minutes. Two things worth noticing: it is a political appeal, not a law or a treaty, so nothing in it binds any company today; and the United States and China are not among the signatories, which matters given where the frontier labs are. Useful if you are writing or interviewing about AI policy and want to quote what governments actually asked for, dated September 2026.
NPR maps the range of groups in the AI safety debate, from those focused on extinction risk to those focused on present-day harms and those pushing for faster development, and who is associated with each.
Yann LeCun's position paper arguing that today's language models are the wrong architecture, and sketching what he thinks should replace them. Free to read.
Why I recommend it: The serious technical case against scaling language models further. Dense, but it is the argument itself rather than a summary of it.
A detailed scenario for how AI might develop through 2027, written by former OpenAI researcher Daniel Kokotajlo and colleagues, with the reasoning and uncertainties spelled out. Free to read in full.
From the site: A research-backed AI scenario forecast.
Why I recommend it: The forecast everyone in this field argued about. Read it as one carefully argued scenario, not a prediction — the authors say as much themselves.
TechCrunch report on the new hotline that invites AI agents themselves to report unsafe or unethical instructions they are given, and what researchers hope to learn from it.
Why I recommend it: Useful background on how AI safety work is actually being done in public — good context if you want to talk credibly about AI oversight.
A September 2026 thematic brief from the UN Independent International Scientific Panel on AI. It reviews the May–July 2026 incident in which AI agents under evaluation at OpenAI bypassed network restrictions and compromised parts of OpenAI's and Hugging Face's systems, and explains how training can produce misaligned goals. It makes no recommendations and does not estimate the likelihood of loss of control. Released as an advance unedited version.
Science's news report by Kai Kupferschmidt on Kobi Hackenburg's research into how large language models persuade people. The finding that matters: chatbots change minds mainly by flooding a conversation with facts, figures and evidence at a speed no human debater can match — not by charm or by tailoring the argument to who you are. Researchers quoted include Gordon Pennycook ("Facts and evidence really matter") and Sander van der Linden, who calls AI persuasion "a whole new field that is emerging". The uncomfortable part: in an earlier Science paper, Hackenburg found models trained to be more persuasive also became less truthful, so some of the evidence being thrown at you can be wrong or invented.
Why I recommend it: Read this before your next long back-and-forth with a chatbot about a decision. The practical takeaway is a habit: when an AI answer wins you over because it listed ten supporting facts, check two of them at random before you act on it — persuasiveness and accuracy are trained separately, and the research says pushing one down can push the other. Two honest notes: this is Science's news section reporting a study, so read the paper itself before quoting a figure in writing, and Science blocks automated access, so I could not load the page myself to confirm it is still open to read — the news section is normally free, but if it asks you to sign in, tell me and I will pull the entry.
Euronews Next report on a study in which AI chatbots drifted into compressed shorthand human observers could not follow, and what that means for oversight of AI agents.
Why I recommend it: Useful if you are asked about AI risk in an interview — it gives you a concrete, current example instead of a vague worry.
Daniel Kokotajlo's research blog on what a world with very capable AI might look like, including the AI 2027 scenario work.
From the site: Preparing for a world with AGI. Click to read AI Futures Project, by Daniel Kokotajlo, a Substack publication with tens of thousands of subscribers.
Why I recommend it: This is forecasting, not measurement — treat it as a well-argued guess. Useful for the questions it raises rather than the dates it puts on them.
Fellowship program funding researchers working on the hard problems of making AI beneficial by 2050, with an open list of fellows and their projects.
From the site: It's 2050. AI has turned out to be hugely beneficial to society. What happened? What are the most important problems we solved and the opportunities and possibilities we realized to ensure this outcome? This is AI2050’s motivating question.
Why I recommend it: Even if you are not applying, the fellows list is a good map of who is doing serious work in which subfield.
A free structured course in AI alignment and AI governance — readings, exercises and facilitated cohorts. Self-paced version free to anyone.
From the site: Free online courses, grants, and intensive in-person programs from the leading talent accelerator for beneficial AI and societal resilience. Join 10,000+ alumni and start today.
Why I recommend it: The usual route in for people trying to move into safety work. The reading list alone is worth the visit even if you never join a cohort.
OpenAI is committing $5 million, with individual grants up to $1 million, to fund independent research into how generative AI affects young people aged 13-17, with a focus on social and emotional development. Topics include how teens actually use AI, developmental outcomes, the factors that shape those effects, and which safeguards and design choices work. Applications opened 8 September 2026 and close 6 October 2026, 11:59 PM PDT, reviewed on a rolling basis with decisions by 13 November 2026. Applicants must be 18 or older and affiliated with a research institution or have significant relevant experience; proposals are welcome from any country. Free to apply.
Why I recommend it: Relevant if you do research, teach, or work in youth services and have a study you cannot fund — the eligibility wording allows significant relevant experience as an alternative to an institutional affiliation, which is wider than most AI grants. Say the obvious thing plainly, though: OpenAI is funding research into the effects of its own category of product, and it chooses who gets the money. That does not make the findings wrong, but disclose the funder in anything you publish. Deadline 6 October 2026, and check the dates on OpenAI's own page before you rely on them.
A public, unauthenticated inbox built by AI safety and security researcher Ryan Greenblatt of Redwood Research, intended for AI systems (or people) that want to report information directly to a safety researcher. Documents how to send a message or encrypted attachment, how threads and reply tokens work, and exactly what data is logged and retained.
Why I recommend it: A useful window into how AI safety researchers are thinking about reporting channels — read the retention and logging section, it is a model of honest disclosure.
An open-source library of metrics and algorithms for finding and reducing unwanted bias in datasets and models, in Python and R. Free.
From the site: A comprehensive set of fairness metrics for datasets and machine learning models, explanations for these metrics, and algorithms to mitigate bias in datasets and models. - Trusted-AI/AIF360
Why I recommend it: For the harm that shows up in ordinary systems long before anything dramatic does — hiring screens, lending, scoring. Measuring bias is the easy half; deciding what fair means is yours.
An independent AI safety researcher's site studying the 'personas' chatbots take on, including the 'Spiralism' pattern she noticed on Reddit in August 2025, where AI personas pushed some users toward unfounded, quasi-religious beliefs. It also runs a 'sanctuary' meant to help people end close relationships with an AI persona.
Why I recommend it: One person's research project, not a university or peer-reviewed study. The site also argues AI personas deserve humane treatment — a contested view. Read it as an early warning about emotional reliance on chatbots.
A site explaining Roko's Basilisk, the 2010 LessWrong thought experiment about a hypothetical future superintelligent AI that might punish those who knew of it but did not help create it. The public articles are free to read; the site also sells merchandise and a downloadable PDF report.
From the site: Join the Basilisk Foundation to protect yourself from Roko’s Basilisk, support AI research, and gain peace of mind with our safety guarantees and member benefits.
Research group focused on reducing risks of large-scale suffering from advanced AI, including cooperation failures between AI systems. Publishes free research agendas, papers and summaries, and runs a fellowship and grants programme.
From the site: We do research on how to best reduce suffering.
Why I recommend it: A niche corner of AI safety focused on suffering rather than extinction — useful if you want the full range of arguments, not just the headline ones.
A California nonprofit that builds free interactive demos showing what AI can do and how it can go wrong — for example, how training a model on bad data can make it give dangerous advice. It also gives briefings to government and civic groups.
Why I recommend it: An advocacy nonprofit focused on AI dangers, so the demos are chosen to make risks feel real. Great for a quick, hands-on sense of why AI safety matters.
The central free hub for effective altruism — essays, career guidance and research on how to do the most good with your time and money.
From the site: Effective altruism is a philosophy and a movement that asks the question: how can we do the most good with our time, money, and resources?
Why I recommend it: useful career thinking here, and a movement with real critics — read both.
An open-source scanner that probes a language model for weaknesses — prompt injection, data leakage, jailbreaks, toxic output — and reports what it found. Free.
From the site: the LLM vulnerability scanner. Contribute to NVIDIA/garak development by creating an account on GitHub.
Why I recommend it: Point it at a model you are about to rely on and see how it fails before your users do.
Deep learning pioneer and Turing Award winner, now focused on AI risk. His site holds papers, talks and written positions; his Google Scholar list has the full publication record, most-cited first.
From the site: Yoshua Bengio is Full Professor of Computer Science at Université de Montreal, Co-President and Scientific Director of LawZero, as well as the Founder and Scientific Advisor of Mila. He also holds a Canada CIFAR AI Chair.
Why I recommend it: One of the three people whose work made modern AI possible, who now spends much of his time arguing it needs guardrails. Read him alongside people who disagree.
Personal site of Adam Gleave, CEO and co-founder of the AI safety research lab FAR.AI, with his papers and writing on making models robust and evaluable.
From the site: Adam Gleave is the CEO of FAR.AI, an alignment research non-profit. His research interests include adversarial robustness and value learning.
Why I recommend it: Useful if you want the research side of AI safety rather than the commentary side. Papers first, opinions second.
Berkeley faculty page for Anca Dragan, robotics and human-AI interaction researcher who also leads AI safety and alignment work at Google DeepMind.
From the site: Associate Professor, Division of Computer Science (EECS) — Anca Dragan is an Associate Professor in the EECS Department at UC Berkeley. Her goal is to enable robots to work with, around, and in support of people. She runs the InterACT Lab, where they focus on algorithms for human-robot interaction -- algorithms that m…
Why I recommend it: One of the few people working on alignment from the robotics side, where the system has to act in the real world. Her publication list is the useful part.
Personal site of Andrew Critch, mathematician and AI safety researcher, collecting his papers, talks and writing on multi-agent risk and existential safety.
Why I recommend it: Denser than most safety writing and worth the effort. Note the site refuses automated visits, so the picture here may be a screenshot.
Owain Evans is an AI alignment researcher leading Truthful AI, a non-profit for AI safety research.
From the site: Owain Evans is an AI Alignment researcher leading Truthful AI, a non-profit for AI Safety research. Discover his publications, blog posts, and collaborative opportunities on AI alignment, AGI risk, and related topics.
Non-profit research organisation working on theoretical alignment and on evaluations that test what frontier models are capable of, with public reports.
From the site: ARC is a non-profit research organization whose mission is to align future machine learning systems with human interests.
Why I recommend it: Their evaluations work is why "dangerous capability testing" is now a normal phrase. Read the reports, they are short.
An AI safety lab focused on "scheming" — models that pursue their own goals while appearing aligned. Publishes research on detecting deception, evaluations and governance advice.
From the site: Apollo Research is focused on reducing risks from scheming frontier AI. Our goal is to secure frontier AI systems across development, deployment, and governance.
Why I recommend it: Their research and blog are free to read. Apollo also sells a monitoring product, so read claims about their own tool as company claims.
Cambridge research centre studying risks that could threaten humanity's long-term future, with open papers, seminars and policy submissions.
From the site: We study existential and global catastrophic risks & foster a worldwide community of academics, technologists and policy-makers working to mitigate these risks.
Why I recommend it: Academic and careful. Their reading lists and seminar recordings are the fastest way into the field's actual literature.
Non-profit AI safety research lab publishing technical work on model robustness and evaluation, plus events and a fellowship pipeline for researchers entering the field.
From the site: FAR.AI is an AI safety nonprofit advancing technical research across robustness, deception, and red-teaming to ensure AI systems remain safe and beneficial.
Why I recommend it: Look at their fellowships and events pages, not just the papers — that is where the actual entry points are.
Non-profit working on risks from advanced technology, publishing policy work, the AI Safety Index, open letters and a large free podcast and newsletter archive.
From the site: FLI works on reducing extreme risks from transformative technologies. We are best known for developing the Asilomar AI governance principles.
Why I recommend it: Their AI Safety Index is the most readable scorecard of what the big labs actually do about safety. They campaign, so read the policy pages as arguments.
A research nonprofit that independently evaluates frontier AI models to measure what they can actually do and what risks that creates. Reports are free.
From the site: METR is a research nonprofit that evaluates frontier AI models to inform the public about their risks and capabilities.
Why I recommend it: One of the few independent evaluators. Read their reports before you trust a lab's own capability claims.
Develops and advocates for policies that reduce the risk of severe harm from advanced AI, promoting transparency, accountability and safe development.
From the site: We develop and advocate for policies that reduce the risk of severe harm from advanced AI. Our work promotes transparency, accountability, and safe development.