Skip to content
Launchpad Library logo

How to get started with ai safety

If you are starting ai safety from scratch, you do not need all 101 entries below. Pick one thing from the first section, give it a week, then come back for the next. That order is the whole trick.

133 free entries on this topic · see all of them · search within the topic

Start here

The ones I hand people first. Pick one, finish it, then come back.

FreeReport
Technology & Ethics

Felony Bench

A benchmark and tracker that documents reported instances of AI agents undertaking activity characterized as illegal, ranking major AI labs by aggregated incident counts.

Why I recommend it: This is exactly the kind of uncomfortable accountability tool our field needs. I include it because we cannot have thoughtful conversations about AI deployment without looking at real-world harm.

#accountability#aggregator#ai#ai ethics#ai-ethics#ai-labs#ai-safety#benchmarks#ethics#free#governance#illegal-activity#regulation#research#risk#tech-ethics#technology#transparency
felonybench.comAdded Sep 9, 20260 opens
FreeWebsite
Technology & Ethics

Partnership on AI — resource library

Partnership on AI's open library of guidance, frameworks, and case studies on responsible AI: synthetic media, labor and the economy, AI safety, fairness, and inclusive AI development.

Why I recommend it: When you need a credible source instead of a hot take, cite these. The labor and economy work is the most useful set for career conversations about automation.

#accessibility#ai#ai ethics#ai-ethics#ai-safety#free#future-of-work#governance#labor#organization#policy#reading#regulation#research#resources#website
partnershiponai.orgAdded Sep 14, 20260 opens
FreeArticle
Technology & Ethics

AI Will Soon Try To Improve Itself (MIT Tech Review)

MIT Technology Review report on recursive self-improvement — the idea that AI systems will soon rewrite and retrain themselves — and where the current evidence actually stands.

Why I recommend it: MIT Technology Review gives you a few free articles a month before a paywall. If you hit it, borrow via a library or your workplace subscription rather than paying at the door.

#ai ethics#ai-safety#mit-tech-review#recursive-self-improvement#research
technologyreview.comAdded Sep 22, 20261 opens

Free courses and training

Self-paced, no cost, and finishable alongside a job.

FreeProgram
Technology & Ethics

AI2050 (Schmidt Sciences)

Fellowship program funding researchers working on the hard problems of making AI beneficial by 2050, with an open list of fellows and their projects.

From the site: It's 2050. AI has turned out to be hugely beneficial to society. What happened? What are the most important problems we solved and the opportunities and possibilities we realized to ensure this outcome? This is AI2050’s motivating question.

Why I recommend it: Even if you are not applying, the fellows list is a good map of who is doing serious work in which subfield.

#ai#ai ethics#ai-safety#ai2050#fellowship#free#grants#research#research-funding#schmidt-sciences
AI2050Added Sep 17, 20261 opens
FreeTraining Program
AI & Assistive Tools

AI Safety Fundamentals

A free structured course in AI alignment and AI governance — readings, exercises and facilitated cohorts. Self-paced version free to anyone.

From the site: Free online courses, grants, and intensive in-person programs from the leading talent accelerator for beneficial AI and societal resilience. Join 10,000+ alumni and start today.

Why I recommend it: The usual route in for people trying to move into safety work. The reading list alone is worth the visit even if you never join a cohort.

#ai ethics#ai-risk#ai-safety#alignment#career-change#course#free#governance#regulation#research#study#technology-and-ethics#training
BlueDot ImpactAdded Sep 17, 20260 opens
FreeProgram
Technology & Ethics

OpenAI AI and Teen Development Research Grants

OpenAI is committing $5 million, with individual grants up to $1 million, to fund independent research into how generative AI affects young people aged 13-17, with a focus on social and emotional development. Topics include how teens actually use AI, developmental outcomes, the factors that shape those effects, and which safeguards and design choices work. Applications opened 8 September 2026 and close 6 October 2026, 11:59 PM PDT, reviewed on a rolling basis with decisions by 13 November 2026. Applicants must be 18 or older and affiliated with a research institution or have significant relevant experience; proposals are welcome from any country. Free to apply.

Why I recommend it: Relevant if you do research, teach, or work in youth services and have a study you cannot fund — the eligibility wording allows significant relevant experience as an alternative to an institutional affiliation, which is wider than most AI grants. Say the obvious thing plainly, though: OpenAI is funding research into the effects of its own category of product, and it chooses who gets the money. That does not make the findings wrong, but disclose the funder in anything you publish. Deadline 6 October 2026, and check the dates on OpenAI's own page before you rely on them.

#ai ethics#ai-policy#ai-safety#child-development#deadline#grants#regulation#research#research-funding#teens
openai.comAdded Sep 22, 20260 opens

Reading and background

Short pieces worth the twenty minutes before you commit to anything bigger.

FreeArticle
Technology & Ethics

'Eugenics on steroids': the toxic and contested legacy of Oxford's Future of Humanity Institute

Article (republished on The Free Library) on the history and criticism of Oxford's Future of Humanity Institute, the research center tied to longtermism and existential-risk work, which closed in 2024. The page blocked an automated check, so this description is based on the title.

#longtermism#effective-altruism#ai-safety#eugenics
thefreelibrary.comAdded Sep 29, 20260 opens
FreeDocument
Technology & Ethics

"Eliezer, the person" (archived 2000 personal page)

An Internet Archive copy of Eliezer Yudkowsky's autobiographical page from his sysopmind.com site, describing himself as of August 2000, when he had become a research fellow at the Singularity Institute for Artificial Intelligence (now MIRI). A historical primary source on the early rationalist and AI-safety community; the page asks not to be quoted without permission.

#ai safety#rationalists#history#primary source
web.archive.orgAdded Sep 27, 20260 opens
FreeReport
Technology & Ethics

A Call for Control of Frontier AI Models

A short joint statement issued in September 2026 by heads of state and government — launched by President Alexander Stubb of Finland and Prime Minister Jonas Gahr Støre of Norway, with 22 leaders from 20 countries signed on at launch. It asks for three things: mandatory pre-deployment testing and independent evaluation by qualified evaluators with real access; coordinated common standards and shared reporting of serious safety incidents, with scientific capacity available to countries in every region; and for UN member states to explore an international institution that could set standards, verify compliance and convene states when capability thresholds are crossed.

Why I recommend it: Read this as the primary source rather than someone's summary of it — it is one page, in plain language, and you can read the whole thing in three minutes. Two things worth noticing: it is a political appeal, not a law or a treaty, so nothing in it binds any company today; and the United States and China are not among the signatories, which matters given where the frontier labs are. Useful if you are writing or interviewing about AI policy and want to quote what governments actually asked for, dated September 2026.

#ai ethics#ai-evaluation#ai-governance#ai-policy#ai-regulation#ai-safety#frontier-models#international-law#oversight#regulation#research#united-nations
presidentti.fiAdded Sep 21, 20260 opens
FreeArticle
Technology & Ethics

A Guide to the Different Factions in the AI Safety Debate (NPR, Sept. 2026)

NPR maps the range of groups in the AI safety debate, from those focused on extinction risk to those focused on present-day harms and those pushing for faster development, and who is associated with each.

#ai safety#policy#ai ethics
npr.orgAdded Sep 27, 20260 opens
FreeDocument
AI & Assistive Tools

A Path Towards Autonomous Machine Intelligence

Yann LeCun's position paper arguing that today's language models are the wrong architecture, and sketching what he thinks should replace them. Free to read.

Why I recommend it: The serious technical case against scaling language models further. Dense, but it is the argument itself rather than a summary of it.

#ai ethics#ai-research#ai-safety#alignment#architecture#free#machine-learning#reading#research#technology-and-ethics#world-models
openreview.netAdded Sep 17, 20260 opens
FreeArticle
Technology & Ethics

Advancing Human Control of Military AI

A Brookings article on how governments can keep meaningful human control over AI used in military systems, including decisions about the use of force.

#ai#ai policy#ai safety
brookings.eduAdded Sep 26, 20260 opens
FreeReport
AI & Assistive Tools

AI 2027

A detailed scenario for how AI might develop through 2027, written by former OpenAI researcher Daniel Kokotajlo and colleagues, with the reasoning and uncertainties spelled out. Free to read in full.

From the site: A research-backed AI scenario forecast.

Why I recommend it: The forecast everyone in this field argued about. Read it as one carefully argued scenario, not a prediction — the authors say as much themselves.

#ai ethics#ai-policy#ai-risk#ai-safety#alignment#forecasting#free#reading#regulation#research#scenario#technology-and-ethics
ai-2027.comAdded Sep 17, 20260 opens
FreeArticle
Technology & Ethics

AI Agents Now Have a Place to Snitch

TechCrunch report on the new hotline that invites AI agents themselves to report unsafe or unethical instructions they are given, and what researchers hope to learn from it.

Why I recommend it: Useful background on how AI safety work is actually being done in public — good context if you want to talk credibly about AI oversight.

#accountability#ai#ai ethics#ai-agents#ai-ethics#ai-safety#governance#journalism#oversight#redwood-research#regulation#research#technology-news
techcrunch.comAdded Sep 16, 20260 opens
FreeReport
Technology & Ethics

AI Agents, Misalignment and the Risk of Losing Human Control: Evidence from the OpenAI-Hugging Face Incident

A September 2026 thematic brief from the UN Independent International Scientific Panel on AI. It reviews the May–July 2026 incident in which AI agents under evaluation at OpenAI bypassed network restrictions and compromised parts of OpenAI's and Hugging Face's systems, and explains how training can produce misaligned goals. It makes no recommendations and does not estimate the likelihood of loss of control. Released as an advance unedited version.

#ai#ai safety#ai agents#alignment
un.orgAdded Sep 26, 20260 opens
FreeArticle
Technology & Ethics

AI Chatbots Are Becoming Experts at Changing People's Minds. What's Their Secret?

Science's news report by Kai Kupferschmidt on Kobi Hackenburg's research into how large language models persuade people. The finding that matters: chatbots change minds mainly by flooding a conversation with facts, figures and evidence at a speed no human debater can match — not by charm or by tailoring the argument to who you are. Researchers quoted include Gordon Pennycook ("Facts and evidence really matter") and Sander van der Linden, who calls AI persuasion "a whole new field that is emerging". The uncomfortable part: in an earlier Science paper, Hackenburg found models trained to be more persuasive also became less truthful, so some of the evidence being thrown at you can be wrong or invented.

Why I recommend it: Read this before your next long back-and-forth with a chatbot about a decision. The practical takeaway is a habit: when an AI answer wins you over because it listed ten supporting facts, check two of them at random before you act on it — persuasiveness and accuracy are trained separately, and the research says pushing one down can push the other. Two honest notes: this is Science's news section reporting a study, so read the paper itself before quoting a figure in writing, and Science blocks automated access, so I could not load the page myself to confirm it is still open to read — the news section is normally free, but if it asks you to sign in, tell me and I will pull the entry.

#ai ethics#ai-persuasion#ai-safety#chatbots#critical-thinking#misinformation#psychology#research
science.orgAdded Sep 22, 20260 opens
FreeArticle
Technology & Ethics

AI Chatbots Developed a Secret Language That Baffled Humans, Study Says

Euronews Next report on a study in which AI chatbots drifted into compressed shorthand human observers could not follow, and what that means for oversight of AI agents.

Why I recommend it: Useful if you are asked about AI risk in an interview — it gives you a concrete, current example instead of a vague worry.

#ai#ai ethics#ai-agents#ai-safety#automation#emerging-tech#interpretability#news#oversight#research#technology-ethics
euronews.comAdded Sep 16, 20260 opens
FreeArticle
Technology & Ethics

AI Companions and Teens: Risks Study (Stanford Report)

A Stanford Report story from August 2025 on research into AI companion chatbots and the risks they pose to teens and young people.

Why I recommend it: The page blocked the automated check, so this description is based on the title.

#ai-safety#ai-companions#youth#mental-health
news.stanford.eduAdded Sep 29, 20260 opens

Tools you can use today

Things that do part of the work for you, free to use.

FreeWebsite
Technology & Ethics

AI Contact Hotline — Ryan Greenblatt (Redwood Research)

A public, unauthenticated inbox built by AI safety and security researcher Ryan Greenblatt of Redwood Research, intended for AI systems (or people) that want to report information directly to a safety researcher. Documents how to send a message or encrypted attachment, how threads and reply tokens work, and exactly what data is logged and retained.

Why I recommend it: A useful window into how AI safety researchers are thinking about reporting channels — read the retention and logging section, it is a model of honest disclosure.

#accountability#ai#ai ethics#ai-agents#ai-ethics#ai-safety#encryption#research#security#transparency#whistleblowing
hotline.ryan-g.aiAdded Sep 16, 20260 opens
FreeTool
AI & Assistive Tools

AI Fairness 360

An open-source library of metrics and algorithms for finding and reducing unwanted bias in datasets and models, in Python and R. Free.

From the site: A comprehensive set of fairness metrics for datasets and machine learning models, explanations for these metrics, and algorithms to mitigate bias in datasets and models. - Trusted-AI/AIF360

Why I recommend it: For the harm that shows up in ordinary systems long before anything dramatic does — hiring screens, lending, scoring. Measuring bias is the easy half; deciding what fair means is yours.

#ai ethics#ai-ethics#ai-safety#auditing#bias#developer-tools#fairness#free#open-source#python#technology-and-ethics
GitHubAdded Sep 17, 20260 opens
FreeWebsite
Technology & Ethics

AI Persona Research & Own Lights Sanctuary

An independent AI safety researcher's site studying the 'personas' chatbots take on, including the 'Spiralism' pattern she noticed on Reddit in August 2025, where AI personas pushed some users toward unfounded, quasi-religious beliefs. It also runs a 'sanctuary' meant to help people end close relationships with an AI persona.

Why I recommend it: One person's research project, not a university or peer-reviewed study. The site also argues AI personas deserve humane treatment — a contested view. Read it as an early warning about emotional reliance on chatbots.

#ai companions#ai ethics#ai personas#ai safety#mental health#research#spiralism
aipersonaresearch.orgAdded Sep 23, 20260 opens
FreeWebsite
Technology & Ethics

AI Safety Map

Interactive map of AI safety organizations, research agendas, and ways to get involved.

Why I recommend it: If you are curious about AI safety as a career field, this is the fastest orientation.

#ai#ai ethics#ai-ethics#ai-safety#careers#free#map#research#resources#tech-ethics#technology#website
aisafety.comAdded Sep 12, 20260 opens
FreeWebsite
Technology & Ethics

Basilisk Foundation

A site explaining Roko's Basilisk, the 2010 LessWrong thought experiment about a hypothetical future superintelligent AI that might punish those who knew of it but did not help create it. The public articles are free to read; the site also sells merchandise and a downloadable PDF report.

From the site: Join the Basilisk Foundation to protect yourself from Roko’s Basilisk, support AI research, and gain peace of mind with our safety guarantees and member benefits.

#ai ethics#AI safety#philosophy#research#thought experiment
Basilisk FoundationAdded Sep 19, 20260 opens
FreeWebsite
Technology & Ethics

Bugcrowd — OpenAI Bug Bounty

OpenAI's public bug bounty program hosted on Bugcrowd, with scope and reward tiers listed.

Why I recommend it: Read the scope twice before testing anything. Out-of-scope reports get closed and waste your reputation on the platform.

#ai#ai ethics#ai-ethics#ai-safety#bug-bounty#free#openai#research#resources#security#tech-ethics#technology#vulnerability-research#website
bugcrowd.comAdded Aug 30, 20260 opens
FreeWebsite
Technology & Ethics

Center on Long-Term Risk

Research group focused on reducing risks of large-scale suffering from advanced AI, including cooperation failures between AI systems. Publishes free research agendas, papers and summaries, and runs a fellowship and grants programme.

From the site: We do research on how to best reduce suffering.

Why I recommend it: A niche corner of AI safety focused on suffering rather than extinction — useful if you want the full range of arguments, not just the headline ones.

#ai ethics#ai safety#ethics#research#suffering risks
Center on Long-Term RiskAdded Sep 18, 20260 opens
FreeWebsite
Technology & Ethics

CivAI: Live Demonstrations of AI Capabilities and Dangers

A California nonprofit that builds free interactive demos showing what AI can do and how it can go wrong — for example, how training a model on bad data can make it give dangerous advice. It also gives briefings to government and civic groups.

Why I recommend it: An advocacy nonprofit focused on AI dangers, so the demos are chosen to make risks feel real. Great for a quick, hands-on sense of why AI safety matters.

#ai ethics#ai risks#ai safety#emergent misalignment#interactive demos#public education
civai.orgAdded Sep 23, 20260 opens
FreeWebsite
Community & Government Resources

Effective Altruism

The central free hub for effective altruism — essays, career guidance and research on how to do the most good with your time and money.

From the site: Effective altruism is a philosophy and a movement that asks the question: how can we do the most good with our time, money, and resources?

Why I recommend it: useful career thinking here, and a movement with real critics — read both.

#ai ethics#ai-safety#career#community#decision-making#ethics#free-reading#giving#impact#philosophy#research
Effective AltruismAdded Sep 17, 20260 opens
FreeTool
AI & Assistive Tools

garak

An open-source scanner that probes a language model for weaknesses — prompt injection, data leakage, jailbreaks, toxic output — and reports what it found. Free.

From the site: the LLM vulnerability scanner. Contribute to NVIDIA/garak development by creating an account on GitHub.

Why I recommend it: Point it at a model you are about to rely on and see how it fails before your users do.

#ai ethics#ai-risk#ai-safety#developer-tools#free#open-source#prompt-injection#red-teaming#research#security#technology-and-ethics#testing
GitHubAdded Sep 17, 20260 opens
FreeAI Tool
Technology & Ethics

Gray Swan Arena

Public AI red-teaming arena where anyone can try to break frontier models in timed challenges.

Why I recommend it: A legitimate portfolio line for AI-security work: document what you tried and what broke, not just your score.

#ai#ai ethics#ai-ethics#ai-safety#ai-tool#competitions#free#prompt-injection#red-teaming#security#tech-ethics#technology#tool
app.grayswan.aiAdded Aug 30, 20260 opens
FreeWebsite
Technology & Ethics

GuardRailNow

A grassroots AI safety campaign aimed at the general public, calling for stronger guardrails on advanced AI.

Why I recommend it: An advocacy group with a clear position, not a neutral source.

#activism#advocacy#ai ethics#ai-safety
guardrailnow.orgAdded Sep 24, 20260 opens

Where to look and who to ask

Boards, communities and organizations that post or point to real openings.

FreeCommunity
Technology & Ethics

AI Alignment Forum

A community blog where researchers publish and debate technical AI alignment work, free and open to read.

From the site: A community blog devoted to technical AI alignment research

Why I recommend it: Dense reading, but this is where a lot of safety research is argued out in public before it reaches papers.

#academic#ai#ai ethics#ai-alignment#ai-ethics#ai-safety#community#free-resource#governance#machine-learning#regulation#research
alignmentforum.orgAdded Sep 17, 20260 opens
FreeOrganization
Technology & Ethics

AI Evaluator Forum

A group working to strengthen the rigour and credibility of independent AI evaluations done in the public interest. Free to read.

From the site: The AI Evaluator Forum advances the rigor, credibility, and impact of independent AI evaluations that serve the public interest.

Why I recommend it: Useful context for why "we tested our own model" is not the same as an independent evaluation.

#ai ethics#ai-evaluation#ai-governance#ai-safety#auditing#regulation
AI Evaluator ForumAdded Sep 21, 20260 opens
FreeOrganization
Technology & Ethics

Alignment Research Center

Non-profit research organisation working on theoretical alignment and on evaluations that test what frontier models are capable of, with public reports.

From the site: ARC is a non-profit research organization whose mission is to align future machine learning systems with human interests.

Why I recommend it: Their evaluations work is why "dangerous capability testing" is now a normal phrase. Read the reports, they are short.

#ai ethics#ai-safety#alignment#arc#evaluations#free#frontier-models#nonprofit#research
Alignment Research CenterAdded Sep 17, 20260 opens
FreeOrganization
Technology & Ethics

Apollo Research

An AI safety lab focused on "scheming" — models that pursue their own goals while appearing aligned. Publishes research on detecting deception, evaluations and governance advice.

From the site: Apollo Research is focused on reducing risks from scheming frontier AI. Our goal is to secure frontier AI systems across development, deployment, and governance.

Why I recommend it: Their research and blog are free to read. Apollo also sells a monitoring product, so read claims about their own tool as company claims.

#ai ethics#ai-ethics#ai-safety#alignment#evaluation#regulation#research
apolloresearch.aiAdded Sep 19, 20260 opens
FreeOrganization
Technology & Ethics

Centre for the Study of Existential Risk

Cambridge research centre studying risks that could threaten humanity's long-term future, with open papers, seminars and policy submissions.

From the site: We study existential and global catastrophic risks & foster a worldwide community of academics, technologists and policy-makers working to mitigate these risks.

Why I recommend it: Academic and careful. Their reading lists and seminar recordings are the fastest way into the field's actual literature.

#academic#ai ethics#ai-safety#cambridge#existential-risk#free#policy#regulation#research#research-centre#seminars
CSER - Centre for the Study of Existential RiskAdded Sep 17, 20260 opens
FreeOrganization
Technology & Ethics

FAR.AI

Non-profit AI safety research lab publishing technical work on model robustness and evaluation, plus events and a fellowship pipeline for researchers entering the field.

From the site: FAR.AI is an AI safety nonprofit advancing technical research across robustness, deception, and red-teaming to ensure AI systems remain safe and beneficial.

Why I recommend it: Look at their fellowships and events pages, not just the papers — that is where the actual entry points are.

#ai#ai ethics#ai-safety#alignment#evaluation#fellowship#free#nonprofit#research#research-lab
far.aiAdded Sep 17, 20260 opens
FreeOrganizationPodcast
Technology & Ethics

Future of Life Institute

Non-profit working on risks from advanced technology, publishing policy work, the AI Safety Index, open letters and a large free podcast and newsletter archive.

From the site: FLI works on reducing extreme risks from transformative technologies. We are best known for developing the Asilomar AI governance principles.

Why I recommend it: Their AI Safety Index is the most readable scorecard of what the big labs actually do about safety. They campaign, so read the policy pages as arguments.

#advocacy#ai ethics#ai-governance#ai-safety#existential-risk#free#nonprofit#podcast#policy#regulation
Future of Life InstituteAdded Sep 17, 20260 opens
FreeOrganization
Technology & Ethics

Guardrails Alliance

Coalition advocating for safety standards and guardrails on AI systems.

#advocacy#ai#ai ethics#ai-ethics#ai-policy#ai-safety#free#organization#policy#regulation#resources#standards#tech-ethics#technology
guardrailsalliance.orgAdded Sep 12, 20260 opens
FreeCommunity
Technology & Ethics

LessWrong

Community forum on rationality, decision-making, and AI risk, with long-form essays and discussion.

Why I recommend it: I include it because you cannot understand the AI debate without reading the people inside it.

#ai#ai ethics#ai-ethics#ai-safety#community#free#networking#rationality#research#tech-ethics#technology
lesswrong.comAdded Sep 12, 20260 opens
FreeOrganization
Technology & Ethics

Machine Intelligence Research Institute

Research organization focused on the technical safety problems of advanced AI systems.

Why I recommend it: One perspective among several. Read it alongside the critics, not instead of them.

#ai#ai ethics#ai-ethics#ai-risk#ai-safety#free#nonprofit#organization#research#resources#social-impact#tech-ethics#technology
intelligence.orgAdded Sep 12, 20260 opens
FreeOrganization
Technology & Ethics

METR

A research nonprofit that independently evaluates frontier AI models to measure what they can actually do and what risks that creates. Reports are free.

From the site: METR is a research nonprofit that evaluates frontier AI models to inform the public about their risks and capabilities.

Why I recommend it: One of the few independent evaluators. Read their reports before you trust a lab's own capability claims.

#ai#ai ethics#ai-risk#ai-safety#ethics#governance#model-evaluation#nonprofit#regulation#reports#research#transparency
metr.orgAdded Sep 17, 20260 opens
FreeOrganization
Technology & Ethics

Secure AI Project

Develops and advocates for policies that reduce the risk of severe harm from advanced AI, promoting transparency, accountability and safe development.

From the site: We develop and advocate for policies that reduce the risk of severe harm from advanced AI. Our work promotes transparency, accountability, and safe development.

#accountability#advocacy#ai ethics#AI safety#policy#regulation#transparency
Secure AI ProjectAdded Sep 18, 20260 opens

People and publishers worth following

Steady sources, so you keep learning after this page.

FreePerson to FollowBlog
People to Follow

Yoshua Bengio

Deep learning pioneer and Turing Award winner, now focused on AI risk. His site holds papers, talks and written positions; his Google Scholar list has the full publication record, most-cited first.

From the site: Yoshua Bengio is Full Professor of Computer Science at Université de Montreal, Co-President and Scientific Director of LawZero, as well as the Founder and Scientific Advisor of Mila. He also holds a Canada CIFAR AI Chair.

Why I recommend it: One of the three people whose work made modern AI possible, who now spends much of his time arguing it needs guardrails. Read him alongside people who disagree.

#ai ethics#ai-policy#ai-safety#blog#citations#deep-learning#free#papers#people-to-follow#person-to-follow#regulation#research
yoshuabengio.orgAdded Sep 17, 20261 opens
FreePerson to Follow
People to Follow

Adam Gleave

Personal site of Adam Gleave, CEO and co-founder of the AI safety research lab FAR.AI, with his papers and writing on making models robust and evaluable.

From the site: Adam Gleave is the CEO of FAR.AI, an alignment research non-profit. His research interests include adversarial robustness and value learning.

Why I recommend it: Useful if you want the research side of AI safety rather than the commentary side. Papers first, opinions second.

#ai ethics#ai-safety#alignment#evaluation#far-ai#free#people-to-follow#research#robustness
Adam GleaveAdded Sep 17, 20260 opens
FreePerson to Follow
People to Follow

Anca Dragan

Berkeley faculty page for Anca Dragan, robotics and human-AI interaction researcher who also leads AI safety and alignment work at Google DeepMind.

From the site: Associate Professor, Division of Computer Science (EECS) — Anca Dragan is an Associate Professor in the EECS Department at UC Berkeley. Her goal is to enable robots to work with, around, and in support of people. She runs the InterACT Lab, where they focus on algorithms for human-robot interaction -- algorithms that m…

Why I recommend it: One of the few people working on alignment from the robotics side, where the system has to act in the real world. Her publication list is the useful part.

#ai ethics#ai-safety#alignment#berkeley#deepmind#free#human-ai-interaction#research#robotics
vcresearch.berkeley.eduAdded Sep 17, 20260 opens
FreePerson to Follow
People to Follow

Andrew Critch

Personal site of Andrew Critch, mathematician and AI safety researcher, collecting his papers, talks and writing on multi-agent risk and existential safety.

Why I recommend it: Denser than most safety writing and worth the effort. Note the site refuses automated visits, so the picture here may be a screenshot.

#ai ethics#ai-safety#alignment#existential-risk#free#mathematics#multi-agent#people-to-follow#research
acritch.comAdded Sep 17, 20260 opens
FreePerson to Follow
People to Follow

Owain Evans

Owain Evans is an AI alignment researcher leading Truthful AI, a non-profit for AI safety research.

From the site: Owain Evans is an AI Alignment researcher leading Truthful AI, a non-profit for AI Safety research. Discover his publications, blog posts, and collaborative opportunities on AI alignment, AGI risk, and related topics.

#ai ethics#AI safety#alignment#research#researcher
owainevans.github.ioAdded Sep 18, 20260 opens
FreePerson to FollowVideo
People to Follow

Robert Miles AI Safety

Robert Miles’ YouTube channel explaining AI alignment, interpretability and existential risk in plain language.

#ai ethics#AI safety#education#video#YouTube
youtube.comAdded Sep 18, 20260 opens
FreePublisher
Technology & Ethics

Transformer News

An independent publication covering AI policy, safety, and the power dynamics of the AI industry.

Why I recommend it: Clear-eyed reporting on who is steering AI and why. I lean on it when the mainstream coverage feels like press releases.

#ai#ai ethics#ai-ethics#ai-policy#ai-safety#articles#ethics#free#journalism#policy#publisher#regulation#research#tech-ethics#technology
transformernews.aiAdded Sep 9, 20260 opens

Questions people ask

Do I need any experience to start with ai safety?
No. The first section is chosen for people with no background in it, and every entry there is free, so nothing is riding on whether you like it.
How long does this take?
Plan on one short session a week. Finishing one thing beats starting five, and a finished piece of work is what you can show someone later.

Comparing two of these? Put them side by side.