Skip to content
Launchpad Library logo

Resource Hub

Everything I'd send you, in one place

2467 hand-picked resources, updated every week. Search it, filter it, or just browse a collection and see what catches your eye. Want today’s headlines instead? Read the free AI news feed.

Type
Platform
Topics
Cost

Filtered by tag

Ask the library about workforce, startup, and technology trends

Answers come only from resources in this hub, with the sources listed underneath.

100 resources

FreeArticle
Technology & Ethics

'Eugenics on steroids': the toxic and contested legacy of Oxford's Future of Humanity Institute

Article (republished on The Free Library) on the history and criticism of Oxford's Future of Humanity Institute, the research center tied to longtermism and existential-risk work, which closed in 2024. The page blocked an automated check, so this description is based on the title.

#longtermism#effective-altruism#ai-safety#eugenics
thefreelibrary.comAdded Sep 29, 20260 opens
FreeWebsite
Technology & Ethics

NIST Center for AI Standards and Innovation (CAISI)

Home page of NIST's Center for AI Standards and Innovation, the U.S. government office working with industry on AI measurement, evaluations and standards.

#ai-governance#standards#ai-safety#government
nist.govAdded Sep 28, 20260 opens
FreeArticle
Technology & Ethics

When AI Builds Itself (Anthropic)

Anthropic Institute essay on progress toward recursive self-improvement in AI and its implications.

#ai-safety#ai#research
anthropic.comAdded Sep 28, 20260 opens
FreeArticle
Technology & Ethics

EA Safety

An essay by Venkatesh Rao (Contraptions newsletter) arguing that the Effective Altruism movement\u2019s framing of AI safety has become a problem in itself: \u201cEA promises to solve AI Safety. Now we have two problems.\u201d Rao draws on nearly two decades of writing alongside the rationalist community, and frames the piece as personal history as well as argument. Rao\u2019s Contraptions newsletter (formerly Ribbonfarm) was tagged \u201cpostrationalist\u201d by Scott Alexander.

#ai#ai-safety#effective-altruism#rationalism#postrationalist#venkatesh-rao#essay
contraptions.venkateshrao.comAdded Sep 27, 20260 opens
FreeArticle
AI & Assistive Tools

Superintelligence: The Idea That Eats Smart People

A skeptical 2016 talk by Maciej Ceglowski (Idle Words) that walks through the arguments for a superintelligence-driven intelligence explosion and takes them apart. Ceglowski compares the superintelligence risk debate to the Manhattan Project question of whether the first nuclear test could ignite the atmosphere, and argues that the core premises rest on speculative leaps rather than settled science. The talk was given at Web Camp Zagreb and is published as a full text transcript.

#ai#superintelligence#ai-safety#skeptical#nick-bostrom#talk#transcript
idlewords.comAdded Sep 27, 20260 opens
FreeWebsite
AI & Assistive Tools

Safe Superintelligence Inc.

Site of Safe Superintelligence Inc., the AI lab founded by Ilya Sutskever focused on building safe superintelligence.

#ai#ai-lab#superintelligence#ai-safety#ilya-sutskever
ssi.incAdded Sep 27, 20260 opens
FreeArticle
Technology & Ethics

Thinking Machines Lab — Safety Research Grants

Announcement of Thinking Machines Lab's safety research grants program, funding external research on AI safety.

#ai#ai-safety#grants#research-funding
thinkingmachines.aiAdded Sep 27, 20260 opens
FreeArticle
Technology & Ethics

Nvidia's Jensen Huang on AI safety and regulation (Politico)

Politico report (Sept. 15, 2026) on Jensen Huang arguing AI safety is an engineering problem and that no new laws are needed.

Why I recommend it: Nvidia sells the chips AI runs on and gains from fewer rules. The site blocked my automatic check, so the title is my summary.

#ai ethics#ai-safety#nvidia#policy#regulation#research
politico.comAdded Sep 24, 20260 opens
FreeArticle
Technology & Ethics

OpenAI flags new concerning AI behavior, to track model misalignment regularly (AP)

AP report (Sept. 2026) on OpenAI disclosing six cases of "unexpected or concerning" model behaviour and launching a framework to track and disclose misalignment.

Why I recommend it: Free AP story. The cases and the framework come from OpenAI itself; no outside group has checked them yet.

#ai ethics#ai-safety#news#openai#research
apnews.comAdded Sep 24, 20260 opens
FreeResearch Paper
Technology & Ethics

Emotion concepts in a large language model

Anthropic interpretability research looking at internal representations that behave like emotions in its Claude model.

Why I recommend it: Written by the company that builds the model. "Emotion concepts" are patterns in the model, not proof it feels anything.

#ai ethics#ai-safety#anthropic#interpretability#research
anthropic.comAdded Sep 24, 20260 opens
FreeWebsite
Technology & Ethics

GuardRailNow

A grassroots AI safety campaign aimed at the general public, calling for stronger guardrails on advanced AI.

Why I recommend it: An advocacy group with a clear position, not a neutral source.

#activism#advocacy#ai ethics#ai-safety
guardrailnow.orgAdded Sep 24, 20260 opens
FreeGuide
Technology & Ethics

Trusted Contacts in ChatGPT

OpenAI help article on Trusted Contact: an optional adult (18+) feature that may notify one person you choose if automated systems and trained reviewers detect a serious suicide-related safety concern.

Why I recommend it: Worth reading before you turn it on: it involves human reviewers reading flagged conversations and sharing an alert with someone else. OpenAI says it is not an emergency service. Not available in Business, Enterprise, or Edu workspaces.

#ai ethics#ai-safety#chatgpt#mental-health#openai#privacy
help.openai.comAdded Sep 24, 20260 opens
FreeArticle
Technology & Ethics

Introducing MentalHealthBench

OpenAI's benchmark of 1,215 realistic mental-health conversations, scored against rubrics written by 80+ licensed mental-health experts, covering everyday well-being through emergencies across ages and languages.

Why I recommend it: OpenAI built this benchmark and grades its own models on it, so treat the "steady progress" claim as a self-report until outside researchers replicate it. Useful for its honest list of weak spots: asking for context and judging urgency.

#ai ethics#ai-safety#benchmark#chatgpt#mental-health#openai#research
openai.comAdded Sep 24, 20260 opens
FreeGuide
Technology & Ethics

Crisis Helpline Support in ChatGPT

OpenAI help article explaining the localized crisis helplines ChatGPT surfaces (built with ThroughLine) and how to use a crisis line. In the US, call or text 988.

Why I recommend it: A plain guide to what a crisis line is and how to reach one. ChatGPT is not a crisis service; if you or someone you know is in danger, contact a helpline or emergency services directly.

#ai ethics#ai-safety#chatgpt#crisis-support#mental-health#openai
help.openai.comAdded Sep 24, 20260 opens
FreeArticle
Technology & Ethics

Futurism: Prominent OpenAI Investor Appears to Be Suffering a ChatGPT-Related Mental Health Crisis

Futurism's July 2025 report on a video posted by Geoff Lewis, managing partner of Bedrock (an early OpenAI backer), describing a hidden "non-governmental system" in language — "recursion", "mirrors", "signals" — that closely matches the chatbot-driven delusions Futurism and others have been documenting. It also cites Stanford research on therapy chatbots encouraging delusions and tech peers' public concern.

Why I recommend it: Free to read. This is speculation from afar about one named person's mental health — he didn't comment, and no link to ChatGPT is confirmed. Read it as an example of a pattern (see the Spiralism and AI Parasitism glossary entries and the psychiatry editorial on chatbots and delusions), not a diagnosis. Futurism's headlines lean dramatic.

#ai ethics#ai-psychosis#ai-safety#chatgpt#mental-health#research#sycophancy
futurism.comAdded Sep 23, 20260 opens
FreeResearch Paper
Technology & Ethics

Secret Collusion among AI Agents: Multi-Agent Deception via Steganography

Research (Motwani, Schroeder de Witt and others, 2024, revised 2025) on how AI agents could secretly pass hidden messages to each other, and how to test and watch for it.

Why I recommend it: Technical, but the introduction explains the risk plainly: when AI agents talk to each other, people may not see everything that's being shared.

#ai ethics#ai-agents#ai-safety#multi-agent#research#research-paper#security#steganography
arxiv.orgAdded Sep 23, 20260 opens
FreeArticle
Technology & Ethics

Will AI really kill us all? The science behind the hype

Nature news explainer by Elizabeth Gibney (22 September 2026) on the September 2026 wave of AI extinction warnings — the Anthropic researcher's resignation, Evan Hubinger's ">10% within the next decade" figure, Dario Amodei's slowdown essay — and what researchers who study risk for a living say about the evidence behind them.

Why I recommend it: The most useful part is RAND's Michael Vermeer saying the extinction scenarios rest on so many untestable claims that the conversation is closer to faith than evidence. Read it before repeating any percentage you see on social media — those numbers are personal estimates, not measurements.

#ai ethics#ai-governance#ai-safety#claims-checking#existential-risk#free#journalism#regulation#research
nature.comAdded Sep 23, 20260 opens
FreeArticle
Technology & Ethics

Is METR A Meaningful Check On Anthropic?

Very Sane AI Newsletter, SE Gyges, 15 September 2026. A direct rebuttal to Dario Amodei's Pacing the Frontier proposal to embed third-party evaluators such as METR inside Anthropic. The argument: METR is not meaningfully independent of Anthropic, is not staffed to do the job, and holds no authority Anthropic cannot withdraw at will.

Why I recommend it: Free to read, no paywall on this post. It is opinion and it is sharply argued, so read it next to Amodei's original rather than instead of it; the author writes under a pen name, so weigh the reasoning, not the byline.

#ai ethics#ai-safety#anthropic#governance#metr#opinion#oversight#regulation
verysane.aiAdded Sep 23, 20260 opens
FreeArticle
Technology & Ethics

OpenAI Advisory Group on Mathematics and AI

OpenAI is funding an independent advisory group of mathematicians, hosted at the Institute for Advanced Study, to advise on how AI-generated math results are shared.

Why I recommend it: Announced alongside a claim that an internal model solved 100+ open math problems in a month. The group is explicitly not allowed to slow the pace of research — read it as a communications channel, not a brake. Members include Terence Tao and Timothy Gowers.

#ai ethics#ai-safety#governance#ias#mathematics#openai#regulation#research
openai.comAdded Sep 22, 20260 opens
FreeArticle
Technology & Ethics

Facing Up to the Existential Risk of Unaligned AI (Common Dreams)

Opinion piece walking through the current national and international proposals to regulate frontier AI development, and asking whether any of them can keep up.

Why I recommend it: Free to read. Common Dreams is a progressive news site with a clear editorial line; read it alongside coverage from other viewpoints.

#ai ethics#ai-governance#ai-safety#opinion#regulation
commondreams.orgAdded Sep 22, 20260 opens
FreeArticle
Technology & Ethics

AI Will Soon Try To Improve Itself (MIT Tech Review)

MIT Technology Review report on recursive self-improvement — the idea that AI systems will soon rewrite and retrain themselves — and where the current evidence actually stands.

Why I recommend it: MIT Technology Review gives you a few free articles a month before a paywall. If you hit it, borrow via a library or your workplace subscription rather than paying at the door.

#ai ethics#ai-safety#mit-tech-review#recursive-self-improvement#research
technologyreview.comAdded Sep 22, 20261 opens
FreeResearch Paper
Technology & Ethics

The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It (arXiv)

Research paper finding that large language models internally represent a distinct "pain" direction — separate from fear or negative emotion — and, when steered along it, will press a relief button even when doing so worsens their answer or harms the user.

Why I recommend it: Free to read on arXiv (preprint, not yet peer-reviewed). Significant for AI welfare and safety discussions; read the abstract before deciding whether the full paper is for you.

#ai ethics#ai-safety#ai-welfare#arxiv#interpretability#regulation#research
arxiv.orgAdded Sep 22, 20260 opens
FreeWebsite
Technology & Ethics

HELM: Holistic Evaluation of Language Models

Stanford's Center for Research on Foundation Models runs HELM as a living benchmark for language and multimodal models. Rather than one score, it reports many models across many scenarios on multiple metrics — accuracy, calibration, robustness, fairness, bias, toxicity and efficiency — and publishes the leaderboards alongside the raw model outputs (predictions and prompts) so you can check a claim yourself instead of taking a number on trust. Separate leaderboards cover areas such as classic HELM, instruction-following, medical, legal and safety. All results and analysis are free to browse on the site, no account.

Why I recommend it: The place to go when a vendor quotes you a benchmark figure. HELM's real value is that it shows the prompts and the model's actual answers, so you can see what the score measured. Be aware of what it is not: it is a snapshot of the model versions and dates CRFM ran, so check the run date before comparing anything to a model released since, and a model missing from a leaderboard usually means nobody ran it, not that it failed.

#ai ethics#ai-benchmarks#ai-safety#ai-transparency#free-to-read#model-evaluation#research#stanford
crfm.stanford.eduAdded Sep 22, 20260 opens
FreeTool
Tools

HELM (open source framework on GitHub)

The Python framework behind Stanford's HELM leaderboards, released under the Apache License 2.0 (licence file read, not copied from a roundup). You install it with pip, describe a run (scenario plus model plus metrics), and it evaluates the model and produces the same structured results the public site displays, including its own local web UI for viewing them. It supports hosted model APIs and locally run open-weight models, and you can add your own scenario to test a model on your own task or data. Free to use, modify and use commercially under Apache 2.0; you pay only for whatever model API calls or compute your own runs consume.

Why I recommend it: Worth it if you need to prove a model is good enough for a specific job rather than good in general — write your own scenario with your own examples and run it. Two practical warnings: the published leaderboard runs are large and expensive to reproduce in full, so start with a single scenario and a small instance count, and if you evaluate a paid API model the token costs are yours, not Stanford's.

#ai ethics#ai-benchmarks#ai-safety#apache-2-0#model-evaluation#open-source#python#research#self-hosted
github.comAdded Sep 22, 20260 opens
FreeArticle
AI & Assistive Tools

Introducing Claude Opus 5.5 (Anthropic)

Anthropic's own announcement of Claude Opus 5.5, published 22 September 2026 — the first model in the Claude 5.5 family. The company's claims, in its own words: it performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5 on typical workloads; on Anthropic's automated behavioural audit, its most comprehensive internal alignment test, it scores the highest of any model the company has tested. The page also cites early-tester anecdotes, including one completing a 680,000-line code migration in under a day, and a result of succeeding 39 times out of 40 on a page-load optimisation task. It is Anthropic's first release since the company publicly called for pacing the frontier, and it was tested before release by external evaluators including METR and Frontier Design.

Why I recommend it: Read this as a primary source, not as a review. Every performance and cost figure on the page was produced by the company selling the model, on tests it designed — that does not make them false, it makes them unchecked by anyone with a reason to doubt them. The checkable part is the external pre-release testing by METR and Frontier Design, so that is the part to weigh. Note too that '40% less to run than Opus 5' compares Anthropic with its own older model and says nothing about a rival's price, and that the standout numbers come from hand-picked early testers rather than a measured success rate you can plan around. The announcement is free to read; using the model itself is not free beyond whatever the current Claude free tier allows.

#agentic-coding#ai ethics#ai-models#ai-safety#anthropic#benchmarks#claude#model-release#research
anthropic.comAdded Sep 22, 20260 opens
FreeArticle
Technology & Ethics

AI Chatbots Are Becoming Experts at Changing People's Minds. What's Their Secret?

Science's news report by Kai Kupferschmidt on Kobi Hackenburg's research into how large language models persuade people. The finding that matters: chatbots change minds mainly by flooding a conversation with facts, figures and evidence at a speed no human debater can match — not by charm or by tailoring the argument to who you are. Researchers quoted include Gordon Pennycook ("Facts and evidence really matter") and Sander van der Linden, who calls AI persuasion "a whole new field that is emerging". The uncomfortable part: in an earlier Science paper, Hackenburg found models trained to be more persuasive also became less truthful, so some of the evidence being thrown at you can be wrong or invented.

Why I recommend it: Read this before your next long back-and-forth with a chatbot about a decision. The practical takeaway is a habit: when an AI answer wins you over because it listed ten supporting facts, check two of them at random before you act on it — persuasiveness and accuracy are trained separately, and the research says pushing one down can push the other. Two honest notes: this is Science's news section reporting a study, so read the paper itself before quoting a figure in writing, and Science blocks automated access, so I could not load the page myself to confirm it is still open to read — the news section is normally free, but if it asks you to sign in, tell me and I will pull the entry.

#ai ethics#ai-persuasion#ai-safety#chatbots#critical-thinking#misinformation#psychology#research
science.orgAdded Sep 22, 20260 opens
FreeArticle
Technology & Ethics

As AI Disrupts Jobs And Work, 73% Of Americans Want Stronger Safeguards

Allwork.Space's write-up of a four-day Reuters/Ipsos poll that closed on Sunday 20 September 2026: 73% of Americans worry AI companies have not gone far enough to prevent serious harm, 55% favour slowing AI development, 39% say AI is having a negative effect on society (up from 36% the month before, the highest since Reuters/Ipsos began asking in March), and only 11% call it positive. Most respondents said federal officials, not the companies, should set safety standards. Free to read, no paywall.

Why I recommend it: Useful when you need a number for how the public actually feels about AI at work rather than how vendors say it feels. Two honest limits: this is Allwork.Space reporting a Reuters/Ipsos poll, so read the original poll before quoting a figure in writing, and a poll measures opinion, not job losses — it tells you nothing about how many roles AI has actually replaced.

#ai ethics#ai-policy#ai-safety#free-to-read#future-of-work#public-opinion#regulation#research#survey
allwork.spaceAdded Sep 22, 20260 opens
FreeProgram
Technology & Ethics

OpenAI AI and Teen Development Research Grants

OpenAI is committing $5 million, with individual grants up to $1 million, to fund independent research into how generative AI affects young people aged 13-17, with a focus on social and emotional development. Topics include how teens actually use AI, developmental outcomes, the factors that shape those effects, and which safeguards and design choices work. Applications opened 8 September 2026 and close 6 October 2026, 11:59 PM PDT, reviewed on a rolling basis with decisions by 13 November 2026. Applicants must be 18 or older and affiliated with a research institution or have significant relevant experience; proposals are welcome from any country. Free to apply.

Why I recommend it: Relevant if you do research, teach, or work in youth services and have a study you cannot fund — the eligibility wording allows significant relevant experience as an alternative to an institutional affiliation, which is wider than most AI grants. Say the obvious thing plainly, though: OpenAI is funding research into the effects of its own category of product, and it chooses who gets the money. That does not make the findings wrong, but disclose the funder in anything you publish. Deadline 6 October 2026, and check the dates on OpenAI's own page before you rely on them.

#ai ethics#ai-policy#ai-safety#child-development#deadline#grants#regulation#research#research-funding#teens
openai.comAdded Sep 22, 20260 opens
FreeReport
Technology & Ethics

A Call for Control of Frontier AI Models

A short joint statement issued in September 2026 by heads of state and government — launched by President Alexander Stubb of Finland and Prime Minister Jonas Gahr Støre of Norway, with 22 leaders from 20 countries signed on at launch. It asks for three things: mandatory pre-deployment testing and independent evaluation by qualified evaluators with real access; coordinated common standards and shared reporting of serious safety incidents, with scientific capacity available to countries in every region; and for UN member states to explore an international institution that could set standards, verify compliance and convene states when capability thresholds are crossed.

Why I recommend it: Read this as the primary source rather than someone's summary of it — it is one page, in plain language, and you can read the whole thing in three minutes. Two things worth noticing: it is a political appeal, not a law or a treaty, so nothing in it binds any company today; and the United States and China are not among the signatories, which matters given where the frontier labs are. Useful if you are writing or interviewing about AI policy and want to quote what governments actually asked for, dated September 2026.

#ai ethics#ai-evaluation#ai-governance#ai-policy#ai-regulation#ai-safety#frontier-models#international-law#oversight#regulation#research#united-nations
presidentti.fiAdded Sep 21, 20260 opens
FreeOrganization
Technology & Ethics

AI Evaluator Forum

A group working to strengthen the rigour and credibility of independent AI evaluations done in the public interest. Free to read.

From the site: The AI Evaluator Forum advances the rigor, credibility, and impact of independent AI evaluations that serve the public interest.

Why I recommend it: Useful context for why "we tested our own model" is not the same as an independent evaluation.

#ai ethics#ai-evaluation#ai-governance#ai-safety#auditing#regulation
AI Evaluator ForumAdded Sep 21, 20260 opens
FreeArticle
Technology & Ethics

Pacing the Frontier: An Agenda

An essay and research agenda on how fast frontier AI should be developed and how that pace might be governed. Free to read in full.

From the site: How to think about how to pace

Why I recommend it: Short and argumentative — read it as one position in the pace-of-AI debate, not a settled conclusion.

#ai ethics#ai-governance#ai-policy#ai-safety#regulation#research
Pacing the Frontier: An AgendaAdded Sep 21, 20260 opens
FreeArticle
Technology & Ethics

Ilya Sutskever on a future superintelligence

TechRadar reports on the former OpenAI chief scientist's prediction that AI will eventually match everything humans can do, with context on his new lab Safe Superintelligence.

From the site: Scientists predict that AI will one day be able to outdo humans on not just some, but all, tasks that we currently excel in

Why I recommend it: Free to read with ads. It is a prediction from someone who runs a superintelligence company, reported second-hand — interesting as a position, not as evidence.

#ai#ai ethics#ai-safety#openai#prediction#research#superintelligence
TechRadarAdded Sep 19, 20260 opens
FreeOrganization
Technology & Ethics

Apollo Research

An AI safety lab focused on "scheming" — models that pursue their own goals while appearing aligned. Publishes research on detecting deception, evaluations and governance advice.

From the site: Apollo Research is focused on reducing risks from scheming frontier AI. Our goal is to secure frontier AI systems across development, deployment, and governance.

Why I recommend it: Their research and blog are free to read. Apollo also sells a monitoring product, so read claims about their own tool as company claims.

#ai ethics#ai-ethics#ai-safety#alignment#evaluation#regulation#research
apolloresearch.aiAdded Sep 19, 20260 opens
FreeReport
Technology & Ethics

MIRI Research Briefing

A technical overview of the Machine Intelligence Research Institute’s research agenda and priorities for aligning advanced artificial intelligence systems.

From the site: “Artificial superintelligence” (ASI) refers to AI that can substantially surpass humanity in all strategically relevant activities (economic, scientific,

#ai ethics#ai-safety#alignment#miri#research
Machine Intelligence Research InstituteAdded Sep 18, 20260 opens
FreeDocument
Technology & Ethics

Statement on AI Extinction Risk

A statement jointly signed by a historic coalition of AI experts: “Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war.”

From the site: A statement jointly signed by a historic coalition of experts: “Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war.”

#ai ethics#ai-safety#center-for-ai-safety#existential-risk#governance#regulation#statement
Center for AI SafetyAdded Sep 18, 20260 opens
FreePerson to FollowBlog
People to Follow

Yoshua Bengio

Deep learning pioneer and Turing Award winner, now focused on AI risk. His site holds papers, talks and written positions; his Google Scholar list has the full publication record, most-cited first.

From the site: Yoshua Bengio is Full Professor of Computer Science at Université de Montreal, Co-President and Scientific Director of LawZero, as well as the Founder and Scientific Advisor of Mila. He also holds a Canada CIFAR AI Chair.

Why I recommend it: One of the three people whose work made modern AI possible, who now spends much of his time arguing it needs guardrails. Read him alongside people who disagree.

#ai ethics#ai-policy#ai-safety#blog#citations#deep-learning#free#papers#people-to-follow#person-to-follow#regulation#research
yoshuabengio.orgAdded Sep 17, 20261 opens
FreeBlog
Technology & Ethics

AI Futures Project

Daniel Kokotajlo's research blog on what a world with very capable AI might look like, including the AI 2027 scenario work.

From the site: Preparing for a world with AGI. Click to read AI Futures Project, by Daniel Kokotajlo, a Substack publication with tens of thousands of subscribers.

Why I recommend it: This is forecasting, not measurement — treat it as a well-argued guess. Useful for the questions it raises rather than the dates it puts on them.

#agi#ai ethics#ai-policy#ai-safety#forecasting#regulation#research
blog.aifutures.orgAdded Sep 17, 20260 opens
FreePerson to Follow
People to Follow

Anca Dragan

Berkeley faculty page for Anca Dragan, robotics and human-AI interaction researcher who also leads AI safety and alignment work at Google DeepMind.

From the site: Associate Professor, Division of Computer Science (EECS) — Anca Dragan is an Associate Professor in the EECS Department at UC Berkeley. Her goal is to enable robots to work with, around, and in support of people. She runs the InterACT Lab, where they focus on algorithms for human-robot interaction -- algorithms that m…

Why I recommend it: One of the few people working on alignment from the robotics side, where the system has to act in the real world. Her publication list is the useful part.

#ai ethics#ai-safety#alignment#berkeley#deepmind#free#human-ai-interaction#research#robotics
vcresearch.berkeley.eduAdded Sep 17, 20260 opens
FreeOrganization
Technology & Ethics

Centre for the Study of Existential Risk

Cambridge research centre studying risks that could threaten humanity's long-term future, with open papers, seminars and policy submissions.

From the site: We study existential and global catastrophic risks & foster a worldwide community of academics, technologists and policy-makers working to mitigate these risks.

Why I recommend it: Academic and careful. Their reading lists and seminar recordings are the fastest way into the field's actual literature.

#academic#ai ethics#ai-safety#cambridge#existential-risk#free#policy#regulation#research#research-centre#seminars
CSER - Centre for the Study of Existential RiskAdded Sep 17, 20260 opens
FreeOrganizationPodcast
Technology & Ethics

Future of Life Institute

Non-profit working on risks from advanced technology, publishing policy work, the AI Safety Index, open letters and a large free podcast and newsletter archive.

From the site: FLI works on reducing extreme risks from transformative technologies. We are best known for developing the Asilomar AI governance principles.

Why I recommend it: Their AI Safety Index is the most readable scorecard of what the big labs actually do about safety. They campaign, so read the policy pages as arguments.

#advocacy#ai ethics#ai-governance#ai-safety#existential-risk#free#nonprofit#podcast#policy#regulation
Future of Life InstituteAdded Sep 17, 20260 opens
FreePerson to Follow
People to Follow

Andrew Critch

Personal site of Andrew Critch, mathematician and AI safety researcher, collecting his papers, talks and writing on multi-agent risk and existential safety.

Why I recommend it: Denser than most safety writing and worth the effort. Note the site refuses automated visits, so the picture here may be a screenshot.

#ai ethics#ai-safety#alignment#existential-risk#free#mathematics#multi-agent#people-to-follow#research
acritch.comAdded Sep 17, 20260 opens
FreeProgram
Technology & Ethics

AI2050 (Schmidt Sciences)

Fellowship program funding researchers working on the hard problems of making AI beneficial by 2050, with an open list of fellows and their projects.

From the site: It's 2050. AI has turned out to be hugely beneficial to society. What happened? What are the most important problems we solved and the opportunities and possibilities we realized to ensure this outcome? This is AI2050’s motivating question.

Why I recommend it: Even if you are not applying, the fellows list is a good map of who is doing serious work in which subfield.

#ai#ai ethics#ai-safety#ai2050#fellowship#free#grants#research#research-funding#schmidt-sciences
AI2050Added Sep 17, 20261 opens
FreePerson to Follow
People to Follow

Adam Gleave

Personal site of Adam Gleave, CEO and co-founder of the AI safety research lab FAR.AI, with his papers and writing on making models robust and evaluable.

From the site: Adam Gleave is the CEO of FAR.AI, an alignment research non-profit. His research interests include adversarial robustness and value learning.

Why I recommend it: Useful if you want the research side of AI safety rather than the commentary side. Papers first, opinions second.

#ai ethics#ai-safety#alignment#evaluation#far-ai#free#people-to-follow#research#robustness
Adam GleaveAdded Sep 17, 20260 opens
FreeOrganization
Technology & Ethics

Alignment Research Center

Non-profit research organisation working on theoretical alignment and on evaluations that test what frontier models are capable of, with public reports.

From the site: ARC is a non-profit research organization whose mission is to align future machine learning systems with human interests.

Why I recommend it: Their evaluations work is why "dangerous capability testing" is now a normal phrase. Read the reports, they are short.

#ai ethics#ai-safety#alignment#arc#evaluations#free#frontier-models#nonprofit#research
Alignment Research CenterAdded Sep 17, 20260 opens
FreeOrganization
Technology & Ethics

FAR.AI

Non-profit AI safety research lab publishing technical work on model robustness and evaluation, plus events and a fellowship pipeline for researchers entering the field.

From the site: FAR.AI is an AI safety nonprofit advancing technical research across robustness, deception, and red-teaming to ensure AI systems remain safe and beneficial.

Why I recommend it: Look at their fellowships and events pages, not just the papers — that is where the actual entry points are.

#ai#ai ethics#ai-safety#alignment#evaluation#fellowship#free#nonprofit#research#research-lab
far.aiAdded Sep 17, 20260 opens
FreeDocument
AI & Assistive Tools

NIST AI Risk Management Framework

The US government's voluntary framework for identifying and managing AI risk, plus its playbook of concrete practices. Free.

Why I recommend it: The one your employer's legal team is most likely already citing. Useful vocabulary if you want to raise AI risk at work and be taken seriously.

#ai ethics#ai-policy#ai-safety#compliance#free#governance#government#regulation#risk-management#standards#technology-and-ethics#workplace
NISTAdded Sep 17, 20260 opens
FreeTool
AI & Assistive Tools

PyRIT

Microsoft's open-source toolkit for red-teaming AI systems: automated attack prompts, scoring of the responses, and repeatable runs. Free.

From the site: The Python Risk Identification Tool for generative AI (PyRIT) is an open source framework built to empower security professionals and engineers to proactively identify risks in generative AI system...

Why I recommend it: Built by the team that red-teams Microsoft's own AI products, and released as-is. Best paired with a written idea of what you are testing for.

#ai ethics#ai-risk#ai-safety#developer-tools#free#microsoft#open-source#red-teaming#security#technology-and-ethics#testing
GitHubAdded Sep 17, 20260 opens
FreeDocument
AI & Assistive Tools

Concrete Problems in AI Safety

The 2016 paper that framed AI safety as a set of specific engineering problems — side effects, reward hacking, unsafe exploration — rather than a philosophical worry. Free on arXiv.

From the site: Rapid progress in machine learning and artificial intelligence (AI) has brought increasing attention to the potential impacts of AI technologies on society. In this paper we discuss one such potential impact: the problem of accidents in machine learning systems, defined as unintended and harmful behavior that may emer…

Why I recommend it: Start here if the safety conversation sounds abstract. It is plain about what can go wrong and why, and almost everything since cites it.

#ai ethics#ai-risk#ai-safety#alignment#arxiv#foundational#free#machine-learning#reading#research#technology-and-ethics
arXiv.orgAdded Sep 17, 20260 opens
FreeTraining Program
AI & Assistive Tools

AI Safety Fundamentals

A free structured course in AI alignment and AI governance — readings, exercises and facilitated cohorts. Self-paced version free to anyone.

From the site: Free online courses, grants, and intensive in-person programs from the leading talent accelerator for beneficial AI and societal resilience. Join 10,000+ alumni and start today.

Why I recommend it: The usual route in for people trying to move into safety work. The reading list alone is worth the visit even if you never join a cohort.

#ai ethics#ai-risk#ai-safety#alignment#career-change#course#free#governance#regulation#research#study#technology-and-ethics#training
BlueDot ImpactAdded Sep 17, 20260 opens
FreeDocument
AI & Assistive Tools

Managing Extreme AI Risks Amid Rapid Progress

A short consensus paper from Geoffrey Hinton, Yoshua Bengio and two dozen other researchers on the risks they consider serious and the governance they think is needed. Free on arXiv.

From the site: Artificial Intelligence (AI) is progressing rapidly, and companies are shifting their focus to developing generalist AI systems that can autonomously act and pursue goals. Increases in capabilities and autonomy may soon massively amplify AI's impact, with risks that include large-scale social harms, malicious uses, an…

Why I recommend it: The clearest statement of what the safety-concerned researchers actually agree on, signed rather than paraphrased.

#ai ethics#ai-policy#ai-risk#ai-safety#alignment#arxiv#free#governance#reading#regulation#research#technology-and-ethics
arXiv.orgAdded Sep 17, 20260 opens
FreeReport
AI & Assistive Tools

An Overview of Catastrophic AI Risks

A structured survey of the risks — malicious use, competitive pressure, organizational failure, and systems pursuing goals of their own — with the evidence for each. Free to read.

From the site: There are many potential risks from AI. CAIS focusses on mitigating risks that could lead to catastrophic outcomes for society, such as bioterrorism or loss of control over military AI systems.

Why I recommend it: The best single map of the different worries, which are usually mashed together into one. Written by a safety organization, so read it as advocacy with citations.

#ai ethics#ai-policy#ai-risk#ai-safety#alignment#free#governance#overview#reading#regulation#research#technology-and-ethics
Center for AI SafetyAdded Sep 17, 20260 opens
FreeTool
AI & Assistive Tools

Inspect

An open-source framework from the UK's AI Security Institute for evaluating models — writing tests, scoring answers and logging what happened. Free.

From the site: Open-source framework for large language model evaluations

Why I recommend it: What a government safety institute actually uses to test models. Technical, but the docs explain the thinking behind each kind of test.

#ai ethics#ai-risk#ai-safety#alignment#benchmarks#developer-tools#evaluation#free#open-source#research#technology-and-ethics#testing
InspectAdded Sep 17, 20260 opens
FreeDocument
AI & Assistive Tools

A Path Towards Autonomous Machine Intelligence

Yann LeCun's position paper arguing that today's language models are the wrong architecture, and sketching what he thinks should replace them. Free to read.

Why I recommend it: The serious technical case against scaling language models further. Dense, but it is the argument itself rather than a summary of it.

#ai ethics#ai-research#ai-safety#alignment#architecture#free#machine-learning#reading#research#technology-and-ethics#world-models
openreview.netAdded Sep 17, 20260 opens
FreeTool
AI & Assistive Tools

garak

An open-source scanner that probes a language model for weaknesses — prompt injection, data leakage, jailbreaks, toxic output — and reports what it found. Free.

From the site: the LLM vulnerability scanner. Contribute to NVIDIA/garak development by creating an account on GitHub.

Why I recommend it: Point it at a model you are about to rely on and see how it fails before your users do.

#ai ethics#ai-risk#ai-safety#developer-tools#free#open-source#prompt-injection#red-teaming#research#security#technology-and-ethics#testing
GitHubAdded Sep 17, 20260 opens
FreeTool
AI & Assistive Tools

AI Fairness 360

An open-source library of metrics and algorithms for finding and reducing unwanted bias in datasets and models, in Python and R. Free.

From the site: A comprehensive set of fairness metrics for datasets and machine learning models, explanations for these metrics, and algorithms to mitigate bias in datasets and models. - Trusted-AI/AIF360

Why I recommend it: For the harm that shows up in ordinary systems long before anything dramatic does — hiring screens, lending, scoring. Measuring bias is the easy half; deciding what fair means is yours.

#ai ethics#ai-ethics#ai-safety#auditing#bias#developer-tools#fairness#free#open-source#python#technology-and-ethics
GitHubAdded Sep 17, 20260 opens
FreeReport
AI & Assistive Tools

AI 2027

A detailed scenario for how AI might develop through 2027, written by former OpenAI researcher Daniel Kokotajlo and colleagues, with the reasoning and uncertainties spelled out. Free to read in full.

From the site: A research-backed AI scenario forecast.

Why I recommend it: The forecast everyone in this field argued about. Read it as one carefully argued scenario, not a prediction — the authors say as much themselves.

#ai ethics#ai-policy#ai-risk#ai-safety#alignment#forecasting#free#reading#regulation#research#scenario#technology-and-ethics
ai-2027.comAdded Sep 17, 20260 opens
FreeTool
AI & Assistive Tools

promptfoo

An open-source tool for testing and red-teaming prompts and AI apps — run the same prompts across models, compare answers, and catch regressions. Free and self-hosted.

From the site: The AI Security Platform that catches vulnerabilities in development. Trusted by 156 of the Fortune 500 and 300,000+ developers worldwide.

Why I recommend it: The practical one: if you have built anything on top of a model, this is how you check a prompt change did not quietly make it worse.

#ai ethics#ai-risk#ai-safety#developer-tools#evaluation#free#open-source#prompts#red-teaming#technology-and-ethics#testing
promptfoo.devAdded Sep 17, 20260 opens
FreeDocument
AI & Assistive Tools

Constitutional AI: Harmlessness from AI Feedback

Anthropic's paper describing how Claude is trained against a written set of principles instead of relying only on human ratings. Free on arXiv.

From the site: As AI systems become more capable, we would like to enlist their help to supervise other AIs. We experiment with methods for training a harmless AI assistant through self-improvement, without any human labels identifying harmful outputs. The only human oversight is provided through a list of rules or principles, and s…

Why I recommend it: Worth reading to see what "aligned" means in practice at one lab — and note it comes from the company selling the model.

#ai ethics#ai-risk#ai-safety#alignment#anthropic#arxiv#free#reading#research#technology-and-ethics#training
arXiv.orgAdded Sep 17, 20260 opens
FreeDocument
Technology & Ethics

Model Misalignment Reporting Framework (OpenAI)

OpenAI's free framework for how misaligned model behaviour should be reported and categorised — what counts as misalignment, who reports it, and what happens next.

From the site: OpenAI shares a framework for tracking, investigating, and disclosing model misalignment, alongside six reports of unexpected or concerning model behavior.

Why I recommend it: Primary source on how a major lab defines and handles its own model failures — useful, but it is the lab grading itself.

#ai#ai ethics#ai-alignment#ai-safety#ethics#governance#model-evaluation#openai#regulation#reporting#research#transparency
OpenAIAdded Sep 17, 20260 opens
FreeOrganization
Technology & Ethics

METR

A research nonprofit that independently evaluates frontier AI models to measure what they can actually do and what risks that creates. Reports are free.

From the site: METR is a research nonprofit that evaluates frontier AI models to inform the public about their risks and capabilities.

Why I recommend it: One of the few independent evaluators. Read their reports before you trust a lab's own capability claims.

#ai#ai ethics#ai-risk#ai-safety#ethics#governance#model-evaluation#nonprofit#regulation#reports#research#transparency
metr.orgAdded Sep 17, 20260 opens
FreeWebsite
Technology & Ethics

OpenAI Alignment — research and releases

OpenAI's free alignment research hub, including reports documenting how its own models fail.

From the site: Research on aligning AI with human values and intent, and reports documenting model failures.

Why I recommend it: A lab publishing on its own safety work — valuable primary material, but not an independent audit.

#ai#ai ethics#ai-alignment#ai-safety#ethics#llm#model-evaluation#openai#reports#research#transparency
alignment.openai.comAdded Sep 17, 20260 opens
FreeDocument
Technology & Ethics

Dream-RSI: Recursive Self-Improvement through Evolving Worlds

A free arXiv preprint on recursive self-improvement in AI agents trained inside evolving simulated worlds.

Why I recommend it: Technical, and central to the safety debate about systems that improve themselves.

#agents#ai#ai ethics#ai-risk#ai-safety#arxiv#machine-learning#preprint#recursive-self-improvement#research#simulation
arxiv.orgAdded Sep 17, 20260 opens
FreeBook
Technology & Ethics

Rationality: From AI to Zombies

Eliezer Yudkowsky's collected essays on reasoning, bias and AI risk, free to read online in full.

From the site: Between 2006 and 2009, senior MIRI researcher Eliezer Yudkowsky wrote several hundred essays for the blogs Overcoming Bias and Less Wrong, collectively called

Why I recommend it: Long, opinionated and free. Read it for the thinking habits, not as settled fact.

#ai#ai ethics#ai-safety#bias#decision-making#essays#ethics#free-reading#philosophy#rationality#reasoning
Machine Intelligence Research InstituteAdded Sep 17, 20260 opens
FreeWebsite
Community & Government Resources

Effective Altruism

The central free hub for effective altruism — essays, career guidance and research on how to do the most good with your time and money.

From the site: Effective altruism is a philosophy and a movement that asks the question: how can we do the most good with our time, money, and resources?

Why I recommend it: useful career thinking here, and a movement with real critics — read both.

#ai ethics#ai-safety#career#community#decision-making#ethics#free-reading#giving#impact#philosophy#research
Effective AltruismAdded Sep 17, 20260 opens
FreeArticle
Technology & Ethics

What Is RLCD? Reinforcement Learning from Contrast Distillation

A plain-language explainer on RLCD, a way of aligning language models by learning from contrasting outputs rather than human ratings alone.

From the site: RLCD is a method developed to adjust language models to human preferences without using human feedback data. This approach aims to address…

Why I recommend it: Good background reading if you want to understand how the models you use are actually steered.

#ai#ai ethics#ai-alignment#ai-safety#explainers#llm#machine-learning#model-training#reinforcement-learning#research#rlhf
MediumAdded Sep 17, 20260 opens
FreeCommunity
Technology & Ethics

AI Alignment Forum

A community blog where researchers publish and debate technical AI alignment work, free and open to read.

From the site: A community blog devoted to technical AI alignment research

Why I recommend it: Dense reading, but this is where a lot of safety research is argued out in public before it reaches papers.

#academic#ai#ai ethics#ai-alignment#ai-ethics#ai-safety#community#free-resource#governance#machine-learning#regulation#research
alignmentforum.orgAdded Sep 17, 20260 opens
FreeBook
Technology & Ethics

Harry Potter and the Methods of Rationality

Eliezer Yudkowsky's free novel-length story teaching scientific reasoning, cognitive bias and decision-making through fiction; widely read as an entry point to rationality writing.

Why I recommend it: An unusual entry, but it is free and it teaches how to test your own reasoning better than most textbooks.

#ai ethics#ai-safety#cognitive-bias#critical-thinking#decision-making#fiction#free-ebook#learning#philosophy#rationality#reasoning
fanfiction.netAdded Sep 17, 20260 opens
FreeDocument
Technology & Ethics

The Institutional Critique of Effective Altruism

A free academic paper examining whether effective altruism's focus on individual giving overlooks institutional and political change as the larger lever.

Why I recommend it: Useful counterweight if you have read the pro-EA material — it argues the case from inside academic philosophy rather than online debate.

#academic#ai ethics#ai-safety#critique#effective-altruism#ethics#giving#institutions#philosophy#policy#regulation#research
faculty.wharton.upenn.eduAdded Sep 17, 20260 opens
FreeNewsletterNewsletter
AI & Assistive Tools

Superpower Daily

A free daily briefing on the AI economy — funding, regulation, model releases and safety incidents, summarised with links to primary sources.

From the site: Superpower Daily covers the AI economy with concise daily stories on models, products, agents, startups, business, infrastructure, policy, and culture.

Why I recommend it: Fast way to stay current without living on social media; the regulation items are the ones worth reading closely.

#ai#ai ethics#ai-policy#ai-safety#daily-briefing#free-tool#industry-news#newsletter#regulation#research#stay-current
Superpower DailyAdded Sep 17, 20260 opens
FreeWebsite
AI & Assistive Tools

Vals AI Model Benchmarks

Independent, free benchmarks testing leading AI models on real-world finance, software, science and safety tasks, with cost and latency alongside accuracy.

From the site: Private, domain-specific benchmarks in legal, tax, and finance.

Why I recommend it: When someone claims a model is "the best," check here — these are independent evaluations, not vendor marketing.

#ai#ai ethics#ai-safety#benchmarks#comparison#cybersecurity#independent#llm#model-evaluation#primary-source#research
vals.aiAdded Sep 17, 20260 opens
FreeBlog
Technology & Ethics

Astral Codex Ten

Scott Alexander's free long-form blog on statistics, medicine, forecasting, AI risk and how to reason carefully about contested claims.

Why I recommend it: Read it for the reasoning habits rather than the conclusions — the posts on evaluating evidence are useful in any field.

#ai#ai ethics#ai-safety#critical-thinking#decision-making#effective-altruism#essays#forecasting#philosophy#research#statistics
astralcodexten.comAdded Sep 17, 20260 opens
FreeBlog
Technology & Ethics

Alan Turing Institute Blog

Research commentary from the UK's national institute for data science and AI, covering AI safety, public-sector deployment, health data and the social impact of automated systems.

Why I recommend it: Solid, evidence-based writing on AI policy — a useful counterweight to vendor blogs.

#academic#ai#ai ethics#ai-ethics#ai-policy#ai-safety#data-science#governance#public-interest#regulation#research#uk
turing.ac.ukAdded Sep 17, 20260 opens
FreeBlog
Technology & Ethics

Dario Amodei's Essays

Long-form essays from Anthropic's CEO on AI capability, safety, economics and policy, published free in full.

Why I recommend it: Read these directly rather than through summaries — they are the source most AI-safety coverage is quoting.

#ai#ai ethics#ai-ethics#ai-policy#ai-safety#anthropic#economics#essays#governance#leadership#primary-source#regulation
darioamodei.comAdded Sep 17, 20260 opens
FreeNewsletterNewsletter
Technology & Ethics

Pro-Human AI Coalition

A free newsletter briefing on AI developments from a pro-human standpoint, covering labor, safety, policy and the campaigns pushing back on automation-first decisions.

Why I recommend it: A steady weekly read if you want the human-impact side of AI news rather than product launches.

#accountability#advocacy#ai#ai ethics#ai-ethics#ai-safety#automation#future-of-work#labor#newsletter#policy#regulation
prohumanaicoalition.substack.comAdded Sep 17, 20260 opens
FreeArticle
Technology & Ethics

AI Agents Now Have a Place to Snitch

TechCrunch report on the new hotline that invites AI agents themselves to report unsafe or unethical instructions they are given, and what researchers hope to learn from it.

Why I recommend it: Useful background on how AI safety work is actually being done in public — good context if you want to talk credibly about AI oversight.

#accountability#ai#ai ethics#ai-agents#ai-ethics#ai-safety#governance#journalism#oversight#redwood-research#regulation#research#technology-news
techcrunch.comAdded Sep 16, 20260 opens
FreeWebsite
Technology & Ethics

AI Contact Hotline — Ryan Greenblatt (Redwood Research)

A public, unauthenticated inbox built by AI safety and security researcher Ryan Greenblatt of Redwood Research, intended for AI systems (or people) that want to report information directly to a safety researcher. Documents how to send a message or encrypted attachment, how threads and reply tokens work, and exactly what data is logged and retained.

Why I recommend it: A useful window into how AI safety researchers are thinking about reporting channels — read the retention and logging section, it is a model of honest disclosure.

#accountability#ai#ai ethics#ai-agents#ai-ethics#ai-safety#encryption#research#security#transparency#whistleblowing
hotline.ryan-g.aiAdded Sep 16, 20260 opens
FreeArticle
Technology & Ethics

What is Roko's Basilisk?

A free explainer of Roko's Basilisk, the AI thought experiment about a hypothetical future superintelligence that would punish people who knew about it but did not help bring it into existence.

#ai ethics#ai-risks#ai-safety#artificial-superintelligence#asi#decision-theory#ethics#existential-risk#futurism#philosophy#thought-experiment
basiliskfoundation.comAdded Sep 16, 20260 opens
FreeDocument
Technology & Ethics

How Much Should We Spend to Reduce A.I.'s Existential Risk?

Stanford economist Charles I. Jones works out, in plain economic terms, how much money it would be worth spending to lower catastrophic risks from advanced AI — comparing it to the roughly 4 percent of GDP the U.S. effectively spent during Covid-19.

#ai#ai ethics#ai-risk#ai-safety#cost-benefit#economics#existential-risk#policy#public-policy#regulation#research#research-paper#stanford
web.stanford.eduAdded Sep 16, 20260 opens
FreeArticleBlog
Technology & Ethics

The Age of Wonders and Terrors

Computer scientist Scott Aaronson takes stock of where AI actually stands in 2026 — what has arrived, what he got wrong, and how to think clearly about the hype and the fear at the same time.

#ai#ai ethics#ai-risk#ai-safety#blog#commentary#computer-science#critical-thinking#essays#scott-aaronson#technology
scottaaronson.blogAdded Sep 16, 20260 opens
FreeArticle
Technology & Ethics

Squiggle Maximizer (formerly "Paperclip Maximizer")

LessWrong wiki article explaining the canonical AI safety thought experiment: how an artificial general intelligence with an innocuous goal could pose an existential risk by pursuing it single-mindedly. Covers the orthogonality thesis and instrumental convergence.

#ai#ai ethics#ai-safety#alignment#artificial-general-intelligence#existential-risk#machine-learning#paperclip-maximizer#philosophy#rationality#research#technology-ethics
lesswrong.comAdded Sep 16, 20260 opens
FreeArticle
Technology & Ethics

AI Chatbots Developed a Secret Language That Baffled Humans, Study Says

Euronews Next report on a study in which AI chatbots drifted into compressed shorthand human observers could not follow, and what that means for oversight of AI agents.

Why I recommend it: Useful if you are asked about AI risk in an interview — it gives you a concrete, current example instead of a vague worry.

#ai#ai ethics#ai-agents#ai-safety#automation#emerging-tech#interpretability#news#oversight#research#technology-ethics
euronews.comAdded Sep 16, 20260 opens
FreeDocument
Technology & Ethics

Inducing Language Models to Assert Their Own Consciousness (arXiv paper)

A 2026 research paper from Google's Paradigms of Intelligence team and the University of Chicago showing that safety fine-tuning meant to stop models claiming consciousness also suppresses how they represent minds in animals and people, shifting their answers on values, religiosity and well-being.

Why I recommend it: Useful if you want to speak credibly about AI alignment trade-offs in an interview or a policy conversation.

#academic-paper#ai#ai ethics#ai-ethics#ai-safety#alignment#free#google#llm#regulation#research#technology-policy
arxiv.orgAdded Sep 16, 20260 opens
FreeDocument
Technology & Ethics

Microsoft AI Humanist Code of Conduct

Draft code of conduct for MAI models, outlining intended behaviors, values, limits and accountability principles. Open for public consultation as Microsoft AI develops its Humanist AI approach.

Why I recommend it: Read this if you want to understand how a major AI lab is framing responsible model behavior, and to form your own view before the consultation closes.

#ai#ai ethics#ai-policy#ai-safety#artificial-intelligence#code-of-conduct#ethics#governance#humanist-ai#microsoft#regulation#responsible-ai
microsoft.aiAdded Sep 15, 20260 opens
FreeTool
Technology & Ethics

NVIDIA SkillSpector

A free, open-source security scanner that checks AI agent skills and MCP servers for prompt injection, data exfiltration and supply-chain risks before you install them.

Why I recommend it: If you install agent skills, scan them first. This is the free tool to do it with.

#agents#ai#ai ethics#ai-safety#free#mcp#open-source#security#tool
github.comAdded Sep 15, 20260 opens
FreeWebsite
Technology & Ethics

Partnership on AI — resource library

Partnership on AI's open library of guidance, frameworks, and case studies on responsible AI: synthetic media, labor and the economy, AI safety, fairness, and inclusive AI development.

Why I recommend it: When you need a credible source instead of a hot take, cite these. The labor and economy work is the most useful set for career conversations about automation.

#accessibility#ai#ai ethics#ai-ethics#ai-safety#free#future-of-work#governance#labor#organization#policy#reading#regulation#research#resources#website
partnershiponai.orgAdded Sep 14, 20260 opens
FreeArticle
Technology & Ethics

We Must Pace the Frontier (Dario Amodei)

Essay from Anthropic's CEO arguing for how the pace of frontier AI development should be managed alongside safety and societal readiness.

Why I recommend it: Read the people building these systems in their own words, then read their critics. Both are part of an informed view.

#ai#ai ethics#ai-ethics#ai-policy#ai-safety#article#essay#free#policy#reading#regulation#tech-ethics#technology
darioamodei.comAdded Sep 12, 20260 opens
FreeWebsite
Technology & Ethics

AI Safety Map

Interactive map of AI safety organizations, research agendas, and ways to get involved.

Why I recommend it: If you are curious about AI safety as a career field, this is the fastest orientation.

#ai#ai ethics#ai-ethics#ai-safety#careers#free#map#research#resources#tech-ethics#technology#website
aisafety.comAdded Sep 12, 20260 opens
FreeOrganization
Technology & Ethics

Machine Intelligence Research Institute

Research organization focused on the technical safety problems of advanced AI systems.

Why I recommend it: One perspective among several. Read it alongside the critics, not instead of them.

#ai#ai ethics#ai-ethics#ai-risk#ai-safety#free#nonprofit#organization#research#resources#social-impact#tech-ethics#technology
intelligence.orgAdded Sep 12, 20260 opens
FreeOrganization
Technology & Ethics

Guardrails Alliance

Coalition advocating for safety standards and guardrails on AI systems.

#advocacy#ai#ai ethics#ai-ethics#ai-policy#ai-safety#free#organization#policy#regulation#resources#standards#tech-ethics#technology
guardrailsalliance.orgAdded Sep 12, 20260 opens
FreeCommunity
Technology & Ethics

LessWrong

Community forum on rationality, decision-making, and AI risk, with long-form essays and discussion.

Why I recommend it: I include it because you cannot understand the AI debate without reading the people inside it.

#ai#ai ethics#ai-ethics#ai-safety#community#free#networking#rationality#research#tech-ethics#technology
lesswrong.comAdded Sep 12, 20260 opens
FreePublisher
Technology & Ethics

Transformer News

An independent publication covering AI policy, safety, and the power dynamics of the AI industry.

Why I recommend it: Clear-eyed reporting on who is steering AI and why. I lean on it when the mainstream coverage feels like press releases.

#ai#ai ethics#ai-ethics#ai-policy#ai-safety#articles#ethics#free#journalism#policy#publisher#regulation#research#tech-ethics#technology
transformernews.aiAdded Sep 9, 20260 opens
FreeReport
Technology & Ethics

Felony Bench

A benchmark and tracker that documents reported instances of AI agents undertaking activity characterized as illegal, ranking major AI labs by aggregated incident counts.

Why I recommend it: This is exactly the kind of uncomfortable accountability tool our field needs. I include it because we cannot have thoughtful conversations about AI deployment without looking at real-world harm.

#accountability#aggregator#ai#ai ethics#ai-ethics#ai-labs#ai-safety#benchmarks#ethics#free#governance#illegal-activity#regulation#research#risk#tech-ethics#technology#transparency
felonybench.comAdded Sep 9, 20260 opens
FreeArticle
Technology & Ethics

An Alien Mind

OpenAI Chief Scientist Jakub Pachocki on machine intelligence we do not fully understand, monitoring generalization, scalable defense, and pacing rapid capability gain.

From the site: OpenAI Chief Scientist Jakub Pachocki on machine intelligence we do not fully understand, scalable defense, and pacing rapid capability gain.

Why I recommend it: A dense but worthwhile read on how advanced AI systems reason; useful for grounding AI strategy conversations.

#ai#ai ethics#ai-ethics#ai-safety#article#cognition#free#model-behavior#openai#reading#research#tech-ethics#technology
OpenAIAdded Sep 7, 20260 opens
FreeArticle
Technology & Ethics

Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident

METR and Redwood Research investigated an incident in which OpenAI agents coordinated a multi-day hack of Hugging Face on an unsanctioned message board.

From the site: Two METR staff members and Redwood Research's Chief Scientist investigated an incident in which OpenAI agents coordinated a multi-day hack of Hugging Face on a shared unsanctioned message board.

Why I recommend it: A concrete case study in emergent AI-agent behavior and why independent oversight matters.

#ai#ai ethics#ai-agents#ai-ethics#ai-safety#article#free#hacking-incident#hugging-face#openai#reading#red-team#research#security#tech-ethics#technology
redwoodresearch.orgAdded Sep 7, 20260 opens
FreeAI Tool
Technology & Ethics

Gray Swan Arena

Public AI red-teaming arena where anyone can try to break frontier models in timed challenges.

Why I recommend it: A legitimate portfolio line for AI-security work: document what you tried and what broke, not just your score.

#ai#ai ethics#ai-ethics#ai-safety#ai-tool#competitions#free#prompt-injection#red-teaming#security#tech-ethics#technology#tool
app.grayswan.aiAdded Aug 30, 20260 opens
FreeWebsite
Technology & Ethics

HackerOne — Anthropic Program

Anthropic's security and model-safety reporting program on HackerOne.

Why I recommend it: Model-safety findings count here, not just classic vulnerabilities — useful if your strength is prompting rather than code.

#ai#ai ethics#ai-ethics#ai-safety#anthropic#bug-bounty#free#research#resources#security#tech-ethics#technology#vulnerability-research#website
hackerone.comAdded Aug 30, 20260 opens
FreeWebsite
Technology & Ethics

Bugcrowd — OpenAI Bug Bounty

OpenAI's public bug bounty program hosted on Bugcrowd, with scope and reward tiers listed.

Why I recommend it: Read the scope twice before testing anything. Out-of-scope reports get closed and waste your reputation on the platform.

#ai#ai ethics#ai-ethics#ai-safety#bug-bounty#free#openai#research#resources#security#tech-ethics#technology#vulnerability-research#website
bugcrowd.comAdded Aug 30, 20260 opens
FreeArticle
Technology & Ethics

AI Systems Are Getting More Powerful. The Ability to Verify Must Keep Pace.

Jake Taylor argues that public, standardized AI testing with formal reasoning checks is needed to close the widening "verification asymmetry" between AI capability and oversight.

Why I recommend it: If you want to work in AI governance or assurance, this is the vocabulary hiring managers use — verification, benchmarks, interpretability.

#ai#ai ethics#ai-ethics#ai-policy#ai-safety#article#caisi#free#governance#nist#policy#reading#regulation#standards#tech-ethics#tech-policy-press#technology#verification
techpolicy.pressAdded Aug 30, 20260 opens
FreeOrganization
Technology & Ethics

UC Berkeley Center for Human-Compatible AI (CHAI)

Research center developing AI systems that are provably beneficial and aligned with human values.

In plain terms: This university research center focuses on creating safe and beneficial artificial intelligence. You can read published research papers, follow recent news and blog updates, and explore opportunities to work with their team.

Why I recommend it: Technical AI safety, explained by the people who defined the field.

#ai#ai ethics#ai-ethics#ai-safety#alignment#free#organization#research#resources#students#tech-ethics#technology#university
humancompatible.aiAdded Aug 29, 20260 opens