Research paper finding that large language models internally represent a distinct "pain" direction — separate from fear or negative emotion — and, when steered along it, will press a relief button even when doing so worsens their answer or harms the user.
Why I recommend it: Free to read on arXiv (preprint, not yet peer-reviewed). Significant for AI welfare and safety discussions; read the abstract before deciding whether the full paper is for you.
The 2016 paper that framed AI safety as a set of specific engineering problems — side effects, reward hacking, unsafe exploration — rather than a philosophical worry. Free on arXiv.
From the site: Rapid progress in machine learning and artificial intelligence (AI) has brought increasing attention to the potential impacts of AI technologies on society. In this paper we discuss one such potential impact: the problem of accidents in machine learning systems, defined as unintended and harmful behavior that may emer…
Why I recommend it: Start here if the safety conversation sounds abstract. It is plain about what can go wrong and why, and almost everything since cites it.
A short consensus paper from Geoffrey Hinton, Yoshua Bengio and two dozen other researchers on the risks they consider serious and the governance they think is needed. Free on arXiv.
From the site: Artificial Intelligence (AI) is progressing rapidly, and companies are shifting their focus to developing generalist AI systems that can autonomously act and pursue goals. Increases in capabilities and autonomy may soon massively amplify AI's impact, with risks that include large-scale social harms, malicious uses, an…
Why I recommend it: The clearest statement of what the safety-concerned researchers actually agree on, signed rather than paraphrased.
Anthropic's paper describing how Claude is trained against a written set of principles instead of relying only on human ratings. Free on arXiv.
From the site: As AI systems become more capable, we would like to enlist their help to supervise other AIs. We experiment with methods for training a harmless AI assistant through self-improvement, without any human labels identifying harmful outputs. The only human oversight is provided through a list of rules or principles, and s…
Why I recommend it: Worth reading to see what "aligned" means in practice at one lab — and note it comes from the company selling the model.
A free explorer for new arXiv research with plain-language paper summaries, topic pages and video overviews, so you can follow AI research without reading raw papers.
From the site: Your first stop to discover and learn about new arXiv research. Detailed paper summaries, video overviews, and more — no prompting required.
Why I recommend it: The fastest way I know to keep up with AI research when you are not a researcher. Free to browse.
A research paper describing a software library whose repository holds almost no code: plain-language design documents are the durable artifact, and AI coding agents regenerate the implementation from those docs on every update.
Why I recommend it: The takeaway for non-engineers is bigger than the paper: clear written thinking is becoming the valuable skill, and the code is what gets generated from it.
The underlying working paper by Jeremy Yang and co-authors, using Perplexity data to model tasks as discrete steps and compare fixed vs. marginal costs of chatbots versus autonomous agents.
Why I recommend it: If the HBS summary hooks you, go to the source. Skim the task-cost framework and use it to audit your own week: which tasks are high-step and repeatable? Those are the ones to hand to an agent first.