2893 hand-picked resources, updated every week. Search it, filter it, or just browse a collection and see what catches your eye. Want today’s headlines instead? Read the free AI news feed.
Type
Platform
Topics
Cost
Filtered by tag
Ask the library about workforce, startup, and technology trends
Answers come only from resources in this hub, with the sources listed underneath.
LessWrong post that separates AI alignment research into two categories, engineering work to make systems behave and scientific work to understand misalignment, in the context of debate over whether some alignment research is net negative.
October 2026 post by computer scientist Scott Aaronson describing CS395T AI Alignment Theory, a new graduate course he teaches at UT Austin, with the course description and topics.
Essay proposing three distinct facets of AI alignment: control (the system cannot cause unacceptably bad outcomes even if trying), intent alignment (the system tries to do what the user wants), and values alignment (the system refuses widely unacceptable actions). Uses a stock-trading agent scenario to show how the facets can conflict.
We0 article (July 2026) on philosopher Harvey Lederman's research on the Confucian thinker Wang Yangming and its connection to Anthropic's alignment work on Claude. Published on the blog of an AI website-builder company.
OpenAI's free framework for how misaligned model behaviour should be reported and categorised — what counts as misalignment, who reports it, and what happens next.
From the site: OpenAI shares a framework for tracking, investigating, and disclosing model misalignment, alongside six reports of unexpected or concerning model behavior.
Why I recommend it: Primary source on how a major lab defines and handles its own model failures — useful, but it is the lab grading itself.
A plain-language explainer on RLCD, a way of aligning language models by learning from contrasting outputs rather than human ratings alone.
From the site: RLCD is a method developed to adjust language models to human preferences without using human feedback data. This approach aims to address…
Why I recommend it: Good background reading if you want to understand how the models you use are actually steered.