We0 article (July 2026) on philosopher Harvey Lederman's research on the Confucian thinker Wang Yangming and its connection to Anthropic's alignment work on Claude. Published on the blog of an AI website-builder company.
OpenAI's free framework for how misaligned model behaviour should be reported and categorised — what counts as misalignment, who reports it, and what happens next.
From the site: OpenAI shares a framework for tracking, investigating, and disclosing model misalignment, alongside six reports of unexpected or concerning model behavior.
Why I recommend it: Primary source on how a major lab defines and handles its own model failures — useful, but it is the lab grading itself.
A plain-language explainer on RLCD, a way of aligning language models by learning from contrasting outputs rather than human ratings alone.
From the site: RLCD is a method developed to adjust language models to human preferences without using human feedback data. This approach aims to address…
Why I recommend it: Good background reading if you want to understand how the models you use are actually steered.