Skip to content
Launchpad Library logo

ai alignment

6 free resources on this topic. Everything here is free and hand-picked. You can also search within this topic.

FreeWebsite
Technology & Ethics

Softmax

Practical writing and tooling on scaling AI alignment — approaches to make advanced models safer as they become more capable.

From the site: A universe of multiplayer games where humans and their coding agents compete, cooperate, and interact.

#AI alignment#ai ethics#safety#scaling
SoftmaxAdded Sep 18, 20260 opens
FreeDocument
Technology & Ethics

Model Misalignment Reporting Framework (OpenAI)

OpenAI's free framework for how misaligned model behaviour should be reported and categorised — what counts as misalignment, who reports it, and what happens next.

From the site: OpenAI shares a framework for tracking, investigating, and disclosing model misalignment, alongside six reports of unexpected or concerning model behavior.

Why I recommend it: Primary source on how a major lab defines and handles its own model failures — useful, but it is the lab grading itself.

#ai#ai ethics#ai-alignment#ai-safety#ethics#governance#model-evaluation#openai#regulation#reporting#research#transparency
OpenAIAdded Sep 17, 20260 opens
FreeWebsite
Technology & Ethics

OpenAI Alignment — research and releases

OpenAI's free alignment research hub, including reports documenting how its own models fail.

From the site: Research on aligning AI with human values and intent, and reports documenting model failures.

Why I recommend it: A lab publishing on its own safety work — valuable primary material, but not an independent audit.

#ai#ai ethics#ai-alignment#ai-safety#ethics#llm#model-evaluation#openai#reports#research#transparency
alignment.openai.comAdded Sep 17, 20260 opens
FreeArticle
Technology & Ethics

What Is RLCD? Reinforcement Learning from Contrast Distillation

A plain-language explainer on RLCD, a way of aligning language models by learning from contrasting outputs rather than human ratings alone.

From the site: RLCD is a method developed to adjust language models to human preferences without using human feedback data. This approach aims to address…

Why I recommend it: Good background reading if you want to understand how the models you use are actually steered.

#ai#ai ethics#ai-alignment#ai-safety#explainers#llm#machine-learning#model-training#reinforcement-learning#research#rlhf
MediumAdded Sep 17, 20260 opens
FreeCommunity
Technology & Ethics

AI Alignment Forum

A community blog where researchers publish and debate technical AI alignment work, free and open to read.

From the site: A community blog devoted to technical AI alignment research

Why I recommend it: Dense reading, but this is where a lot of safety research is argued out in public before it reaches papers.

#academic#ai#ai ethics#ai-alignment#ai-ethics#ai-safety#community#free-resource#governance#machine-learning#regulation#research
alignmentforum.orgAdded Sep 17, 20260 opens