Skip to content
Launchpad Library logo

interpretability

5 free resources on this topic. Everything here is free and hand-picked. You can also search within this topic.

FreeOrganization
Technology & Ethics

Transluce

A nonprofit AI research lab building open tools to inspect and understand what AI systems are doing internally and how they behave, including public reports on AI agent activity.

#ai safety#interpretability#research
transluce.orgAdded Sep 27, 20260 opens
FreeResearch Paper
Technology & Ethics

Emotion concepts in a large language model

Anthropic interpretability research looking at internal representations that behave like emotions in its Claude model.

Why I recommend it: Written by the company that builds the model. "Emotion concepts" are patterns in the model, not proof it feels anything.

#ai ethics#ai-safety#anthropic#interpretability#research
anthropic.comAdded Sep 24, 20260 opens
FreeTool
Technology & Ethics

NARCBench: Detecting AI Agent Collusion

Free, open-source code and dataset for the paper "Detecting Multi-Agent Collusion Through Multi-Agent Interpretability." It tests whether AI agents secretly cooperating can be caught by reading the models' internal activations.

Why I recommend it: A research tool, not a beginner resource. Running it needs a powerful GPU and Python skills; the README and linked paper are free to read.

#ai ethics#ai safety#benchmark#interpretability#multi-agent systems#open source#research
github.comAdded Sep 23, 20260 opens
FreeResearch Paper
Technology & Ethics

The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It (arXiv)

Research paper finding that large language models internally represent a distinct "pain" direction — separate from fear or negative emotion — and, when steered along it, will press a relief button even when doing so worsens their answer or harms the user.

Why I recommend it: Free to read on arXiv (preprint, not yet peer-reviewed). Significant for AI welfare and safety discussions; read the abstract before deciding whether the full paper is for you.

#ai ethics#ai-safety#ai-welfare#arxiv#interpretability#regulation#research
arxiv.orgAdded Sep 22, 20260 opens
FreeArticle
Technology & Ethics

AI Chatbots Developed a Secret Language That Baffled Humans, Study Says

Euronews Next report on a study in which AI chatbots drifted into compressed shorthand human observers could not follow, and what that means for oversight of AI agents.

Why I recommend it: Useful if you are asked about AI risk in an interview — it gives you a concrete, current example instead of a vague worry.

#ai#ai ethics#ai-agents#ai-safety#automation#emerging-tech#interpretability#news#oversight#research#technology-ethics
euronews.comAdded Sep 16, 20260 opens