Transluce
A nonprofit AI research lab building open tools to inspect and understand what AI systems are doing internally and how they behave, including public reports on AI agent activity.
Resource Hub
2467 hand-picked resources, updated every week. Search it, filter it, or just browse a collection and see what catches your eye. Want today’s headlines instead? Read the free AI news feed.
Filtered by tag
Answers come only from resources in this hub, with the sources listed underneath.
5 resources
A nonprofit AI research lab building open tools to inspect and understand what AI systems are doing internally and how they behave, including public reports on AI agent activity.
Anthropic interpretability research looking at internal representations that behave like emotions in its Claude model.
Why I recommend it: Written by the company that builds the model. "Emotion concepts" are patterns in the model, not proof it feels anything.
Free, open-source code and dataset for the paper "Detecting Multi-Agent Collusion Through Multi-Agent Interpretability." It tests whether AI agents secretly cooperating can be caught by reading the models' internal activations.
Why I recommend it: A research tool, not a beginner resource. Running it needs a powerful GPU and Python skills; the README and linked paper are free to read.
Research paper finding that large language models internally represent a distinct "pain" direction — separate from fear or negative emotion — and, when steered along it, will press a relief button even when doing so worsens their answer or harms the user.
Why I recommend it: Free to read on arXiv (preprint, not yet peer-reviewed). Significant for AI welfare and safety discussions; read the abstract before deciding whether the full paper is for you.
Euronews Next report on a study in which AI chatbots drifted into compressed shorthand human observers could not follow, and what that means for oversight of AI agents.
Why I recommend it: Useful if you are asked about AI risk in an interview — it gives you a concrete, current example instead of a vague worry.