FreeTool
Technology & Ethics
NARCBench: Detecting AI Agent Collusion
Free, open-source code and dataset for the paper "Detecting Multi-Agent Collusion Through Multi-Agent Interpretability." It tests whether AI agents secretly cooperating can be caught by reading the models' internal activations.
Why I recommend it: A research tool, not a beginner resource. Running it needs a powerful GPU and Python skills; the README and linked paper are free to read.
#ai ethics#ai safety#benchmark#interpretability#multi-agent systems#open source#research
github.comAdded Sep 23, 20260 opens
