OpenAI's benchmark of 1,215 realistic mental-health conversations, scored against rubrics written by 80+ licensed mental-health experts, covering everyday well-being through emergencies across ages and languages.
Why I recommend it: OpenAI built this benchmark and grades its own models on it, so treat the "steady progress" claim as a self-report until outside researchers replicate it. Useful for its honest list of weak spots: asking for context and judging urgency.
Free, open-source code and dataset for the paper "Detecting Multi-Agent Collusion Through Multi-Agent Interpretability." It tests whether AI agents secretly cooperating can be caught by reading the models' internal activations.
Why I recommend it: A research tool, not a beginner resource. Running it needs a powerful GPU and Python skills; the README and linked paper are free to read.
Public leaderboard and open-source benchmark that drops AI agents into realistic business environments with 47 real tools across sales, marketing, operations, support, finance, and HR. Scores are based on final environment state, not an LLM-as-judge.
Why I recommend it: The leaderboard and the benchmark code are free; running it yourself means paying the model APIs at the costs shown. The test design is based on Zapier's own task data, so it's a realistic lens on agent work, but Zapier also sells automation tools — treat the benchmark as a useful public dataset, not a neutral referee.