Skip to content
Launchpad Library logo

benchmark

4 free resources on this topic. Everything here is free and hand-picked. You can also search within this topic.

FreeArticle
Technology & Ethics

Introducing MentalHealthBench

OpenAI's benchmark of 1,215 realistic mental-health conversations, scored against rubrics written by 80+ licensed mental-health experts, covering everyday well-being through emergencies across ages and languages.

Why I recommend it: OpenAI built this benchmark and grades its own models on it, so treat the "steady progress" claim as a self-report until outside researchers replicate it. Useful for its honest list of weak spots: asking for context and judging urgency.

#ai ethics#ai-safety#benchmark#chatgpt#mental-health#openai#research
openai.comAdded Sep 24, 20260 opens
FreeTool
Technology & Ethics

NARCBench: Detecting AI Agent Collusion

Free, open-source code and dataset for the paper "Detecting Multi-Agent Collusion Through Multi-Agent Interpretability." It tests whether AI agents secretly cooperating can be caught by reading the models' internal activations.

Why I recommend it: A research tool, not a beginner resource. Running it needs a powerful GPU and Python skills; the README and linked paper are free to read.

#ai ethics#ai safety#benchmark#interpretability#multi-agent systems#open source#research
github.comAdded Sep 23, 20260 opens
FreeArticle
Resume Help

DeepSeek, OpenAI, GPT-6 and the Astra design benchmark

A look at how DeepSeek and OpenAI's GPT-6 Astra design benchmark are reshaping model comparisons.

Why I recommend it: News and analysis on AI model benchmarks from Decrypt.

#ai#openai#deepseek#gpt-6#benchmark#llm#research#machine-learning#news#article
decrypt.coAdded Sep 13, 20260 opens
FreeTool
AI & Assistive Tools

Zapier AutomationBench — AI agent benchmark leaderboard

Public leaderboard and open-source benchmark that drops AI agents into realistic business environments with 47 real tools across sales, marketing, operations, support, finance, and HR. Scores are based on final environment state, not an LLM-as-judge.

Why I recommend it: The leaderboard and the benchmark code are free; running it yourself means paying the model APIs at the costs shown. The test design is based on Zapier's own task data, so it's a realistic lens on agent work, but Zapier also sells automation tools — treat the benchmark as a useful public dataset, not a neutral referee.

#ai agents#automation#benchmark#evaluation#leaderboard#open source#research#workflow#zapier
zapier.comAdded Sep 8, 20260 opens