Stanford's Center for Research on Foundation Models runs HELM as a living benchmark for language and multimodal models. Rather than one score, it reports many models across many scenarios on multiple metrics — accuracy, calibration, robustness, fairness, bias, toxicity and efficiency — and publishes the leaderboards alongside the raw model outputs (predictions and prompts) so you can check a claim yourself instead of taking a number on trust. Separate leaderboards cover areas such as classic HELM, instruction-following, medical, legal and safety. All results and analysis are free to browse on the site, no account.
Why I recommend it: The place to go when a vendor quotes you a benchmark figure. HELM's real value is that it shows the prompts and the model's actual answers, so you can see what the score measured. Be aware of what it is not: it is a snapshot of the model versions and dates CRFM ran, so check the run date before comparing anything to a model released since, and a model missing from a leaderboard usually means nobody ran it, not that it failed.
The Python framework behind Stanford's HELM leaderboards, released under the Apache License 2.0 (licence file read, not copied from a roundup). You install it with pip, describe a run (scenario plus model plus metrics), and it evaluates the model and produces the same structured results the public site displays, including its own local web UI for viewing them. It supports hosted model APIs and locally run open-weight models, and you can add your own scenario to test a model on your own task or data. Free to use, modify and use commercially under Apache 2.0; you pay only for whatever model API calls or compute your own runs consume.
Why I recommend it: Worth it if you need to prove a model is good enough for a specific job rather than good in general — write your own scenario with your own examples and run it. Two practical warnings: the published leaderboard runs are large and expensive to reproduce in full, so start with a single scenario and a small instance count, and if you evaluate a paid API model the token costs are yours, not Stanford's.
OpenAI's free framework for how misaligned model behaviour should be reported and categorised — what counts as misalignment, who reports it, and what happens next.
From the site: OpenAI shares a framework for tracking, investigating, and disclosing model misalignment, alongside six reports of unexpected or concerning model behavior.
Why I recommend it: Primary source on how a major lab defines and handles its own model failures — useful, but it is the lab grading itself.
A research nonprofit that independently evaluates frontier AI models to measure what they can actually do and what risks that creates. Reports are free.
From the site: METR is a research nonprofit that evaluates frontier AI models to inform the public about their risks and capabilities.
Why I recommend it: One of the few independent evaluators. Read their reports before you trust a lab's own capability claims.
Independent, free benchmarks testing leading AI models on real-world finance, software, science and safety tasks, with cost and latency alongside accuracy.
From the site: Private, domain-specific benchmarks in legal, tax, and finance.
Why I recommend it: When someone claims a model is "the best," check here — these are independent evaluations, not vendor marketing.