Stanford's Center for Research on Foundation Models runs HELM as a living benchmark for language and multimodal models. Rather than one score, it reports many models across many scenarios on multiple metrics — accuracy, calibration, robustness, fairness, bias, toxicity and efficiency — and publishes the leaderboards alongside the raw model outputs (predictions and prompts) so you can check a claim yourself instead of taking a number on trust. Separate leaderboards cover areas such as classic HELM, instruction-following, medical, legal and safety. All results and analysis are free to browse on the site, no account.
Why I recommend it: The place to go when a vendor quotes you a benchmark figure. HELM's real value is that it shows the prompts and the model's actual answers, so you can see what the score measured. Be aware of what it is not: it is a snapshot of the model versions and dates CRFM ran, so check the run date before comparing anything to a model released since, and a model missing from a leaderboard usually means nobody ran it, not that it failed.
The Python framework behind Stanford's HELM leaderboards, released under the Apache License 2.0 (licence file read, not copied from a roundup). You install it with pip, describe a run (scenario plus model plus metrics), and it evaluates the model and produces the same structured results the public site displays, including its own local web UI for viewing them. It supports hosted model APIs and locally run open-weight models, and you can add your own scenario to test a model on your own task or data. Free to use, modify and use commercially under Apache 2.0; you pay only for whatever model API calls or compute your own runs consume.
Why I recommend it: Worth it if you need to prove a model is good enough for a specific job rather than good in general — write your own scenario with your own examples and run it. Two practical warnings: the published leaderboard runs are large and expensive to reproduce in full, so start with a single scenario and a small instance count, and if you evaluate a paid API model the token costs are yours, not Stanford's.