HELM (open source framework on GitHub)
The Python framework behind Stanford's HELM leaderboards, released under the Apache License 2.0 (licence file read, not copied from a roundup). You install it with pip, describe a run (scenario plus model plus metrics), and it evaluates the model and produces the same structured results the public site displays, including its own local web UI for viewing them. It supports hosted model APIs and locally run open-weight models, and you can add your own scenario to test a model on your own task or data. Free to use, modify and use commercially under Apache 2.0; you pay only for whatever model API calls or compute your own runs consume.
Why I recommend it: Worth it if you need to prove a model is good enough for a specific job rather than good in general — write your own scenario with your own examples and run it. Two practical warnings: the published leaderboard runs are large and expensive to reproduce in full, so start with a single scenario and a small instance count, and if you evaluate a paid API model the token costs are yours, not Stanford's.
