The claim
On HELM's Capabilities leaderboard, GPT-5 mini has the highest mean score of the 68 models listed, at 0.819 across five test sets.
Stanford HELM Capabilities, release v1.15.0 · check it yourself
What it’s worth: A real, checkable result from people with nothing to sell — and a narrow one. It says that on five specific test sets, this model answered more questions correctly. It does not say it is the best model for your work.
Question 1
Who produced the number?
Stanford's Center for Research on Foundation Models. Not a model vendor: HELM evaluates everyone's models on the same tests and publishes results that embarrass companies as readily as they flatter them.
Question 2
Compared with what, exactly?
Against the other 67 models in the same release, run through the same five test sets with the same prompts. That is the comparison you want, and it is rarer than it should be. The mean score is an unweighted average of the five — so a model can lead on the average while losing badly on the one scenario you care about.
Question 3
When was it measured?
HELM release v1.15.0; this table was downloaded on 23 September 2026. HELM does not publish a run date per model row, so treat every score as "as of this release", not as of today. New releases add models and can change ranks.
Question 4
Can you see the actual questions?
Yes — this is the reason to use HELM at all. Every prompt sent, and every answer the model gave, is published and readable. You can open a question, read the model's reasoning, and decide for yourself whether the grader was fair.
RLVR (Reinforcement Learning with Verifiable Rewards)Training / Training DataReasoning Model
Question 5
How many, and which, people or runs?
Countable, and published per test set: 1,000 MMLU-Pro questions, 446 GPQA, 541 IFEval, 1,000 WildBench, 1,000 Omni-MATH. No human testers and no anecdotes. One caveat: WildBench is scored by a judge model rather than against a right answer, so that column is a machine's opinion of quality.
Question 6
Does it say "up to", "as much as" or "can"?
No. These are plain percentages of questions answered correctly, not best cases. The one soft edge is that scores are for one run per model — a different run can land a point or two either way.
Question 7
Who paid for the study, and who benefits?
Stanford CRFM is academic and publishes its funders. Model providers give access and sometimes credits, which is how the models get evaluated at all; it is also a relationship worth knowing about when a table is read as a ranking.
Question 8
Is it measuring opinion, or measuring what happened?
Measured, not surveyed. Nobody was asked how they felt — a program answered questions and the answers were graded.
