Question 1 of 5
A model tops the leaderboard. Can it do the job?
This is the gap that costs people money. A score measures answers to questions that have right answers. A job is mostly neither.
What the scores say
Stanford HELM Capabilities, release v1.15.0, read on 23 September 2026.
- GPT-5 mini (2025-08-07) leads the 68 models listed with a mean of 0.819 across five test sets.
- Every prompt and every answer behind those scores is published, so the grading can be checked by anyone.
- The five sets are graduate-level science questions, hard multiple choice, instruction following, open-ended chat quality and competition math.
What the record shows
Documented cases from the impacts page, each with a named source and date.
- Klarna said its AI assistant was doing the work of 700 agents, then began rehiring humans for customer service after quality complaints.
- Announced US job cuts attributing the decision to AI are employer explanations, not audits of what a model actually replaced.
- No test set on the HELM board contains a distressed customer, an ambiguous refund policy, or a colleague who will be blamed if the answer is wrong.
What the two together support
Both are true at once, and they are not in conflict. The score says the model answers exam questions better than the others on that release. The record says exam performance did not survive contact with customers at one well-documented company. What neither supports is the leap people make in meetings — that a leaderboard position predicts whether a deployment will hold up.
Which evidence is weaker
The record side, clearly. One company's reversal is an anecdote, and Klarna has commercial reasons for both the original claim and the walk-back. The score is the sturdier number here; it is just answering a narrower question than the one being asked of it.
