Skip to content
Launchpad Library logo

Measured, not marketed

Real model scores

Three groups publish numbers about AI models that you can actually check: Stanford’s HELM, which publishes every prompt and every answer behind a score; METR, which measures how long a task an agent can finish; and Artificial Analysis, a for-profit company that tests models through their live APIs. The HELM and METR figures were read straight out of their own published data files on 23 September 2026; the Artificial Analysis figures were checked on its leaderboard on September 28, 2026. Nothing here is a vendor claim, and nothing here is my estimate.

All three are free to read. HELM and METR publish their raw data; Artificial Analysis publishes its methodology and scores. Below each table, the headline claim gets the same eight-question checklist this site applies to anyone else’s number — including theirs.

A score is not an outcome. These tables say what models did on tests. For what has actually happened to people so far, and an honest argument about how the two fit together, read Scores vs. Reality.

HELM Capabilities: 68 models, five test sets

HELM runs every model through the same tests with the same prompts, and publishes the questions and the answers so you can judge the grading yourself. The mean score is the plain average of the five columns — useful for a rough ranking, and worth ignoring in favor of the one column that matches your work.

Source: Stanford HELM Capabilities, release v1.15.0. Scores are shares of questions answered correctly, except WildBench, which is quality scored by a judge model.

MMLU-Pro · 1,000 questions
Multiple-choice exam questions across 14 subjects, answered with reasoning shown. Scored on the share answered correctly.
GPQA · 446 questions
Graduate-level physics, chemistry and biology questions written to be hard to look up. Scored on the share answered correctly.
IFEval · 541 questions
Whether the model does exactly what the instruction said (word counts, formats, forbidden words). Scored strictly.
WildBench · 1,000 questions
Real user requests scored for quality by a judge model, not by a right answer. The softest of the five.
Omni-MATH · 1,000 questions
Olympiad-level math problems. Scored on the share solved.
HELM Capabilities scores, release v1.15.0, highest mean score first
ModelMeanMMLU-ProGPQAIFEvalWildBenchOmni-MATH
GPT-5 mini (2025-08-07)81.9%83.5%75.6%92.7%85.5%72.2%
o4-mini (2025-04-16)81.2%82.0%73.5%92.9%85.4%72.0%
o3 (2025-04-16)81.1%85.9%75.3%86.9%86.1%71.4%
GPT-5 (2025-08-07)80.7%86.3%79.1%87.5%85.7%64.7%
Gemini 3 Pro (Preview)79.9%90.3%80.3%87.6%85.9%55.5%
Qwen3 235B A22B Instruct 2507 FP879.8%84.4%72.6%83.5%86.6%71.8%
Grok 4 (0709)78.5%85.1%72.6%94.9%79.7%60.3%
Claude 4 Opus (20250514, extended thinking)78.0%87.5%70.9%84.9%85.2%61.6%
gpt-oss-120b77.0%79.5%68.4%83.6%84.5%68.8%
Kimi K2 Instruct76.8%81.9%65.2%85.0%86.2%65.4%
Claude 4 Sonnet (20250514, extended thinking)76.6%84.3%70.6%84.0%83.8%60.2%
Claude 4.5 Sonnet (20250929)76.2%86.9%68.6%85.0%85.4%55.3%
Claude 4 Opus (20250514)75.7%85.9%66.6%91.8%83.3%51.1%
GPT-5 nano (2025-08-07)74.8%77.8%67.9%93.2%80.6%54.7%
Gemini 2.5 Pro (03-25 preview)74.5%86.3%74.9%84.0%85.7%41.6%
Claude 4 Sonnet (20250514)73.3%84.3%64.3%83.9%82.5%51.2%
Grok 3 Beta72.7%78.8%65.0%88.4%84.9%46.4%
GPT-4.1 (2025-04-14)72.7%81.1%65.9%83.8%85.4%47.1%
Qwen3 235B A22B FP8 Throughput72.6%81.7%62.3%81.6%82.8%54.8%
GPT-4.1 mini (2025-04-14)72.6%78.3%61.4%90.4%83.8%49.1%

Showing 20 of 68 models. Read the prompts and answers on HELM.

METR time horizons: how long a task, and how often

METR measures something more concrete than a test score: the length of software task — in how long it takes a person — that a model finishes reliably. Read the two columns together. The 50% column is a coin flip; the 80% column is closer to what you would accept from a colleague, and it is always far shorter.

Source: METR Time Horizon 1.1, from their published results file. Ranges are the confidence interval METR publishes for the 50% figure.

METR Time Horizon 1.1 time horizons, longest first
ModelReleasedFinishes half the timeRangeFinishes 80% of the time
Claude Mythos Preview (early)2026-04-0717 hr8.5 hr – 2.3 days3.1 hr
Claude Opus 4.62026-02-0512 hr5.3 hr – 2.5 days1.2 hr
Gemini 3.1 Pro2026-02-196.4 hr3.9 hr – 12 hr1.5 hr
GPT-5.2 (high)2025-12-115.9 hr3.3 hr – 14 hr1.1 hr
GPT-5.3-Codex (high)2026-02-055.8 hr3.2 hr – 14 hr55 min
GPT-5.4 (xhigh)2026-03-055.7 hr3.1 hr – 13 hr54 min
Claude Opus 4.52025-11-244.9 hr2.7 hr – 10 hr49 min
Gemini 3 Pro2025-11-183.7 hr2.3 hr – 6.3 hr54 min
GPT-5.1-Codex-Max2025-11-193.7 hr2.2 hr – 6.6 hr51 min
GPT-52025-08-073.4 hr1.9 hr – 6.8 hr38 min
o32025-04-162.0 hr1.2 hr – 3.2 hr30 min
Claude Opus 4.12025-08-051.7 hr59 min – 2.7 hr23 min
Claude Opus 42025-05-221.7 hr60 min – 2.7 hr20 min
Claude 3.7 Sonnet2025-02-241.0 hr33 min – 1.7 hr12 min
o12024-12-0539 min21 min – 1.1 hr7.1 min
Claude 3.5 Sonnet (Oct 2024)2024-10-2221 min10 min – 41 min2.6 min
o1-preview2024-09-1220 min12 min – 33 min4.4 min
Claude 3.5 Sonnet (Jun 2024)2024-06-2011 min5.5 min – 22 min1.7 min
GPT-4o2024-05-137.0 min4.0 min – 13 min1.3 min
GPT-4 (1106)2023-11-064.0 min1.9 min – 8.4 min47 sec
GPT-42023-03-144.0 min1.9 min – 8.0 min53 sec
Claude 3 Opus2024-03-044.0 min1.7 min – 8.8 min38 sec
GPT-4 Turbo2024-04-093.7 min2.0 min – 6.7 min56 sec
GPT-3.5 Turbo Instruct2022-03-1536 sec15 sec – 1.1 min15 sec
davinci-0022020-05-289 sec6 sec – 13 sec3 sec
GPT-22019-02-143 sec1 sec – 9 sec1 sec

METR’s own note: measurements above roughly 16 hours are unreliable with their current task suite, and those points are left out of the trend estimate. Doubling time since 2023: 129 days (range 104–158). METR’s methodology and FAQ.

Artificial Analysis Intelligence Index: one composite score

Artificial Analysis is an independent benchmarking company that tests AI models directly through their live APIs, rather than repeating the numbers labs claim, and combines results from reasoning, coding, agentic and knowledge-work tests into one composite score, the Intelligence Index.

A different kind of independence. HELM is an academic project and METR is a nonprofit. Artificial Analysis is a for-profit company that also sells private benchmarking services directly to the AI labs it ranks in public. Its founders disclose this openly, and its policy walls that work off from the public leaderboard. It is still a real structural difference, and this page does not treat the three sources as equally independent.

Source: Artificial Analysis Intelligence Index, v4.3.2, checked September 28, 2026.

Artificial Analysis Intelligence Index v4.3.2, highest score first
RankModelDeveloperIntelligence Index
1Claude Opus 5.5 (max with fallback)Anthropic58
2Claude Sonnet 5.5 (max with fallback)Anthropic56
3Claude Fable 5.1 (max with fallback)Anthropic53
4GPT-6 Astra (max)OpenAI53
5Muse Spark 1.3 (max)Meta48
6GPT-6 Sol (max)OpenAI48
7Grok 4.7 (xhigh)SpaceXAI (xAI)46
8MiMo-V2.6-ProXiaomi46
9Qwen3.8 Max (0902)Alibaba45
10GLM-5.3 (max)Z.ai45
11Step 5 PreviewStepFun44
12Kimi K3 (max)Kimi44
13GLM-5.3-FlashZ.ai42
14Gemini 3.8 Flash (high)Google41
15DeepSeek V4.1 Flash (max)DeepSeek39

Methodology: Artificial Analysis tests models through live vendor APIs, standardizes tokenization so results compare across providers, and publishes confidence intervals of about ±1%. It states a “no pay-to-rank” policy on its public leaderboard. Artificial Analysis methodology. See the full leaderboard on Artificial Analysis.

The checklist, applied to each claim

These are the best-sourced model numbers I know of, and they still need the checklist. Every question from AI Basics gets an answer for each claim — including the awkward ones about who funds the measuring.

The claim

On HELM's Capabilities leaderboard, GPT-5 mini has the highest mean score of the 68 models listed, at 0.819 across five test sets.

Stanford HELM Capabilities, release v1.15.0 · check it yourself

What it’s worth: A real, checkable result from people with nothing to sell — and a narrow one. It says that on five specific test sets, this model answered more questions correctly. It does not say it is the best model for your work.

  1. Question 1

    Who produced the number?

    Stanford's Center for Research on Foundation Models. Not a model vendor: HELM evaluates everyone's models on the same tests and publishes results that embarrass companies as readily as they flatter them.

    AI WashingInferenceAI Safety

  2. Question 2

    Compared with what, exactly?

    Against the other 67 models in the same release, run through the same five test sets with the same prompts. That is the comparison you want, and it is rarer than it should be. The mean score is an unweighted average of the five — so a model can lead on the average while losing badly on the one scenario you care about.

    InferenceScaling Laws

  3. Question 3

    When was it measured?

    HELM release v1.15.0; this table was downloaded on 23 September 2026. HELM does not publish a run date per model row, so treat every score as "as of this release", not as of today. New releases add models and can change ranks.

    Training / Training DataContext Window

  4. Question 4

    Can you see the actual questions?

    Yes — this is the reason to use HELM at all. Every prompt sent, and every answer the model gave, is published and readable. You can open a question, read the model's reasoning, and decide for yourself whether the grader was fair.

    RLVR (Reinforcement Learning with Verifiable Rewards)Training / Training DataReasoning Model

  5. Question 5

    How many, and which, people or runs?

    Countable, and published per test set: 1,000 MMLU-Pro questions, 446 GPQA, 541 IFEval, 1,000 WildBench, 1,000 Omni-MATH. No human testers and no anecdotes. One caveat: WildBench is scored by a judge model rather than against a right answer, so that column is a machine's opinion of quality.

    Agentic AI / AI AgentAgent HarnessVibe Coding

  6. Question 6

    Does it say "up to", "as much as" or "can"?

    No. These are plain percentages of questions answered correctly, not best cases. The one soft edge is that scores are for one run per model — a different run can land a point or two either way.

    AI WashingInference

  7. Question 7

    Who paid for the study, and who benefits?

    Stanford CRFM is academic and publishes its funders. Model providers give access and sometimes credits, which is how the models get evaluated at all; it is also a relationship worth knowing about when a table is read as a ranking.

    AI EthicsAI Washing

  8. Question 8

    Is it measuring opinion, or measuring what happened?

    Measured, not surveyed. Nobody was asked how they felt — a program answered questions and the answers were graded.

    Doomer (AI context)AccelerationismAI Ethics

The claim

METR's Time Horizon 1.1 measurement puts Claude Mythos Preview (early) at a 50% time horizon of 17 hr — the length of software task, measured in how long humans take, that it finishes about half the time.

METR, Time Horizon 1.1 · check it yourself

What it’s worth: The most useful single number published about agents, and the easiest to overstate. Half the time is a coin flip, the confidence range is enormous, and the tasks are software tasks with a clear right answer — not your job.

  1. Question 1

    Who produced the number?

    METR, a non-profit that evaluates models for dangerous capability, including pre-release testing paid for by the labs whose models it tests. That funding relationship is worth holding in mind; against it, METR publishes its full task data, its code, and results that undercut lab marketing.

    AI WashingInferenceAI Safety

  2. Question 2

    Compared with what, exactly?

    Against human contractors doing the same tasks: the horizon is expressed in human time, so "17 hr" means tasks that took people that long. Across the same suite, GPT-2 sits at 3 sec.

    InferenceScaling Laws

  3. Question 3

    When was it measured?

    Time Horizon 1.1, downloaded 23 September 2026. METR marks its own older write-ups as out of date rather than quietly editing them, which is the behavior you want from a measurement.

    Training / Training DataContext Window

  4. Question 4

    Can you see the actual questions?

    Partly. The methodology, the analysis code and the raw results file are public, and a large part of the task suite is published. Some tasks are deliberately held back so models cannot be trained on them — reasonable, and it means you cannot inspect every single task yourself.

    RLVR (Reinforcement Learning with Verifiable Rewards)Training / Training DataReasoning Model

  5. Question 5

    How many, and which, people or runs?

    Each figure is a curve fitted across many task attempts, which is why METR publishes a confidence interval rather than a single number. For the top model the 50% horizon lands anywhere between 8.5 hr and 2.3 days. Quoting the middle number alone is the most common misuse of this table.

    Agentic AI / AI AgentAgent HarnessVibe Coding

  6. Question 6

    Does it say "up to", "as much as" or "can"?

    Not in the wording, but the 50% figure functions like one. At 80% success — closer to what you would accept from a colleague — the same model drops to 3.1 hr. Both columns are below.

    AI WashingInference

  7. Question 7

    Who paid for the study, and who benefits?

    METR is funded by donations and by paid pre-release evaluation work for AI companies. It says so openly. Read the numbers; be careful reading the framing as independent endorsement in either direction.

    AI EthicsAI Washing

  8. Question 8

    Is it measuring opinion, or measuring what happened?

    Measured, and measured on software tasks with checkable answers — not on messy work with no right answer, and not on anyone's actual job. It says nothing about jobs.

    Doomer (AI context)AccelerationismAI Ethics

The claim

METR estimates the 50% time horizon has been doubling roughly every 129 days for models released since 2023 — faster than the 188-day doubling across the whole record.

METR, doubling-time estimates in the published results file · check it yourself

What it’s worth: A trend fitted to about two dozen points, with a stated range of 104 to 158 days. Worth knowing. Not a schedule, and not a thing to extrapolate five years forward.

  1. Question 1

    Who produced the number?

    METR, from its own results file — the same numbers in the table below, fitted with a regression whose code is public.

    AI WashingInferenceAI Safety

  2. Question 2

    Compared with what, exactly?

    The recent trend is compared with the trend across the full record since 2019. The recent one is faster, which is the finding.

    InferenceScaling Laws

  3. Question 3

    When was it measured?

    From the Time Horizon 1.1 results file, downloaded 23 September 2026. Every doubling estimate has a cut-off date built in, and each new model shifts it.

    Training / Training DataContext Window

  4. Question 4

    Can you see the actual questions?

    The underlying per-model numbers are all published and listed below, so you can fit your own line and see whether you get the same answer.

    RLVR (Reinforcement Learning with Verifiable Rewards)Training / Training DataReasoning Model

  5. Question 5

    How many, and which, people or runs?

    26 models, and points above roughly 16 hours are excluded because METR says its current task suite cannot measure them reliably. A trend line through two dozen points is a sketch, not a law.

    Agentic AI / AI AgentAgent HarnessVibe Coding

  6. Question 6

    Does it say "up to", "as much as" or "can"?

    No "up to", but the single number hides a range: METR publishes 104 to 158 days. Anyone quoting the midpoint as a fact is rounding away the uncertainty.

    AI WashingInference

  7. Question 7

    Who paid for the study, and who benefits?

    Same as above: donations plus paid pre-release evaluation for AI companies, disclosed on their site.

    AI EthicsAI Washing

  8. Question 8

    Is it measuring opinion, or measuring what happened?

    Measured capability on software tasks, extrapolated as a trend. It is not a forecast of employment, and METR does not present it as one.

    Doomer (AI context)AccelerationismAI Ethics

The claim

Claude Opus 5.5 has the highest Artificial Analysis Intelligence Index score of any model, at 58.

Artificial Analysis Intelligence Index v4.3.2, checked September 28, 2026 · check it yourself

What it’s worth: A measured composite from live API tests, published with confidence intervals — and a two-point lead over the next model, near the edge of what the stated ±1% range can separate. The company ranking the models also sells private benchmarking to some of the labs it ranks.

  1. Question 1

    Who produced the number?

    Artificial Analysis, a for-profit benchmarking company. It is not a model maker, and it runs the tests itself through each vendor's live API rather than copying numbers the labs report.

    AI WashingInferenceAI Safety

  2. Question 2

    Compared with what, exactly?

    Against the other models on the same leaderboard, run through the same set of reasoning, coding, agentic and knowledge-work tests. The Index combines those into one composite, so a model can lead the composite while trailing on the single test you care about. The 58 is for the 'max with fallback' setting, which may not be the setting you get in a chat app.

    InferenceScaling Laws

  3. Question 3

    When was it measured?

    Index v4.3.2, checked September 28, 2026. The Index is versioned, and a new version can reweight tests and shift every score, so a number is only comparable with numbers from the same version.

    Training / Training DataContext Window

  4. Question 4

    Can you see the actual questions?

    Partly. The list of tests and the methodology are published, and many of the component tests are public benchmarks. The full prompts and answers behind each score are not published the way HELM publishes them.

    RLVR (Reinforcement Learning with Verifiable Rewards)Training / Training DataReasoning Model

  5. Question 5

    How many, and which, people or runs?

    Artificial Analysis publishes confidence intervals of about ±1% on the Index. Claude Opus 5.5 at 58 and Claude Sonnet 5.5 at 56 are separated by more than that, but not by much; the models tied at 53, 48, 46 and 45 cannot be told apart by this number.

    Agentic AI / AI AgentAgent HarnessVibe Coding

  6. Question 6

    Does it say "up to", "as much as" or "can"?

    No 'up to' wording. The one soft edge is 'highest of any model': it means highest of the models Artificial Analysis has tested, at the settings it tested.

    AI WashingInference

  7. Question 7

    Who paid for the study, and who benefits?

    This is the question to weigh here. Artificial Analysis states a 'no pay-to-rank' policy for its public leaderboard, and says its private work is walled off from it. The company also sells private benchmarking services directly to AI labs, including some of the labs ranked in this table, and its founders disclose that openly. A lab paying the scorekeeper for other work is a real conflict of interest, even with a policy against it, and readers should decide how much independence to credit this number with that in mind.

    AI EthicsAI Washing

  8. Question 8

    Is it measuring opinion, or measuring what happened?

    Measured, not surveyed. Programs answered test questions and completed tasks, and the results were graded. It says nothing about any particular job.

    Doomer (AI context)AccelerationismAI Ethics

What these numbers can’t tell you

  • None of them measures the model you can use for free. Scores are run against the paid API at settings you may not have; a free tier is usually a smaller model, a shorter memory, or a daily cap.
  • None of them measures your work. HELM asks exam questions; METR runs software tasks with a checkable answer. Writing a cover letter, reading a contract, or judging a person is not on any of these lists.
  • A leaderboard is a snapshot. Every table is tied to a release and a check date, and a model can be retired, retrained or renamed without the score changing.
  • Ranks are close. A gap of one or two points between models is usually noise; a gap of twenty is not.