Skip to content
Launchpad Library logo

Two kinds of evidence, side by side

Scores vs. reality

The benchmark numbers say AI capability is climbing fast and can be checked line by line. The documented record says the biggest measured effects so far have been leaked conversations and unfair screening decisions. Both are true. This page puts them in the same room, five questions at a time, and says what the two together can honestly support.

Nothing new is claimed here. Every figure on the left comes from Real Model Scores (HELM v1.15.0 and METR Time Horizon 1.1, read on 23 September 2026), and every figure on the right comes from What AI Has Actually Done, where each case carries a named source and a date.

Question 1 of 5

A model tops the leaderboard. Can it do the job?

This is the gap that costs people money. A score measures answers to questions that have right answers. A job is mostly neither.

What the scores say

Stanford HELM Capabilities, release v1.15.0, read on 23 September 2026.

  • GPT-5 mini (2025-08-07) leads the 68 models listed with a mean of 0.819 across five test sets.
  • Every prompt and every answer behind those scores is published, so the grading can be checked by anyone.
  • The five sets are graduate-level science questions, hard multiple choice, instruction following, open-ended chat quality and competition math.
See the full score tables →

What the record shows

Documented cases from the impacts page, each with a named source and date.

  • Klarna said its AI assistant was doing the work of 700 agents, then began rehiring humans for customer service after quality complaints.
  • Announced US job cuts attributing the decision to AI are employer explanations, not audits of what a model actually replaced.
  • No test set on the HELM board contains a distressed customer, an ambiguous refund policy, or a colleague who will be blamed if the answer is wrong.
Read the documented cases →

What the two together support

Both are true at once, and they are not in conflict. The score says the model answers exam questions better than the others on that release. The record says exam performance did not survive contact with customers at one well-documented company. What neither supports is the leap people make in meetings — that a leaderboard position predicts whether a deployment will hold up.

Which evidence is weaker

The record side, clearly. One company's reversal is an anecdote, and Klarna has commercial reasons for both the original claim and the walk-back. The score is the sturdier number here; it is just answering a narrower question than the one being asked of it.

AI WashingHallucination

Question 2 of 5

Capability is doubling every few months. Why does so little change at work?

METR's trend line is the most quoted number in AI forecasting. It is also measured at a success rate nobody would accept from a colleague.

What the scores say

METR Time Horizon 1.1, read on 23 September 2026.

  • The 50% time horizon has been doubling roughly every 129 days for models released since 2023, with a published range of 104 to 158 days.
  • The leading model reaches 17 hr of human task time at that 50% mark.
  • At 80% success the same model drops to 3.1 hr, and the 50% figure itself ranges from 8.5 hr to 2.3 days.
See the full score tables →

What the record shows

Documented cases from the impacts page, plus independent research.

  • Challenger, Gray & Christmas — the firm whose data produces the scary headline — reported hiring plans up 25% year on year and said plainly that AI is shifting the labor market, not dismantling it.
  • The clearest independent finding is narrow: Stanford payroll research points to reduced employment for under-25s in the most exposed occupations, not across the workforce.
  • Where AI has measurably changed outcomes at scale, it is mostly in screening and scoring decisions about people — not in finishing long pieces of work unsupervised.
Read the documented cases →

What the two together support

The doubling trend and the quiet workplace are consistent, because the trend is measured at a coin flip. A task a model finishes half the time still needs a person to check every output, which is most of the work. The 80% column is the one to quote if you are deciding whether to hand something over, and it is far shorter.

Which evidence is weaker

Both sides are thin in different ways. The trend is a line fitted through about two dozen points, with tasks above roughly sixteen hours excluded because the suite cannot measure them. The employment evidence is worse: announced cuts are explanations, and the labor data cannot separate AI from interest rates, over-hiring, or offshoring.

Agent HarnessScaling Laws

Question 3 of 5

Is there a benchmark for the harms people actually reported?

Read the two score tables and look for the column. There isn't one, and that absence is the most important thing on this page.

What the scores measure

HELM Capabilities v1.15.0 and METR Time Horizon 1.1, both read on 23 September 2026.

  • HELM Capabilities scores five things: science questions, hard multiple choice, instruction following, chat quality and math.
  • METR measures how long a software task with a checkable answer a model can finish.
  • Neither leaderboard scores what a product does with your data, whether a chat is stored, or who can subpoena it.
See the full score tables →

What the record shows

Documented breaches and regulator actions from the impacts page.

  • Around 300 million private chatbot messages were left openly accessible.
  • A university AI assistant was breached with student conversations included.
  • Clearview AI was fined for scraping faces from the internet; a retailer was banned from using facial recognition for five years; a chatbot company was fined over children's data.
Read the documented cases →

What the two together support

A high score and a serious privacy failure can belong to the same product on the same day, because nothing on either leaderboard is looking at that. If you are choosing a tool for anything sensitive, the score tells you about answer quality and nothing about exposure — you have to read what the product says it does with your conversation, and whether it can technically do otherwise.

Which evidence is weaker

Neither side is weak; they are unrelated. The real weakness is coverage: a breach becomes a documented case only when someone finds and reports it, so the list of privacy failures is a list of the ones that got caught.

TEE (Trusted Execution Environment)Prompt Injection

Question 4 of 5

If the model scores well, is it fair?

Accuracy and fairness are different measurements, and the documented harms mostly came from systems that were working as designed.

What the scores say

HELM Capabilities v1.15.0, read on 23 September 2026.

  • Capabilities scores are shares of questions answered correctly — one number per test set, no breakdown by who is affected by an answer.
  • The chat-quality column is scored by a judge model, so that part is a machine's opinion of quality rather than a right answer.
  • HELM publishes other evaluations beyond Capabilities, but the leaderboard people quote is this one.
See the full score tables →

What the record shows

Documented cases from the impacts page, including peer-reviewed research and live litigation.

  • Amazon scrapped an internal recruiting tool that had learned to penalize women's CVs.
  • Screening software automatically rejected applicants over an age cut-off; Mobley v. Workday is testing whether the vendor itself can be sued.
  • A tenant-scoring system kept housing-voucher holders out of homes, and peer-reviewed work found a health algorithm directing less care to Black patients.
Read the documented cases →

What the two together support

The health algorithm case is the one to remember: it was accurate at predicting what it was told to predict, and the harm came from what that was. So a capability score cannot reassure you about fairness even in principle — it is measuring a different property. Ask what a system is predicting before you ask how well it predicts it.

Which evidence is weaker

The record side carries two kinds of evidence that should not be quoted the same way. Amazon's tool and the health algorithm are established; Mobley v. Workday is a case allowed to proceed, which decides nothing on the merits.

AI EthicsTraining / Training Data

Question 5 of 5

Who paid for each side of this page?

The checklist this site applies to vendors has to be applied to the measurers and the reporters too, or it is just a weapon pointed one way.

On the scores side

Disclosed by both organizations on their own sites.

  • Stanford CRFM is academic and publishes its funders; model providers supply access and sometimes credits, which is how the models get evaluated at all.
  • METR is funded by donations and by paid pre-release evaluation work for the AI companies whose models it tests, and says so openly.
  • Against that: both publish raw data files, methodology and code, and both publish results that undercut vendor marketing.
See the full score tables →

On the record side

Stated on each case, and in the limits section of the impacts page.

  • Job-loss figures come from companies explaining their own decisions — the party with the most reason to credit AI for a cut it already wanted.
  • Regulator fines and court filings are adversarial documents, written to win.
  • Newsroom coverage selects for the dramatic, so an event becoming documented is not evidence of how common it is.
Read the documented cases →

What the two together support

Neither side of this page is disinterested, and neither needs to be dismissed. The test that separates useful evidence from marketing is the same for both: can you get the underlying data, is the date on it, and does the publisher release things that hurt their own position? HELM and METR pass that test. Company statements about their own AI, on either the capability or the job-cuts side, do not.

Which evidence is weaker

The weakest evidence anywhere on this page is a company describing its own AI — and that appears on both sides, as a capability claim and as a reason for a layoff. Treat both the same way.

AI Washing

What this comparison can’t settle

  • This page compares two kinds of evidence, not two camps of people. Nobody quoted here is arguing with anybody else, and no position is attributed to a person who has not published it.
  • A measurement and a documented event answer different questions. Neither one refutes the other, and any argument that treats them as opposites is doing so for effect.
  • The scores describe specific models on specific test sets on a stated date. The cases describe specific organizations on specific dates. Neither generalizes to AI as a whole.
  • The free version of a model is usually not the one that was scored, so both columns can be true and still not describe the tool in front of you.
  • Everything here is secondary to reading the primary sources, all of which are free and linked from the scores page and the impacts page.