Artificial Analysis benchmark write-up of Anthropic's Claude Sonnet 5.5: scores 56 on the Intelligence Index, two points behind Opus 5.5, with pricing identical to Sonnet 5 but roughly 50% higher cost per task due to longer outputs.
Epoch AI report on a spatial reasoning benchmark built around finding mistakes in photos of partially assembled IKEA furniture — a proxy for economically important visual reasoning tasks like repair work. Top scores jumped from 28% to 80% in 10 months; Chinese open-weight models trail the frontier by about 7 months.
A recommendation page from Artificial Analysis that suggests AI models for a given use case, drawing on the site's benchmark data across intelligence, speed, and cost. Free to use.
Why I recommend it: A quick way to turn "we need a model for X" into a shortlist with cost and speed tradeoffs attached. Useful in the first meeting; verify with your own tests before committing.
A running log of new model evaluations, benchmark methodology updates, and platform changes at Artificial Analysis, including Intelligence Index version changes and newly benchmarked models. Free to read.
Why I recommend it: The fastest way to see which models were benchmarked in the last week and when the scoring methodology changed. Bookmark it if you track the model landscape.
A platform from Artificial Analysis for building custom benchmarks from your own files, agent traces, or coding environment, then running them across leading models to compare quality, cost per task, and time per task. Benchmarks can be graded against objective rubrics or pairwise judging. Optima is a commercial product; the public announcement and product overview are free to read.
Why I recommend it: Standard benchmarks tell you which model is best in general; they cannot tell you which is best for your workload. If you are choosing a model for a real product, a custom benchmark on your own tasks is the right move — this is one way to do it without building the harness yourself.
A leaderboard comparing more than 250 AI language models across intelligence, price, output speed, latency, and context window, with per-model provider analysis. Free to read.
Why I recommend it: The single table to open when someone claims one model is "the best." Sort by cost per task or speed and the answer often changes — a useful reality check in vendor conversations.
A page listing benchmark scores (HLE, GPQA Diamond, MMLU-Pro, AI BENCHY) for a model called space-bunny-alpha, with confidence intervals.
Why I recommend it: Unverified: I couldn't find who runs this site or who made the model, and several scores are on subsets of the tests. Treat the numbers as unconfirmed until an independent lab reproduces them.
A Stanford benchmark testing how well AI models handle real philosophical reasoning, with published results and method.
Why I recommend it: A new academic benchmark. Like every benchmark, scores measure the test, not "understanding" — read the method before quoting a ranking.
Independent benchmarks comparing AI models and API providers on quality, speed, price and context length, with free charts and leaderboards.
Why I recommend it: Free to browse; the company sells API access to its data. Rankings depend on which tests they choose — compare with a second benchmark source before deciding.
Independent, domain-specific AI benchmarks in legal, tax and finance, testing how models actually perform on professional work rather than academic tests.
Why I recommend it: Free leaderboards; the company sells private benchmarking to enterprises. Domain-specific results are more useful than general scores if you work in these fields.
Benchmarks of open-source AI models by use case, with quality, cost and speed trade-offs.
Why I recommend it: Free to browse, but it is built by Together AI, which sells hosting for these same open models. Treat "where models run fastest" as a vendor showcase and cross-check with independent benchmarks.
Research and product writing from the team behind the Arena model leaderboards: how coding-agent harnesses change cost and success rates, how the agent leaderboards are built, and their academic partnership calls.
Why I recommend it: Free to read. The clearest writing anywhere on why two people using the same model get very different results — the tool wrapped around the model changes the cost and the outcome. They run the leaderboards they write about, so treat their rankings as one measurement, not the verdict.
Anthropic's own announcement of Claude Opus 5.5, published 22 September 2026 — the first model in the Claude 5.5 family. The company's claims, in its own words: it performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5 on typical workloads; on Anthropic's automated behavioural audit, its most comprehensive internal alignment test, it scores the highest of any model the company has tested. The page also cites early-tester anecdotes, including one completing a 680,000-line code migration in under a day, and a result of succeeding 39 times out of 40 on a page-load optimisation task. It is Anthropic's first release since the company publicly called for pacing the frontier, and it was tested before release by external evaluators including METR and Frontier Design.
Why I recommend it: Read this as a primary source, not as a review. Every performance and cost figure on the page was produced by the company selling the model, on tests it designed — that does not make them false, it makes them unchecked by anyone with a reason to doubt them. The checkable part is the external pre-release testing by METR and Frontier Design, so that is the part to weigh. Note too that '40% less to run than Opus 5' compares Anthropic with its own older model and says nothing about a rival's price, and that the standout numbers come from hand-picked early testers rather than a measured success rate you can plan around. The announcement is free to read; using the model itself is not free beyond whatever the current Claude free tier allows.
Raji and co-authors show that benchmarks claiming to measure general ability measure something much narrower, and that the gap is how overclaiming happens.
From the site: There is a tendency across different subfields in AI to valorize a small collection of influential benchmarks. These benchmarks operate as stand-ins for a range of anointed common problems that are frequently framed as foundational milestones on the path towards flexible and generalizable AI systems. State-of-the-art…
Why I recommend it: Read this before you trust a benchmark chart in a launch post. Free on arXiv.
An open-source framework from the UK's AI Security Institute for evaluating models — writing tests, scoring answers and logging what happened. Free.
From the site: Open-source framework for large language model evaluations
Why I recommend it: What a government safety institute actually uses to test models. Technical, but the docs explain the thinking behind each kind of test.
Research-grade tracking of what AI models can do and how that has changed over time, with the data and methods published. Free.
From the site: Our hub for benchmark results, featuring the performance of leading AI models on challenging tasks. It includes results from benchmarks administered internally by Epoch AI as well as data collected from external sources. Explore trends in AI capabilities across time, by benchmark, or by model.
Why I recommend it: For the longer view rather than this week's launch — they show their working, which most leaderboards do not.
Independent benchmarking of the major AI models on speed, price and quality, with the numbers side by side. Free to read.
From the site: Comparison and analysis of AI models and API hosting providers. Independent benchmarks across key performance metrics including quality, price, output speed & latency.
Why I recommend it: Where I check what a model actually costs per million words before believing a "cheap" claim.
Head-to-head model comparisons voted on by the public: you see two anonymous answers to the same prompt and pick the better one, and the rankings come from those votes. Free.
From the site: Chat, compare, vote for the world's best AI models. Join the community shaping the public leaderboard for LLMs, image, and code models through real-world evaluation.
Why I recommend it: The closest thing to a fair fight between models on ordinary prompts, instead of marketing claims. Votes are taste as much as accuracy, so read it as popularity with a purpose.
Independent, free benchmarks testing leading AI models on real-world finance, software, science and safety tasks, with cost and latency alongside accuracy.
From the site: Private, domain-specific benchmarks in legal, tax, and finance.
Why I recommend it: When someone claims a model is "the best," check here — these are independent evaluations, not vendor marketing.
A free, regularly updated leaderboard benchmarking how well leading AI models actually search the web, with the methodology and benchmarks published alongside.
Why I recommend it: Check this before assuming your favorite chatbot is the best one for research. The rankings move month to month.
A benchmark and tracker that documents reported instances of AI agents undertaking activity characterized as illegal, ranking major AI labs by aggregated incident counts.
Why I recommend it: This is exactly the kind of uncomfortable accountability tool our field needs. I include it because we cannot have thoughtful conversations about AI deployment without looking at real-world harm.
Hiring-funnel data showing roughly 6% of job views become applications, 3% of applicants reach an interview, and 27% of interviewees are hired - about one hire per 180 applicants, with tech roles needing far more applicants than healthcare.
Why I recommend it: Read this when the silence feels personal. It is not: 97% of applicants are screened out before a human ever sees them. That is the argument for referrals, sourcing, and warm outreach over another 40 cold applications.
In-depth analysis of frontier AI model releases, benchmarks, and capability claims. Hosted by Philip.
In plain terms: This YouTube channel provides detailed breakdowns of new artificial intelligence models and testing benchmarks. You can watch videos hosted by Philip to examine capability claims and learn how new tools perform.
Why I recommend it: The best check against hype when a new model drops.