Skip to content
Launchpad Library logo

benchmarks

26 free resources on this topic. Everything here is free and hand-picked. You can also search within this topic.

FreeArticle
Technology & Ethics

Claude Sonnet 5.5 Reaches #2 on the Artificial Analysis Intelligence Index

Artificial Analysis benchmark write-up of Anthropic's Claude Sonnet 5.5: scores 56 on the Intelligence Index, two points behind Opus 5.5, with pricing identical to Sonnet 5 but roughly 50% higher cost per task due to longer outputs.

#ai-models#benchmarks#anthropic#claude
artificialanalysis.aiAdded Sep 28, 20260 opens
FreeReport
Research & Papers

Can AI Spot Mistakes in IKEA Assembly?

Epoch AI report on a spatial reasoning benchmark built around finding mistakes in photos of partially assembled IKEA furniture — a proxy for economically important visual reasoning tasks like repair work. Top scores jumped from 28% to 80% in 10 months; Chinese open-weight models trail the frontier by about 7 months.

#ai-research#benchmarks#spatial-reasoning#epoch-ai
epoch.aiAdded Sep 28, 20260 opens
Artificial Analysis — Model Recommendations logoArtificial Analysis
FreeFinder Tool
AI & Assistive Tools

Artificial Analysis — Model Recommendations

A recommendation page from Artificial Analysis that suggests AI models for a given use case, drawing on the site's benchmark data across intelligence, speed, and cost. Free to use.

Why I recommend it: A quick way to turn "we need a model for X" into a shortlist with cost and speed tradeoffs attached. Useful in the first meeting; verify with your own tests before committing.

#ai#model-comparison#benchmarks#recommendation#llm#artificial-analysis
artificialanalysis.aiAdded Sep 27, 20260 opens
FreeWebsite
AI & Assistive Tools

Artificial Analysis — Changelog

A running log of new model evaluations, benchmark methodology updates, and platform changes at Artificial Analysis, including Intelligence Index version changes and newly benchmarked models. Free to read.

Why I recommend it: The fastest way to see which models were benchmarked in the last week and when the scoring methodology changed. Bookmark it if you track the model landscape.

#ai#benchmarks#model-comparison#changelog#new-models#artificial-analysis
artificialanalysis.aiAdded Sep 27, 20260 opens
FreeTool
AI & Assistive Tools

Optima — Custom AI Benchmark Builder (Artificial Analysis)

A platform from Artificial Analysis for building custom benchmarks from your own files, agent traces, or coding environment, then running them across leading models to compare quality, cost per task, and time per task. Benchmarks can be graded against objective rubrics or pairwise judging. Optima is a commercial product; the public announcement and product overview are free to read.

Why I recommend it: Standard benchmarks tell you which model is best in general; they cannot tell you which is best for your workload. If you are choosing a model for a real product, a custom benchmark on your own tasks is the right move — this is one way to do it without building the harness yourself.

#ai#benchmarks#evaluation#model-comparison#custom-benchmark#artificial-analysis#enterprise
artificialanalysis.aiAdded Sep 27, 20260 opens
FreeWebsite
AI & Assistive Tools

Artificial Analysis — LLM Leaderboard

A leaderboard comparing more than 250 AI language models across intelligence, price, output speed, latency, and context window, with per-model provider analysis. Free to read.

Why I recommend it: The single table to open when someone claims one model is "the best." Sort by cost per task or speed and the answer often changes — a useful reality check in vendor conversations.

#ai#benchmarks#leaderboard#model-comparison#llm#artificial-analysis
artificialanalysis.aiAdded Sep 27, 20260 opens
FreeWebsite
AI Models & Benchmarks

space-bunny-alpha Benchmarks

A page listing benchmark scores (HLE, GPQA Diamond, MMLU-Pro, AI BENCHY) for a model called space-bunny-alpha, with confidence intervals.

Why I recommend it: Unverified: I couldn't find who runs this site or who made the model, and several scores are on subsets of the tests. Treat the numbers as unconfirmed until an independent lab reproduces them.

#benchmarks#ai
spacebunnyalpha.comAdded Sep 26, 20266 opens
Freeresearch
Technology & Ethics

PhilosophyBench (Stanford)

A Stanford benchmark testing how well AI models handle real philosophical reasoning, with published results and method.

Why I recommend it: A new academic benchmark. Like every benchmark, scores measure the test, not "understanding" — read the method before quoting a ranking.

#research#benchmarks#ai ethics
philosophybench.orgAdded Sep 25, 20260 opens
Freenews
Technology & Ethics

New Study on AI's Philosophical Skills (Daily Nous)

Daily Nous, the philosophy profession's news site, on the PhilosophyBench study and philosophers' reactions to it.

Why I recommend it: Read the comments too — working philosophers push back on what the benchmark can show.

#research#benchmarks
dailynous.comAdded Sep 25, 20260 opens
FreeTool
AI Models & Benchmarks

Artificial Analysis

Independent benchmarks comparing AI models and API providers on quality, speed, price and context length, with free charts and leaderboards.

Why I recommend it: Free to browse; the company sells API access to its data. Rankings depend on which tests they choose — compare with a second benchmark source before deciding.

#benchmarks#models#comparison#research
artificialanalysis.aiAdded Sep 25, 20260 opens
FreeTool
AI Models & Benchmarks

Vals AI

Independent, domain-specific AI benchmarks in legal, tax and finance, testing how models actually perform on professional work rather than academic tests.

Why I recommend it: Free leaderboards; the company sells private benchmarking to enterprises. Domain-specific results are more useful than general scores if you work in these fields.

#benchmarks#legal#finance#research
vals.aiAdded Sep 25, 20260 opens
FreeArticle
AI & Assistive Tools

Grok 4.7 announcement (xAI)

xAI's own announcement of its Grok 4.7 model, with the company's claimed benchmark results.

Why I recommend it: Company announcement: the scores are xAI's own and haven't been checked independently. Some Grok features need a paid plan.

#AI models#benchmarks#Grok#research#xAI
x.aiAdded Sep 24, 20260 opens
FreeWebsite
AI & Assistive Tools

The Open Frontier

Benchmarks of open-source AI models by use case, with quality, cost and speed trade-offs.

Why I recommend it: Free to browse, but it is built by Together AI, which sells hosting for these same open models. Treat "where models run fastest" as a vendor showcase and cross-check with independent benchmarks.

#benchmarks#model-comparison#open-source-models#research
theopenfrontier.comAdded Sep 24, 20260 opens
FreePublication
AI & Assistive Tools

Arena Blog (LLM and Agent Evaluation Research)

Research and product writing from the team behind the Arena model leaderboards: how coding-agent harnesses change cost and success rates, how the agent leaderboards are built, and their academic partnership calls.

Why I recommend it: Free to read. The clearest writing anywhere on why two people using the same model get very different results — the tool wrapped around the model changes the cost and the outcome. They run the leaderboards they write about, so treat their rankings as one measurement, not the verdict.

#ai-evaluation#benchmarks#llm#ai-agents#research#free#ai-and-assistive-tools
arena.aiAdded Sep 22, 20260 opens
FreeArticle
AI & Assistive Tools

Introducing Claude Opus 5.5 (Anthropic)

Anthropic's own announcement of Claude Opus 5.5, published 22 September 2026 — the first model in the Claude 5.5 family. The company's claims, in its own words: it performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5 on typical workloads; on Anthropic's automated behavioural audit, its most comprehensive internal alignment test, it scores the highest of any model the company has tested. The page also cites early-tester anecdotes, including one completing a 680,000-line code migration in under a day, and a result of succeeding 39 times out of 40 on a page-load optimisation task. It is Anthropic's first release since the company publicly called for pacing the frontier, and it was tested before release by external evaluators including METR and Frontier Design.

Why I recommend it: Read this as a primary source, not as a review. Every performance and cost figure on the page was produced by the company selling the model, on tests it designed — that does not make them false, it makes them unchecked by anyone with a reason to doubt them. The checkable part is the external pre-release testing by METR and Frontier Design, so that is the part to weigh. Note too that '40% less to run than Opus 5' compares Anthropic with its own older model and says nothing about a rival's price, and that the standout numbers come from hand-picked early testers rather than a measured success rate you can plan around. The announcement is free to read; using the model itself is not free beyond whatever the current Claude free tier allows.

#agentic-coding#ai ethics#ai-models#ai-safety#anthropic#benchmarks#claude#model-release#research
anthropic.comAdded Sep 22, 20260 opens
FreeResearch Paper
Technology & Ethics

AI and the Everything in the Whole Wide World Benchmark

Raji and co-authors show that benchmarks claiming to measure general ability measure something much narrower, and that the gap is how overclaiming happens.

From the site: There is a tendency across different subfields in AI to valorize a small collection of influential benchmarks. These benchmarks operate as stand-ins for a range of anointed common problems that are frequently framed as foundational milestones on the path towards flexible and generalizable AI systems. State-of-the-art…

Why I recommend it: Read this before you trust a benchmark chart in a launch post. Free on arXiv.

#ai ethics#ai-ethics#benchmarks#free#reading#research#technology-and-ethics
arXiv.orgAdded Sep 19, 20260 opens
FreeTool
AI & Assistive Tools

Inspect

An open-source framework from the UK's AI Security Institute for evaluating models — writing tests, scoring answers and logging what happened. Free.

From the site: Open-source framework for large language model evaluations

Why I recommend it: What a government safety institute actually uses to test models. Technical, but the docs explain the thinking behind each kind of test.

#ai ethics#ai-risk#ai-safety#alignment#benchmarks#developer-tools#evaluation#free#open-source#research#technology-and-ethics#testing
InspectAdded Sep 17, 20260 opens
FreeTool
AI & Assistive Tools

Epoch AI Benchmarking Hub

Research-grade tracking of what AI models can do and how that has changed over time, with the data and methods published. Free.

From the site: Our hub for benchmark results, featuring the performance of leading AI models on challenging tasks. It includes results from benchmarks administered internally by Epoch AI as well as data collected from external sources. Explore trends in AI capabilities across time, by benchmark, or by model.

Why I recommend it: For the longer view rather than this week's launch — they show their working, which most leaderboards do not.

#benchmarks#research#ai-models#free#data#independent#evaluation#ai-trends#transparency#ai-tools
Epoch AIAdded Sep 17, 20260 opens
FreeTool
AI & Assistive Tools

Artificial Analysis

Independent benchmarking of the major AI models on speed, price and quality, with the numbers side by side. Free to read.

From the site: Comparison and analysis of AI models and API hosting providers. Independent benchmarks across key performance metrics including quality, price, output speed & latency.

Why I recommend it: Where I check what a model actually costs per million words before believing a "cheap" claim.

#benchmarks#model-comparison#pricing#ai-tools#free#independent#research#evaluation#ai-models#cost-comparison
artificialanalysis.aiAdded Sep 17, 20260 opens
FreeTool
AI & Assistive Tools

LMArena

Head-to-head model comparisons voted on by the public: you see two anonymous answers to the same prompt and pick the better one, and the rankings come from those votes. Free.

From the site: Chat, compare, vote for the world's best AI models. Join the community shaping the public leaderboard for LLMs, image, and code models through real-world evaluation.

Why I recommend it: The closest thing to a fair fight between models on ordinary prompts, instead of marketing claims. Votes are taste as much as accuracy, so read it as popularity with a purpose.

#benchmarks#model-comparison#ai-tools#free#research#evaluation#leaderboard#ai-models#independent#decision-making
Arena AI: The Official AI Ranking & LLM LeaderboardAdded Sep 17, 20260 opens
FreeWebsite
AI & Assistive Tools

Vals AI Model Benchmarks

Independent, free benchmarks testing leading AI models on real-world finance, software, science and safety tasks, with cost and latency alongside accuracy.

From the site: Private, domain-specific benchmarks in legal, tax, and finance.

Why I recommend it: When someone claims a model is "the best," check here — these are independent evaluations, not vendor marketing.

#ai#ai ethics#ai-safety#benchmarks#comparison#cybersecurity#independent#llm#model-evaluation#primary-source#research
vals.aiAdded Sep 17, 20260 opens
FreeWebsite
AI & Assistive Tools

Parallel Search Capability Leaderboard

A free, regularly updated leaderboard benchmarking how well leading AI models actually search the web, with the methodology and benchmarks published alongside.

Why I recommend it: Check this before assuming your favorite chatbot is the best one for research. The rankings move month to month.

#ai#benchmarks#ai-agents#search#model-comparison#research#evaluation#data#free-tools#ai-tools
parallel.aiAdded Sep 17, 20260 opens
FreeFellowship
Career & Professional Optimization

Vals AI Fellowship

A fellowship with Vals AI, the team that benchmarks how AI models actually perform, for people who want to work on model evaluation.

Why I recommend it: Evaluation work is where a lot of AI hiring is heading — careful, skeptical people do well here even without a PhD.

#ai#ai-jobs#application#benchmarks#early-career#entry-level#evaluation#fellowship#free#job markets#research
app.notion.comAdded Sep 16, 20260 opens
FreeReport
Technology & Ethics

Felony Bench

A benchmark and tracker that documents reported instances of AI agents undertaking activity characterized as illegal, ranking major AI labs by aggregated incident counts.

Why I recommend it: This is exactly the kind of uncomfortable accountability tool our field needs. I include it because we cannot have thoughtful conversations about AI deployment without looking at real-world harm.

#accountability#aggregator#ai#ai ethics#ai-ethics#ai-labs#ai-safety#benchmarks#ethics#free#governance#illegal-activity#regulation#research#risk#tech-ethics#technology#transparency
felonybench.comAdded Sep 9, 20260 opens
FreeArticle
Career & Professional Optimization

Recruitment Funnel Benchmarks: Conversion Rates by Stage

Hiring-funnel data showing roughly 6% of job views become applications, 3% of applicants reach an interview, and 27% of interviewees are hired - about one hire per 180 applicants, with tech roles needing far more applicants than healthcare.

Why I recommend it: Read this when the silence feels personal. It is not: 97% of applicants are screened out before a human ever sees them. That is the argument for referrals, sourcing, and warm outreach over another 40 cold applications.

#applicant-tracking#article#benchmarks#career#career-development#free#hiring-data#job markets#job-search#reading#research
pin.comAdded Sep 4, 20260 opens
FreePerson to FollowVideo
AI & Assistive Tools

AI Explained (YouTube)

In-depth analysis of frontier AI model releases, benchmarks, and capability claims. Hosted by Philip.

In plain terms: This YouTube channel provides detailed breakdowns of new artificial intelligence models and testing benchmarks. You can watch videos hosted by Philip to examine capability claims and learn how new tools perform.

Why I recommend it: The best check against hype when a new model drops.

#ai#ai-news#ai-tools#analysis#benchmarks#expert#free#person-to-follow#research#video#youtube
youtube.comAdded Aug 29, 20260 opens