Research-grade tracking of what AI models can do and how that has changed over time, with the data and methods published. Free.
From the site: Our hub for benchmark results, featuring the performance of leading AI models on challenging tasks. It includes results from benchmarks administered internally by Epoch AI as well as data collected from external sources. Explore trends in AI capabilities across time, by benchmark, or by model.
Why I recommend it: For the longer view rather than this week's launch — they show their working, which most leaderboards do not.
Independent benchmarking of the major AI models on speed, price and quality, with the numbers side by side. Free to read.
From the site: Comparison and analysis of AI models and API hosting providers. Independent benchmarks across key performance metrics including quality, price, output speed & latency.
Why I recommend it: Where I check what a model actually costs per million words before believing a "cheap" claim.
Head-to-head model comparisons voted on by the public: you see two anonymous answers to the same prompt and pick the better one, and the rankings come from those votes. Free.
From the site: Chat, compare, vote for the world's best AI models. Join the community shaping the public leaderboard for LLMs, image, and code models through real-world evaluation.
Why I recommend it: The closest thing to a fair fight between models on ordinary prompts, instead of marketing claims. Votes are taste as much as accuracy, so read it as popularity with a purpose.
Independent, free benchmarks testing leading AI models on real-world finance, software, science and safety tasks, with cost and latency alongside accuracy.
From the site: Private, domain-specific benchmarks in legal, tax, and finance.
Why I recommend it: When someone claims a model is "the best," check here — these are independent evaluations, not vendor marketing.