A recommendation page from Artificial Analysis that suggests AI models for a given use case, drawing on the site's benchmark data across intelligence, speed, and cost. Free to use.
Why I recommend it: A quick way to turn "we need a model for X" into a shortlist with cost and speed tradeoffs attached. Useful in the first meeting; verify with your own tests before committing.
A running log of new model evaluations, benchmark methodology updates, and platform changes at Artificial Analysis, including Intelligence Index version changes and newly benchmarked models. Free to read.
Why I recommend it: The fastest way to see which models were benchmarked in the last week and when the scoring methodology changed. Bookmark it if you track the model landscape.
A platform from Artificial Analysis for building custom benchmarks from your own files, agent traces, or coding environment, then running them across leading models to compare quality, cost per task, and time per task. Benchmarks can be graded against objective rubrics or pairwise judging. Optima is a commercial product; the public announcement and product overview are free to read.
Why I recommend it: Standard benchmarks tell you which model is best in general; they cannot tell you which is best for your workload. If you are choosing a model for a real product, a custom benchmark on your own tasks is the right move — this is one way to do it without building the harness yourself.
A leaderboard comparing more than 250 AI language models across intelligence, price, output speed, latency, and context window, with per-model provider analysis. Free to read.
Why I recommend it: The single table to open when someone claims one model is "the best." Sort by cost per task or speed and the answer often changes — a useful reality check in vendor conversations.
Benchmarks of open-source AI models by use case, with quality, cost and speed trade-offs.
Why I recommend it: Free to browse, but it is built by Together AI, which sells hosting for these same open models. Treat "where models run fastest" as a vendor showcase and cross-check with independent benchmarks.
Independent benchmarking of the major AI models on speed, price and quality, with the numbers side by side. Free to read.
From the site: Comparison and analysis of AI models and API hosting providers. Independent benchmarks across key performance metrics including quality, price, output speed & latency.
Why I recommend it: Where I check what a model actually costs per million words before believing a "cheap" claim.
Head-to-head model comparisons voted on by the public: you see two anonymous answers to the same prompt and pick the better one, and the rankings come from those votes. Free.
From the site: Chat, compare, vote for the world's best AI models. Join the community shaping the public leaderboard for LLMs, image, and code models through real-world evaluation.
Why I recommend it: The closest thing to a fair fight between models on ordinary prompts, instead of marketing claims. Votes are taste as much as accuracy, so read it as popularity with a purpose.
A free, regularly updated leaderboard benchmarking how well leading AI models actually search the web, with the methodology and benchmarks published alongside.
Why I recommend it: Check this before assuming your favorite chatbot is the best one for research. The rankings move month to month.