A leaderboard comparing more than 250 AI language models across intelligence, price, output speed, latency, and context window, with per-model provider analysis. Free to read.
Why I recommend it: The single table to open when someone claims one model is "the best." Sort by cost per task or speed and the answer often changes — a useful reality check in vendor conversations.
Head-to-head model comparisons voted on by the public: you see two anonymous answers to the same prompt and pick the better one, and the rankings come from those votes. Free.
From the site: Chat, compare, vote for the world's best AI models. Join the community shaping the public leaderboard for LLMs, image, and code models through real-world evaluation.
Why I recommend it: The closest thing to a fair fight between models on ordinary prompts, instead of marketing claims. Votes are taste as much as accuracy, so read it as popularity with a purpose.
Public leaderboard and open-source benchmark that drops AI agents into realistic business environments with 47 real tools across sales, marketing, operations, support, finance, and HR. Scores are based on final environment state, not an LLM-as-judge.
Why I recommend it: The leaderboard and the benchmark code are free; running it yourself means paying the model APIs at the costs shown. The test design is based on Zapier's own task data, so it's a realistic lens on agent work, but Zapier also sells automation tools — treat the benchmark as a useful public dataset, not a neutral referee.