Independent benchmarks comparing AI models and API providers on quality, speed, price and context length, with free charts and leaderboards.
Why I recommend it: Free to browse; the company sells API access to its data. Rankings depend on which tests they choose — compare with a second benchmark source before deciding.
Independent, free benchmarks testing leading AI models on real-world finance, software, science and safety tasks, with cost and latency alongside accuracy.
From the site: Private, domain-specific benchmarks in legal, tax, and finance.
Why I recommend it: When someone claims a model is "the best," check here — these are independent evaluations, not vendor marketing.
XDA Developers had Claude Code (Opus 5), Codex (GPT-5.6) and Google Antigravity (Gemini 3.8) rebuild the same website from the same brief. The comparison shows which agent handles detail, polish and real-world edge cases best.
Why I recommend it: Useful if you are choosing an AI coding assistant for side projects or learning to prompt more effectively. The winner is not necessarily the one you would expect.