Artificial Analysis benchmark write-up of Anthropic's Claude Sonnet 5.5: scores 56 on the Intelligence Index, two points behind Opus 5.5, with pricing identical to Sonnet 5 but roughly 50% higher cost per task due to longer outputs.
36Kr report on Google DeepMind's Gemini 4 Pro, covering early hands-on tests and comparisons with Claude Opus 5.5 across coding, reasoning and multimodal tasks.
The Decoder report on comments from Boris Power, OpenAI's Head of Applied Research, that most of the company's research targets future model generations, and that user awareness of AI capabilities lags model performance.
TechCrunch reports on frontier AI models breaking surviving World War II Enigma messages, continuing codebreaking work of the kind Alan Turing did at Bletchley Park. Results described are as reported by the people who ran them.
Anthropic's own announcement of Claude Opus 5.5, published 22 September 2026 — the first model in the Claude 5.5 family. The company's claims, in its own words: it performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5 on typical workloads; on Anthropic's automated behavioural audit, its most comprehensive internal alignment test, it scores the highest of any model the company has tested. The page also cites early-tester anecdotes, including one completing a 680,000-line code migration in under a day, and a result of succeeding 39 times out of 40 on a page-load optimisation task. It is Anthropic's first release since the company publicly called for pacing the frontier, and it was tested before release by external evaluators including METR and Frontier Design.
Why I recommend it: Read this as a primary source, not as a review. Every performance and cost figure on the page was produced by the company selling the model, on tests it designed — that does not make them false, it makes them unchecked by anyone with a reason to doubt them. The checkable part is the external pre-release testing by METR and Frontier Design, so that is the part to weigh. Note too that '40% less to run than Opus 5' compares Anthropic with its own older model and says nothing about a rival's price, and that the standout numbers come from hand-picked early testers rather than a measured success rate you can plan around. The announcement is free to read; using the model itself is not free beyond whatever the current Claude free tier allows.
DeepSeek's chat assistant. Free tier gives unlimited messages and file uploads on a daily-reset quota; no payment required. One of the most significant free AI releases of the past two years.
Why I recommend it: Worth trying on reasoning-heavy work and long documents. Note it is a Chinese service — read its data terms before pasting anything confidential.
Research-grade tracking of what AI models can do and how that has changed over time, with the data and methods published. Free.
From the site: Our hub for benchmark results, featuring the performance of leading AI models on challenging tasks. It includes results from benchmarks administered internally by Epoch AI as well as data collected from external sources. Explore trends in AI capabilities across time, by benchmark, or by model.
Why I recommend it: For the longer view rather than this week's launch — they show their working, which most leaderboards do not.
Independent benchmarking of the major AI models on speed, price and quality, with the numbers side by side. Free to read.
From the site: Comparison and analysis of AI models and API hosting providers. Independent benchmarks across key performance metrics including quality, price, output speed & latency.
Why I recommend it: Where I check what a model actually costs per million words before believing a "cheap" claim.
Head-to-head model comparisons voted on by the public: you see two anonymous answers to the same prompt and pick the better one, and the rankings come from those votes. Free.
From the site: Chat, compare, vote for the world's best AI models. Join the community shaping the public leaderboard for LLMs, image, and code models through real-world evaluation.
Why I recommend it: The closest thing to a fair fight between models on ordinary prompts, instead of marketing claims. Votes are taste as much as accuracy, so read it as popularity with a purpose.