Epoch Capabilities Index
Epoch AI's index combining many benchmarks into a single capability scale for comparing AI models over time, with a release-date graph view.
Simon Willison's annotated slides and notes from his closing keynote at WeAreDevelopers World Congress North America (Sept. 2026): a chronological tour of the year's LLM developments, starting with the November 2025 models that made coding agents start working.
Read / subscribe freeThis is the only entry from 2026 in LLMs (So Far) — Simon Willison in the hub so far.
Other resources that share this publication's topics.
Epoch AI's index combining many benchmarks into a single capability scale for comparing AI models over time, with a release-date graph view.
A public leaderboard where visitors chat with, compare and vote on AI models, producing community rankings for language, image and code models based on real-world use.
An open-source agent skill on GitHub by nicobailon that generates rich HTML pages or slide decks for diagrams, diff reviews, plan audits, data tables, and project recaps.
Artificial Analysis on Solar Mini 4, a proprietary reasoning model from Korean lab Upstage. It scores 24 on the Artificial Analysis Intelligence Index, with Upstage reporting 35B total and 3B active parameters, and is priced at $0.10/$0.40 per million input/output tokens — though it costs about 5x as much per task as GPT-6 Luna (max).
Community-led OWASP project publishing free, open-source guidance on security and safety risks in generative AI applications, including the Top 10 for LLM applications.
arXiv paper (Sept. 2026) introducing SkillGym, a framework that turns human-written agent skills into executable, verifiable training environments so language model agents learn them as built-in capabilities.