Skip to content
Launchpad Library logo
← All publications
2026 in LLMs (So Far) — Simon Willison logo

2026 in LLMs (So Far) — Simon Willison

FreeAI Models & Benchmarkssimonwillison.net

Simon Willison's annotated slides and notes from his closing keynote at WeAreDevelopers World Congress North America (Sept. 2026): a chronological tour of the year's LLM developments, starting with the November 2025 models that made coding agents start working.

Read / subscribe free

From this publication

This is the only entry from 2026 in LLMs (So Far) — Simon Willison in the hub so far.

Related in the library

Other resources that share this publication's topics.

FreeTool
Technology & Ethics

Epoch Capabilities Index

Epoch AI's index combining many benchmarks into a single capability scale for comparing AI models over time, with a release-date graph view.

#ai-models#benchmarks
epoch.aiAdded Oct 2, 20260 opens
FreeTool
Technology & Ethics

Arena AI

A public leaderboard where visitors chat with, compare and vote on AI models, producing community rankings for language, image and code models based on real-world use.

#ai-models#benchmarks#leaderboard
arena.aiAdded Oct 2, 20260 opens
FreeAI Tool
AI Tools & Open Source

Visual Explainer (Agent Skill)

An open-source agent skill on GitHub by nicobailon that generates rich HTML pages or slide decks for diagrams, diff reviews, plan audits, data tables, and project recaps.

#agent-skills#open-source#coding-agents#github
github.comAdded Oct 2, 20260 opens
FreeArticle
AI Models & Benchmarks

Upstage Releases Solar Mini 4 (Artificial Analysis)

Artificial Analysis on Solar Mini 4, a proprietary reasoning model from Korean lab Upstage. It scores 24 on the Artificial Analysis Intelligence Index, with Upstage reporting 35B total and 3B active parameters, and is priced at $0.10/$0.40 per million input/output tokens — though it costs about 5x as much per task as GPT-6 Luna (max).

#ai-models#benchmarks#korea#llm
artificialanalysis.aiAdded Oct 2, 20260 opens
FreeWebsite
Technology & Ethics

OWASP Gen AI Security Project

Community-led OWASP project publishing free, open-source guidance on security and safety risks in generative AI applications, including the Top 10 for LLM applications.

#ai#security#ai safety#llm#owasp
genai.owasp.orgAdded Sep 29, 20260 opens
FreeResearch Paper
Research & Papers

SkillGym: Internalizing Human Skills into LLMs for Real-World Problem Solving

Authors
Zhilong Ge, Yuting Shao, Yutao Yang, Yuxuan Cai, Jie Zhou, et al.
Venue
Preprint, not yet peer-reviewed
Published
September 2026
ID
arXiv:2609.27717

arXiv paper (Sept. 2026) introducing SkillGym, a framework that turns human-written agent skills into executable, verifiable training environments so language model agents learn them as built-in capabilities.

#ai-agents#llm#research#training
arxiv.orgAdded Sep 28, 20260 opens