Side by side
Compare two AI models
Pick two models and see what each costs, how well it works and what it's actually for. The scores are my own judgment on a 1–5 scale from using these — not vendor claims and not benchmark results. Your picks stay in the address bar, so you can send the comparison to someone else.
Choose a second model above to see them side by side.
Frontier flagship models
GPT-6.1 Sol
OpenAI
Visit- Published benchmark scores
- DeepSWE v1.175.2%
share of software-engineering tasks completed successfully · tested: high reasoning effort
OpenAI · 29 September 2026 · maker's own figures
OpenAI reports 74.1% for GPT-6 Astra in the same launch comparison.
- OSWorld 2.0 (offline set)71.4%
share of computer-use tasks completed successfully · tested: max reasoning effort
OpenAI · 29 September 2026 · maker's own figures
OpenAI reports 73.5% for GPT-6 Astra in the same launch comparison.
These launch figures come from OpenAI's own evaluations. Independent leaderboards may test different settings and production systems.
Figures copied from the sources above, checked 29 September 2026.
- Cost
- API $2 in / $10 out per million tokens; available in ChatGPT Work and Codex on eligible paid plans.
- Performance
- OpenAI's September 29 update to the Sol line, aimed at complex coding, computer use and professional workflows. OpenAI says it approaches GPT-6 Astra's capabilities while charging one-fifth of Astra's standard input and output token rates; the launch benchmark figures are the maker's own.
- Use it for
- Complex coding and debugging
- Computer-use and multistep agent workflows
- Long document and professional production tasks
- Keep in mind
- OpenAI classifies the model as Critical for cybersecurity and High for biological and chemical capability under its Preparedness Framework, and applies the same safeguard stack used for GPT-6 Astra.
