Skip to content
Launchpad Library logo

Side by side

Compare two AI models

Pick two models and see what each costs, how well it works and what it's actually for. The scores are my own judgment on a 1–5 scale from using these — not vendor claims and not benchmark results. Your picks stay in the address bar, so you can send the comparison to someone else.

Choose a second model above to see them side by side.

Frontier flagship models

GPT-6.1 Sol

OpenAI

Visit
Published benchmark scores
  • DeepSWE v1.175.2%

    share of software-engineering tasks completed successfully · tested: high reasoning effort

    OpenAI · 29 September 2026 · maker's own figures

    OpenAI reports 74.1% for GPT-6 Astra in the same launch comparison.

  • OSWorld 2.0 (offline set)71.4%

    share of computer-use tasks completed successfully · tested: max reasoning effort

    OpenAI · 29 September 2026 · maker's own figures

    OpenAI reports 73.5% for GPT-6 Astra in the same launch comparison.

These launch figures come from OpenAI's own evaluations. Independent leaderboards may test different settings and production systems.

Figures copied from the sources above, checked 29 September 2026.

Cost
API $2 in / $10 out per million tokens; available in ChatGPT Work and Codex on eligible paid plans.
Performance
OpenAI's September 29 update to the Sol line, aimed at complex coding, computer use and professional workflows. OpenAI says it approaches GPT-6 Astra's capabilities while charging one-fifth of Astra's standard input and output token rates; the launch benchmark figures are the maker's own.
Use it for
  • Complex coding and debugging
  • Computer-use and multistep agent workflows
  • Long document and professional production tasks
Keep in mind
OpenAI classifies the model as Critical for cybersecurity and High for biological and chemical capability under its Preparedness Framework, and applies the same safeguard stack used for GPT-6 Astra.