Side by side
Compare two AI models
Pick two models and see what each costs, how well it works and what it's actually for. The scores are my own judgment on a 1–5 scale from using these — not vendor claims and not benchmark results. Your picks stay in the address bar, so you can send the comparison to someone else.
Example comparison: ChatGPT vs Perplexity. Pick your own two models below, or clear this example to start fresh.
Published benchmark scores
Real measured figures, copied from the sources named in each cell and checked 17 September 2026. A blank means nobody has run that test on it — not a low score.
| Benchmark | ChatGPT | Perplexity |
|---|---|---|
| Terminal-Bench 4.0share of real command-line tasks completed unaided | 58.2% (±2.8)tested: GPT-6 Astra, max reasoning — ranked 1stTerminal-Bench / Snorkel29 August 2026 · independent | Not tested |
| Artificial Analysis Intelligence Index v4.3.2combined score across ten reasoning and knowledge tests | 53tested: GPT-6 Astra, max reasoning — ranked 7th of 216 modelsArtificial Analysis28 September 2026 · independentMeasured at 63.6 tokens per second output speed. | Not tested |
| LMArena Elo (style-controlled)blind head-to-head preference votes from the public | 1,486tested: GPT-5.6 Sol, extra-high reasoningModel Gauntlet (LMArena data)20 July 2026 · independent | Not tested |
| SimpleQA (F-score)factual accuracy on short questions with one checkable answer | Not tested | 85.8tested: Perplexity Sonar ProPerplexity's own Sonar Pro announcement21 January 2025 · maker's own figuresThe maker's own figure. It is the number they lead with, and no independent tester has published a comparable SimpleQA run for Sonar Pro. |
| MMLUbroad multiple-choice general knowledge | Not tested | |
| HumanEvalsmall Python programming problems solved correctly | Not tested | |
| GPQAgraduate-level science questions | Not tested | |
| SciCodescientific coding problems solved correctly | Not tested |
ChatGPT: OpenAI doesn't publish comparable per-model figures on a single page, so these come from independent testers. They test the top reasoning setting, which is not what the free tier gives you.
Perplexity: Perplexity benchmarks its Sonar models rather than the search product you use in the browser, so read these as a floor, not a ceiling. The headline factuality figure is Perplexity's own, from the Sonar Pro launch post.
| Score | ChatGPT | Perplexity |
|---|---|---|
| Free valueHow much you get without paying. | 4/5 | 4/5 |
| PerformanceHow well it does its main job. | 5/5 | 5/5 |
| Range of usesHow many different jobs it suits. | 5/5 | 4/5 |
Chat and everyday writing
ChatGPT
OpenAI
VisitFree value
Performance
Range of uses
The safest first choice for writing work; the free tier is limited but real.
- Published benchmark scores
- Terminal-Bench 4.058.2% (±2.8)
share of real command-line tasks completed unaided · tested: GPT-6 Astra, max reasoning — ranked 1st
Terminal-Bench / Snorkel · 29 August 2026 · independent
- Artificial Analysis Intelligence Index v4.3.253
combined score across ten reasoning and knowledge tests · tested: GPT-6 Astra, max reasoning — ranked 7th of 216 models
Artificial Analysis · 28 September 2026 · independent
Measured at 63.6 tokens per second output speed.
- LMArena Elo (style-controlled)1,486
blind head-to-head preference votes from the public · tested: GPT-5.6 Sol, extra-high reasoning
Model Gauntlet (LMArena data) · 20 July 2026 · independent
OpenAI doesn't publish comparable per-model figures on a single page, so these come from independent testers. They test the top reasoning setting, which is not what the free tier gives you.
Figures copied from the sources above, checked 17 September 2026.
- Cost
- Free plan; Plus $20/month, Pro from $100/month.
- Performance
- The most reliable all-rounder for writing and rewriting. Strong at following a long, fussy instruction; occasionally confident about things it made up.
- Use it for
- Tailoring a resume to a specific posting
- Turning rough notes into clean prose
- Practicing interview answers out loud
- Keep in mind
- Check any fact, name, date or number before you send it to an employer.
Search, answers and research
Perplexity
Perplexity AI
VisitFree value
Performance
Range of uses
Cites its sources, which makes it the one to use for anything factual.
- Published benchmark scores
- SimpleQA (F-score)85.8
factual accuracy on short questions with one checkable answer · tested: Perplexity Sonar Pro
Perplexity's own Sonar Pro announcement · 21 January 2025 · maker's own figures
The maker's own figure. It is the number they lead with, and no independent tester has published a comparable SimpleQA run for Sonar Pro.
- MMLU80.1%
broad multiple-choice general knowledge · tested: Perplexity Sonar
TPS Report · model released 27 January 2025 · independent
- HumanEval72.5%
small Python programming problems solved correctly · tested: Perplexity Sonar
TPS Report · model released 27 January 2025 · independent
- GPQA48.0%
graduate-level science questions · tested: Perplexity Sonar
TPS Report · model released 27 January 2025 · independent
- SciCode22.9%
scientific coding problems solved correctly · tested: Perplexity Sonar
Benchmark List · 21 July 2026 · independent
Perplexity benchmarks its Sonar models rather than the search product you use in the browser, so read these as a floor, not a ceiling. The headline factuality figure is Perplexity's own, from the Sonar Pro launch post.
Figures copied from the sources above, checked 17 September 2026.
- Cost
- Free plan; Pro $20/month, Max $200/month.
- Performance
- Cites its sources on every answer, which makes it far safer than a plain chatbot for anything factual.
- Use it for
- Researching an employer before an interview
- Checking a salary range
- Finding the original source of a claim
- Keep in mind
- Follow the citations. It sometimes summarizes a source more confidently than the source does.
