Side by side
Compare two AI models
Pick two models and see what each costs, how well it works and what it's actually for. The scores are my own judgment on a 1–5 scale from using these — not vendor claims and not benchmark results. Your picks stay in the address bar, so you can send the comparison to someone else.
Choose a second model above to see them side by side.
Chat and everyday writing
Pi
Inflection
VisitFree value
Performance
Range of uses
Free and calming to talk to, but not the one for producing finished work.
- Published benchmark scores
- OpenUGI31.5
willingness and ability to answer awkward open-ended questions · tested: Inflection-3 Pi
Benchmark List · 6 May 2026 · independent
Not a general-capability test; no MMLU, GPQA or LMArena entry exists for Pi.
- Aggregate claim against GPT-4 (no per-test figures released)"over 94% of GPT-4's average performance", on 40% of the training compute
the maker's own summary across a set of benchmarks it chose · tested: Inflection-2.5, the model behind Pi
Inflection's launch announcement, reported by VentureBeat · 7 March 2024 · maker's own figures
An averaged claim, not a score: Inflection never published the per-benchmark table, its own page for the model is gone, and GPT-4 is two generations old. Read it as history.
Pi is barely benchmarked — Inflection stopped competing on leaderboards. The one current figure is from a niche test, so treat it as a curiosity rather than a ranking.
Figures copied from the sources above, checked 17 September 2026.
- Cost
- Free — there is no paid consumer plan to buy.
- Performance
- Conversational and patient rather than sharp — it asks questions back instead of producing walls of text.
- Use it for
- Talking through a career decision
- Rehearsing a difficult conversation
- Getting unstuck when you don't know what to ask
