Skip to content
Launchpad Library logo

AI models, compared honestly

The 18 models and model-based assistants in the library, grouped by what they're for — with what each one really costs, how it performs in plain language, and what I'd actually use it for.

Compare two models side by side

Want to see who's actually building these models? See the Frontier AI Labs directory.

Where these numbers come from

Every price below is the figure the maker publishes on its own pricing page, with a link so you can check it. Each model shows the date its own price was last checked. The benchmark figures are real published scores, each one carrying the test it comes from, who ran it and the date, and marked when the number comes from the company that makes the model — checked 17 September 2026. Where nobody publishes a score, or the entry is a platform rather than a model, it says so instead of guessing. The plain-language performance note is still mine, written from using these tools.

Chat and everyday writing

4 options

General assistants you type at. Best for drafting, rewriting, summarizing and thinking out loud.

  • ChatGPT

    OpenAI
    Published price
    • FreeUnlimited everyday chats, limited uploads, images, voice and deep research$0/month
    • GoMay include ads. OpenAI leaves the figure off its own page and shows it at checkout; $8 is the figure reported by CloudZero's price tracking, September 2026$8/month
    • Plus$20/month
    • Pro$100/month, or $200/month for the higher usage tier
    • Business$25/seat billed monthly, two-seat minimum$20/seat/month billed annually

    OpenAI localizes and A/B-tests these figures and does not print the Go price on the page, so confirm at checkout before paying.

    Source: OpenAI's pricing page · Last checked 17 September 2026

    Published benchmark scores
    • Terminal-Bench 4.058.2% (±2.8)

      share of real command-line tasks completed unaided · tested: GPT-6 Astra, max reasoning — ranked 1st

      Terminal-Bench / Snorkel · 29 August 2026 · independent

    • Artificial Analysis Intelligence Index v4.3.253

      combined score across ten reasoning and knowledge tests · tested: GPT-6 Astra, max reasoning — ranked 7th of 216 models

      Artificial Analysis · 28 September 2026 · independent

      Measured at 63.6 tokens per second output speed.

    OpenAI doesn't publish comparable per-model figures on a single page, so these come from independent testers. They test the top reasoning setting, which is not what the free tier gives you.

    Figures copied from the sources above, checked 17 September 2026.

    1 more published figure on its own page.

    Performance, in my words
    The most reliable all-rounder for writing and rewriting. Strong at following a long, fussy instruction; occasionally confident about things it made up.
    Prompting it from this site
    You can run this hereThree settings are wired up — quick, middle and thinking — and they run on this site's own AI credits, not your ChatGPT plan.
    Use it for
    • Tailoring a resume to a specific posting
    • Turning rough notes into clean prose
    • Practicing interview answers out loud

    Check any fact, name, date or number before you send it to an employer.

  • Kimi

    Moonshot AI
    Published price
    • FreeNormal browser use, no subscription$0
    • Moderato$15/month billed annually$19/month
    • Allegretto$31/month billed annually$39/month
    • Allegro$79/month billed annually$99/month
    • Vivace$159/month billed annually$199/month
    • API — Kimi K3$3.00 in / $15.00 out per million tokens

    Source: Moonshot's membership pricing · Last checked 17 September 2026

    Published benchmark scores
    • Artificial Analysis Intelligence Index v4.3.244

      combined score across ten reasoning and knowledge tests · tested: Kimi K3

      Artificial Analysis · 28 September 2026 · independent

    Figures are for Kimi K3, the current version on Artificial Analysis. No output speed was published for this model, so none is shown.

    Figures copied from the sources above, checked 17 September 2026.

    Performance, in my words
    Handles very long documents better than most free assistants — you can paste a whole report and ask about it.
    Prompting it from this site
    Endpoint wired, waiting on a keyKimi K2.5 runs through OpenRouter. Once an OpenRouter key is saved for this site the choice switches on; until then it's grayed out rather than failing when clicked.
    Use it for
    • Summarizing long PDFs and reports
    • Comparing two long job descriptions
    • Pulling the requirements out of a dense posting
  • Hermes Agent

    Hermes
    Published price
    • SoftwareMIT license, run it yourself$0
    • Model useNo fee to HermesWhatever your own API key costs
    • Managed hosting — 7-day trialIncludes built-in model access, 500 conversations every 5 hours$12 one-off
    • Managed hosting — monthlyThird-party hosts run this; Hostinger's managed plan starts at $5.99/month and renews at $11.99/monthFrom about $6/month
    • Your own server instead$4–$25/month VPS plus your model spend

    The project itself sells nothing; the hosted figures come from the third-party hosts that package it, and the official hosting page prints only the trial price.

    Source: the Hermes Agent site · Last checked 17 September 2026

    Published benchmark scores

    The agent itself has never been benchmarked. These are scores for Hermes 4 405B, the model it runs on, and the headline numbers come from Nous Research's own technical report.

    Figures copied from the sources above, checked 17 September 2026.

    1 more published figure on its own page.

    Performance, in my words
    An open agent you point at a task rather than a chatbot; capable but rougher around the edges than the commercial assistants.
    Prompting it from this site
    Can't be run from hereIt's an agent that works across a whole project, not a single prompt. Run it on your own machine — the setup steps on its page show how.
    Use it for
    • Automating a repeated multi-step task
    • Learning how agents actually work
  • Pi

    Inflection
    Published price
    • Pi (web, iOS, Android, WhatsApp)No subscription, no premium tier and no in-app purchase; rate limits apply to stop abuse$0
    • Inflection for EnterpriseBusiness deployments only — no list price anywhereQuoted by their sales team

    Inflection publishes no consumer price list because the consumer product is entirely free; the company sells to businesses instead.

    Source: Inflection's Pi site · Last checked 17 September 2026

    Published benchmark scores
    • OpenUGI31.5

      willingness and ability to answer awkward open-ended questions · tested: Inflection-3 Pi

      Benchmark List · 6 May 2026 · independent

      Not a general-capability test; no MMLU, GPQA or LMArena entry exists for Pi.

    • Aggregate claim against GPT-4 (no per-test figures released)"over 94% of GPT-4's average performance", on 40% of the training compute

      the maker's own summary across a set of benchmarks it chose · tested: Inflection-2.5, the model behind Pi

      Inflection's launch announcement, reported by VentureBeat · 7 March 2024 · maker's own figures

      An averaged claim, not a score: Inflection never published the per-benchmark table, its own page for the model is gone, and GPT-4 is two generations old. Read it as history.

    Pi is barely benchmarked — Inflection stopped competing on leaderboards. The one current figure is from a niche test, so treat it as a curiosity rather than a ranking.

    Figures copied from the sources above, checked 17 September 2026.

    Performance, in my words
    Conversational and patient rather than sharp — it asks questions back instead of producing walls of text.
    Prompting it from this site
    Can't be run from hereInflection publishes no API for Pi, so there is nothing to send a prompt to.
    Use it for
    • Talking through a career decision
    • Rehearsing a difficult conversation
    • Getting unstuck when you don't know what to ask

Frontier flagship models

3 options

The flagship models from the largest labs, built for demanding reasoning, coding and agent work, and reached through paid plans and APIs rather than free consumer apps.

  • Claude Opus 5.5

    Anthropic
    Published price
    • API — input$4 per 1M tokens
    • API — output$20 per 1M tokens
    • API — cache writes (5 min / 1 hour)$5 / $8 per 1M tokens
    • API — cache hits$0.20 per 1M tokens
    • Claude FreeChat on web, desktop and mobile; usage-window limits apply$0
    • Claude Pro$17/month (annual) or $20/month
    • Claude MaxFrom $100/month

    Consumer plan prices are for Claude.ai overall; model access on each plan follows usage limits rather than a fixed assignment.

    Source: Anthropic's pricing page · Last checked 27 September 2026

    Published benchmark scores
    • Artificial Analysis Intelligence Index v4.3.258 — 1st of 216 models

      combined score across ten reasoning and knowledge tests · tested: max with fallback

      Artificial Analysis · 28 September 2026 · independent

      Measured at 95.2 tokens per second output speed.

    Figures copied from the sources above, checked 17 September 2026.

    Performance, in my words
    Anthropic's model for long-running agentic coding and knowledge work, and the first release since the lab called for pacing the frontier. Its launch post reports it matches Claude Fable 5.1 on most work at 40% lower running cost than Opus 5; benchmark figures are the maker's own, so check independent leaderboards for a specific task.
    Prompting it from this site
    Can't be run from hereThis site has no Claude wired up — the models it prompts run through OpenAI's and Google's endpoints. Use Claude.ai or Anthropic's API instead.
    Use it for
    • Long-running coding and migration tasks
    • Knowledge-work and document reasoning
    • Agentic workflows through Claude paid plans

    Heavier Opus 5.5 access needs a paid Claude plan; the free plan applies a rolling usage window.

  • Claude Sonnet 5.5

    Anthropic
    Published price
    • API — input$2 per 1M tokens
    • API — output$10 per 1M tokens
    • API — cache reads$0.20 per 1M tokens
    • Claude FreeChat on web, desktop and mobile; usage-window limits apply$0
    • Claude Pro$17/month (annual) or $20/month
    • Claude MaxFrom $100/month

    API prices come from the launch announcement; consumer plan prices are for Claude.ai overall, and model access on each plan follows usage limits rather than a fixed assignment.

    Source: Anthropic's launch announcement · Last checked 28 September 2026

    Published benchmark scores
    • Artificial Analysis Intelligence Index v4.3.256 — 3rd of 216 models

      combined score across ten reasoning and knowledge tests · tested: max with fallback

      Artificial Analysis · 28 September 2026 · independent

      Measured at 141.9 tokens per second output speed.

    Figures copied from the sources above, checked 17 September 2026.

    Performance, in my words
    Anthropic's faster, lower-cost complement to Opus 5.5, announced 28 September 2026, aimed at well-scoped everyday tasks — fixing bugs and producing documents, slides and spreadsheets. The launch post reports 70.6% on Terminal-Bench 4.0 agentic coding (up from 10.3% for Sonnet 5), two points below Opus 5.5 on GDPval-AA, and outputs generated 30%+ faster than Sonnet 5; benchmark figures are the maker's own, so check independent leaderboards for a specific task.
    Prompting it from this site
    Can't be run from hereThis site has no Claude wired up — the models it prompts run through OpenAI's and Google's endpoints. Use Claude.ai or Anthropic's API instead.
    Use it for
    • Everyday coding and bug fixes
    • Polished documents, slides and spreadsheets
    • High-volume work that doesn't need Opus 5.5's sustained judgment

    The first Sonnet model to launch with cyber safeguards like those on the lab's most capable models, per the announcement; heavier use on Claude plans follows usage-window limits.

  • GPT-6 Astra

    OpenAI
    Published price
    • Standard (≤272K input)$10 in / $50 out per 1M tokens
    • Cached input$1.00 per 1M tokens
    • Long context (>272K input)The full request is priced at the long-context rate$20 in / $75 out per 1M tokens
    • Batch & Flex50% of standard rates
    • Fast mode2× standard rates
    • In ChatGPTAvailable to Plus, Pro, Business and Enterprise users, per OpenAI's announcementPlus $20, Pro $100/month

    Source: OpenAI's API pricing page · Last checked 27 September 2026

    Published benchmark scores
    • Artificial Analysis Intelligence Index v4.3.253 — 7th of 216 models

      combined score across ten reasoning and knowledge tests · tested: max reasoning

      Artificial Analysis · 28 September 2026 · independent

      Measured at 63.6 tokens per second output speed.

    Figures copied from the sources above, checked 17 September 2026.

    Performance, in my words
    OpenAI's flagship model for complex reasoning, coding and computer use. Its launch announcement reports state-of-the-art scores on several benchmarks; those are the maker's own figures, so check them against an independent leaderboard before relying on a specific claim.
    Prompting it from this site
    You can run this hereAvailable here as the “thinking” ChatGPT setting — it runs on this site's own AI credits, not your ChatGPT plan.
    Use it for
    • Demanding reasoning or research work
    • Agentic computer-use tasks
    • Complex coding through ChatGPT paid plans

    API use is priced per token and adds up fast; the free ChatGPT plan does not include Astra.

Coding assistants

4 options

Models wrapped in a tool that can read and change files. Useful even if you don't write code — this is how this site gets built.

  • Kilo

    Kilo
    Published price
    • IndividualOpen source, free forever$0
    • Model useOr $0 with the free model routingProvider cost, zero markup
    • Teams14-day trial$15/user/month
    • EnterpriseCustom

    Source: Kilo's pricing page · Last checked 17 September 2026

    Published benchmark scores
    • KiloBench (Terminal-Bench 2.0, run in Kilo's harness)76.2% at $87.41 per attempt — 1st

      share of real command-line tasks completed unaided, with the API cost of trying · tested: GPT-5.6 Sol driven by Kilo

      Kilo's own KiloBench board · checked September 2026 · maker's own figures

    • KiloBench (Terminal-Bench 2.0, run in Kilo's harness)75.3% at $112.27 per attempt — 3rd

      share of real command-line tasks completed unaided, with the API cost of trying · tested: Gemini 3.8 Flash driven by Kilo

      Kilo's own KiloBench board · checked September 2026 · maker's own figures

    Kilo now publishes its own board — KiloBench — which runs Terminal-Bench 2.0 through Kilo's actual agent harness and reports the cost of each attempt alongside the pass rate. These are Kilo's own figures, not an independent leaderboard, and each one belongs to the model plugged in rather than to Kilo itself. The 88% SWE-bench figure floating around one comparison blog still has no primary source, so it stays out.

    Figures copied from the sources above, checked 17 September 2026.

    1 more published figure on its own page.

    Performance, in my words
    As good as whichever model you plug into it. Works in VS Code, JetBrains, the terminal and the cloud.
    Prompting it from this site
    Can't be run from hereA coding assistant that edits files in your editor, not a model you send one prompt to.
    Use it for
    • Editing a project without paying a subscription
    • Keeping code private by running a local model
  • OpenCode

    OpenCode
    Published price
    • OpenCodeOpen source, bring your own model key$0
    • Zen gatewayStarts with a $20 balancePay as you go, no subscription
    • Zen — Claude Opus 4.5$5.00 in / $25.00 out per million tokens
    • Zen — Claude Haiku 4.5$1.00 in / $5.00 out per million tokens

    Source: OpenCode Zen pricing · Last checked 17 September 2026

    Published benchmark scores
    • Terminal-Bench 2.051.7% — 66th place

      share of real command-line tasks completed unaided · tested: OpenCode driving Claude Opus 4.5

      Terminal Trove comparison · checked September 2026 · independent

      Third-party comparison site, not the official leaderboard.

    OpenCode publishes no score of its own; the one public figure tests it as a harness driving somebody else's model.

    Figures copied from the sources above, checked 17 September 2026.

    Performance, in my words
    A terminal-first coding agent — fast and direct, less hand-holding than an editor plugin.
    Prompting it from this site
    Can't be run from hereA terminal coding agent — it works on a repository, not a single prompt.
    Use it for
    • Command-line work
    • Batch changes across many files
  • Devin

    Cognition
    Published price
    • FreeLight quota, limited models$0/month
    • Pro$20/month
    • Max$200/month
    • Teams$80/month + $40 per developer seat
    • EnterpriseCustom, billed in compute units

    Source: Devin's pricing page · Last checked 17 September 2026

    Published benchmark scores
    • SWE-bench (original, unassisted)13.86%

      share of real GitHub issues resolved end to end

      Cognition's own technical report · 15 March 2024 · maker's own figures

      The maker's own figure, and now two years old.

    • SWE-bench Verified48.2% or 61.7%, depending on the source

      share of human-checked GitHub issues resolved end to end

      Dataku (48.2%) and Tensorfeed (61.7%) · 9 June 2025 / undated · independent

      The two aggregators contradict each other and neither is the official leaderboard. Treat as unverified.

    Devin's numbers are a mess to cite honestly: Cognition's only first-party report predates the SWE-bench Verified set, and the two aggregators that list a current figure disagree by 13 points. Both are below.

    Figures copied from the sources above, checked 17 September 2026.

    Performance, in my words
    Aims to take a whole ticket end to end rather than answer one question. Impressive on well-scoped tasks, expensive when it wanders.
    Prompting it from this site
    Can't be run from hereRuns only inside Cognition's own cloud, on whole projects.
    Use it for
    • Understanding how autonomous coding agents are designed

    Listed here for the free documentation and free tier — the useful quotas are paid.

  • Herdr

    Herdr
    Published price
    • HerdrApache 2.0, installed and run by you$0
    • Model useWhatever your own API key costs

    Herdr has no pricing page — there is nothing to buy from them.

    Source: the Herdr site · Last checked 17 September 2026

    Published benchmark scores

    No published score. Herdr isn't a model — it runs several agents side by side, so there's no capability score to publish. The only figures its site quotes are speed and processor use against tmux, which measure the terminal, not the AI. A benchmark for this kind of orchestration now exists (ClawArena-Team, July 2026), but nobody has run Herdr on it, so there is still nothing to quote.

    Performance, in my words
    Coordinates several agents on one job instead of running a single assistant.
    Prompting it from this site
    Can't be run from hereA runtime that keeps other agents alive in terminals; there's no model here to prompt.
    Use it for
    • Running more than one agent at once
    • Comparing how models tackle the same task

Search, answers and research

2 options

Models pointed at live sources, so answers come with somewhere to check them.

  • Perplexity

    Perplexity AI
    Published price
    • FreeLimited daily searches, basic models$0/month
    • Pro$17/month billed annually$20/month
    • Max$167/month billed annually$200/month

    Source: Perplexity's pricing page · Last checked 17 September 2026

    Published benchmark scores
    • SimpleQA (F-score)85.8

      factual accuracy on short questions with one checkable answer · tested: Perplexity Sonar Pro

      Perplexity's own Sonar Pro announcement · 21 January 2025 · maker's own figures

      The maker's own figure. It is the number they lead with, and no independent tester has published a comparable SimpleQA run for Sonar Pro.

    • MMLU80.1%

      broad multiple-choice general knowledge · tested: Perplexity Sonar

      TPS Report · model released 27 January 2025 · independent

    Perplexity benchmarks its Sonar models rather than the search product you use in the browser, so read these as a floor, not a ceiling. The headline factuality figure is Perplexity's own, from the Sonar Pro launch post.

    Figures copied from the sources above, checked 17 September 2026.

    3 more published figures on its own page.

    Performance, in my words
    Cites its sources on every answer, which makes it far safer than a plain chatbot for anything factual.
    Prompting it from this site
    Endpoint wired, waiting on a keySonar Pro searches the live web and answers with sources. Runs through OpenRouter once a key is saved.
    Use it for
    • Researching an employer before an interview
    • Checking a salary range
    • Finding the original source of a claim

    Follow the citations. It sometimes summarizes a source more confidently than the source does.

  • Emergent Mind

    Emergent Mind
    Published price
    • Explore (free)All pre-generated paper and topic pages, plus 5 custom searches a week$0
    • Research (paid)Unlimited custom searches. Annual is advertised as 16% cheaper; 15-day money-back guarantee$12/month, or $120/year

    The figures come from the maker's own plan-update post — the pricing page itself renders its numbers after loading, so check them at sign-up.

    Source: Emergent Mind's pricing page · Last checked 17 September 2026

    Published benchmark scores

    No published score. Emergent Mind explains other people's research papers. It isn't a model and has no score of its own — the benchmarks it writes about belong to the papers.

    Performance, in my words
    Turns new arXiv papers into readable summaries and topic pages — no prompting required.
    Prompting it from this site
    Can't be run from hereA research search site rather than a model.
    Use it for
    • Following AI research without reading raw papers
    • Finding the paper behind a headline

Open models you can run yourself

3 options

Published weights you can download and run on your own machine. Free forever, private, and slower than a hosted model.

  • Xiaomi MiMo

    Xiaomi
    Published price
    • WeightsMIT license, commercial use and fine-tuning allowed$0
    • API — MiMo-V2.6-Pro$0.435 in / $0.87 out per million tokens

    Source: Xiaomi's MiMo pricing docs · Last checked 17 September 2026

    Published benchmark scores
    • Artificial Analysis Intelligence Index v4.3.246

      combined score across ten reasoning and knowledge tests · tested: MiMo-V2.6-Pro

      Artificial Analysis · 28 September 2026 · independent

      Measured at 46.9 tokens per second output speed.

    Figures are for MiMo-V2.6-Pro, the current version on Artificial Analysis. The card's two old variants (MiMo-V2-Flash and MiMo-7B-RL-0530) were tested against each other, so they've been replaced with the one current set.

    Figures copied from the sources above, checked 17 September 2026.

    Performance, in my words
    Solid open reasoning models. Nowhere near the biggest hosted models, plenty for summarizing, drafting and code help.
    Prompting it from this site
    Endpoint wired, waiting on a keyMiMo V2.5 through OpenRouter, about $0.14 per million tokens in — the cheapest of the four, though the card now tracks the newer MiMo-V2.6-Pro. Needs an OpenRouter key.
    Use it for
    • Running a capable model on your own hardware
    • Work you don't want sent to anyone's server
  • Hugging Face

    Hugging Face
    Published price
    • Free100GB private storage, 100,000 inference credits a month$0
    • PRO1TB private storage, 2M inference credits$9/month
    • Team$20/user/month
    • Enterprise$50/user/month
    • Spaces hardwareFree CPU tier, GPUs by the hour

    Source: Hugging Face's pricing page · Last checked 17 September 2026

    Published benchmark scores

    No published score. Hugging Face hosts models rather than making one, so there's nothing to score. Worth knowing: its own Open LLM Leaderboard was retired in March 2025, so anyone still citing it is citing something frozen.

    Performance, in my words
    Not one model but nearly all of them — plus the datasets, demos and courses around them.
    Prompting it from this site
    Can't be run from hereA hub holding thousands of models — pick one there and run it yourself; the setup steps on its page show how.
    Use it for
    • Finding an open model for a specific job
    • Trying a model in the browser before installing anything
  • GLM-5.3

    Z.ai
    Published price
    • WeightsOpenly published for download$0
    • API — GLM-5.3Cached input $0.26$1.40 in / $4.40 out per million tokens
    • API — GLM-4.5-Air$0.20 in / $1.10 out per million tokens
    • API — Flash modelsFree
    • Coding plansLite $18, Pro $72, Max $160 per month

    Source: Z.ai's pricing page · Last checked 17 September 2026

    Published benchmark scores
    • Artificial Analysis Intelligence Index v4.3.245

      combined score across ten reasoning and knowledge tests · tested: GLM-5.3

      Artificial Analysis · 28 September 2026 · independent

      Measured at 88.8 tokens per second output speed.

    • Vals AI Legal Research49.04% — 5th

      real legal research tasks answered correctly · tested: GLM-5.3

      Vals AI · 19 August 2026 · independent

    Figures are for GLM-5.3, the current version on Artificial Analysis. The old mix of GLM-5 and GLM-4.6 numbers has been removed so every score here belongs to the same release.

    Figures copied from the sources above, checked 17 September 2026.

    Performance, in my words
    One of the stronger open releases, particularly on reasoning and code, per the maker's own write-up.
    Prompting it from this site
    Endpoint wired, waiting on a keyServed through OpenRouter at roughly $1.40 per million tokens in — the wired endpoint still points at GLM-5.2. Needs an OpenRouter key saved here first.
    Use it for
    • Keeping track of what open models can now do

    The figures in the post are the maker's own. Cross-check against an independent benchmark.

Compare models and prices yourself

2 options

Where to check cost and performance claims — including mine — against measured numbers.

  • OpenRouter

    OpenRouter
    Published price
    • Free models25+ free models, 50 requests a day$0
    • InferenceProvider list price, no markup
    • Credit top-up fee5.5%
    • Own provider key (BYOK)5% of the equivalent OpenRouter cost

    Source: OpenRouter's pricing page · Last checked 17 September 2026

    Published benchmark scores

    No published score. OpenRouter is the shop, not the product: it ranks models by how much real developer traffic they get and republishes Artificial Analysis scores. It makes no capability claim of its own.

    Performance, in my words
    Hundreds of models behind one interface with prices side by side — the quickest way to see what a model actually costs per use.
    Prompting it from this site
    It's the route, not the modelOpenRouter is what the four models above would run through. Save an OpenRouter key here and they all switch on at once.
    Use it for
    • Comparing model prices
    • Trying several models on the same prompt
    • Using free models without a card
  • Vals AI Model Benchmarks

    Vals AI
    Published price
    • Benchmarks and reportsNo paywall, no account needed$0
    • Private evaluations for companiesA consulting engagement, so there is no list price to publishQuoted on request

    Source: the Vals AI benchmarks · Last checked 17 September 2026

    Published benchmark scores

    No published score. Vals AI runs the tests, so it has no score of its own. As an example of what it publishes: the top model on its LegalBench board scored 88.56% when I checked.

    Performance, in my words
    Independent evaluations on real professional tasks — law, finance, medicine — rather than vendor demos.
    Prompting it from this site
    Can't be run from hereA benchmark publisher, not a model.
    Use it for
    • Checking a vendor's performance claim
    • Picking a model for a specific kind of work

Looking for something else? Search the whole library or compare any two resources.