Skip to content
Launchpad Library logo
All models

Chat and everyday writing

Kimi

Made by Moonshot AI

Moonshot's assistant — free in the browser and unusually good with very long documents.

What it is

Kimi is a general assistant like ChatGPT, but the thing worth knowing about it is how much text it will swallow at once. You can paste an entire report, a full benefits document or two long job descriptions and ask questions across all of it without chopping the file up first.

That makes it the one I reach for when the job is comprehension rather than composition: pull the real requirements out of a dense posting, tell me what changed between these two contracts, summarize this 60-page report into something I can act on.

Normal browser use is free with no subscription. The paid tiers and the API are for heavy or programmatic use.

Published price

  • FreeNormal browser use, no subscription$0
  • Moderato$15/month billed annually$19/month
  • Allegretto$31/month billed annually$39/month
  • Allegro$79/month billed annually$99/month
  • Vivace$159/month billed annually$199/month
  • API — Kimi K3$3.00 in / $15.00 out per million tokens

Source: Moonshot's membership pricing · Last checked 17 September 2026

Benchmarks

The published figures for Kimi, copied exactly as their sources give them and checked 17 September 2026. Each one names the test, who ran it and when, and says when the number comes from the maker rather than an independent tester.

  • Artificial Analysis Intelligence Index v4.3.244

    combined score across ten reasoning and knowledge tests · tested: Kimi K3

    Artificial Analysis · 28 September 2026 · independent

Figures are for Kimi K3, the current version on Artificial Analysis. No output speed was published for this model, so none is shown.

Figures copied from the sources above, checked 17 September 2026.

Where to read today's numbers

Scores move, so these are the boards that keep them current.

  • Ranks models by blind head-to-head votes from the public. Good for a feel for general quality, weak on specialist work.

  • Independent testing of quality, speed and price per million tokens across hosted models.

  • Ranks models by how much real developer traffic they actually get — usage, not quality.

How it performs, in my experience

Handles very long documents better than most free assistants — you can paste a whole report and ask about it.

  • Free value

    5 / 5

    How much you get without paying.

  • Performance

    4 / 5

    How well it does its main job.

  • Range of uses

    3 / 5

    How many different jobs it suits.

Free and unusually good with long documents; narrower than a general assistant. These scores are my own judgment, not a measurement — the benchmark links above are the independent version.

Prompting it from this site

Endpoint wired, waiting on a key

Kimi K2.5 runs through OpenRouter. Once an OpenRouter key is saved for this site the choice switches on; until then it's grayed out rather than failing when clicked.

Use it for

  • Summarizing long PDFs and reports
  • Comparing two long job descriptions
  • Pulling the requirements out of a dense posting

Running it yourself

Runs on your own machine

Moonshot publishes Kimi's weights openly, so it can run on your own hardware — but it is a very large model, so a laptop won't do it.

What you need first

  • A Linux machine with server-class GPUs — check the model card on Hugging Face for the exact memory figure for the version you pick
  • Python 3.10+ and a recent CUDA driver
  • Several hundred gigabytes of free disk for the weights

Step by step

  1. 1.Install the serving engine

    pip install vllm huggingface_hub
  2. 2.Download the weights

    hf download moonshotai/Kimi-K2-Instruct --local-dir ./kimi
  3. 3.Serve it as a local API

    vllm serve ./kimi --trust-remote-code --served-model-name kimi
  4. 4.Ask it something

    curl localhost:8000/v1/chat/completions -H 'Content-Type: application/json' -d '{"model":"kimi","messages":[{"role":"user","content":"hello"}]}'

Smaller quantized community builds of the same weights run on far less hardware with some loss of quality — search the Hugging Face hub for GGUF versions and run them with Ollama or llama.cpp.

Others in chat and everyday writing