Chat and everyday writing
Kimi
Made by Moonshot AI
Moonshot's assistant — free in the browser and unusually good with very long documents.
What it is
Kimi is a general assistant like ChatGPT, but the thing worth knowing about it is how much text it will swallow at once. You can paste an entire report, a full benefits document or two long job descriptions and ask questions across all of it without chopping the file up first.
That makes it the one I reach for when the job is comprehension rather than composition: pull the real requirements out of a dense posting, tell me what changed between these two contracts, summarize this 60-page report into something I can act on.
Normal browser use is free with no subscription. The paid tiers and the API are for heavy or programmatic use.
Published price
- FreeNormal browser use, no subscription$0
- Moderato$15/month billed annually$19/month
- Allegretto$31/month billed annually$39/month
- Allegro$79/month billed annually$99/month
- Vivace$159/month billed annually$199/month
- API — Kimi K3$3.00 in / $15.00 out per million tokens
Source: Moonshot's membership pricing · Last checked 17 September 2026
Benchmarks
The published figures for Kimi, copied exactly as their sources give them and checked 17 September 2026. Each one names the test, who ran it and when, and says when the number comes from the maker rather than an independent tester.
- Artificial Analysis Intelligence Index v4.3.244
combined score across ten reasoning and knowledge tests · tested: Kimi K3
Artificial Analysis · 28 September 2026 · independent
Figures are for Kimi K3, the current version on Artificial Analysis. No output speed was published for this model, so none is shown.
Figures copied from the sources above, checked 17 September 2026.
Where to read today's numbers
Scores move, so these are the boards that keep them current.
- LMArena leaderboardIndependent
Ranks models by blind head-to-head votes from the public. Good for a feel for general quality, weak on specialist work.
- Artificial AnalysisIndependent
Independent testing of quality, speed and price per million tokens across hosted models.
- OpenRouter rankingsIndependent
Ranks models by how much real developer traffic they actually get — usage, not quality.
How it performs, in my experience
Handles very long documents better than most free assistants — you can paste a whole report and ask about it.
Free value
5 / 5
How much you get without paying.
Performance
4 / 5
How well it does its main job.
Range of uses
3 / 5
How many different jobs it suits.
Free and unusually good with long documents; narrower than a general assistant. These scores are my own judgment, not a measurement — the benchmark links above are the independent version.
Prompting it from this site
Endpoint wired, waiting on a key
Kimi K2.5 runs through OpenRouter. Once an OpenRouter key is saved for this site the choice switches on; until then it's grayed out rather than failing when clicked.
Use it for
- Summarizing long PDFs and reports
- Comparing two long job descriptions
- Pulling the requirements out of a dense posting
Running it yourself
Runs on your own machine
Moonshot publishes Kimi's weights openly, so it can run on your own hardware — but it is a very large model, so a laptop won't do it.
What you need first
- A Linux machine with server-class GPUs — check the model card on Hugging Face for the exact memory figure for the version you pick
- Python 3.10+ and a recent CUDA driver
- Several hundred gigabytes of free disk for the weights
Step by step
1.Install the serving engine
pip install vllm huggingface_hub2.Download the weights
hf download moonshotai/Kimi-K2-Instruct --local-dir ./kimi3.Serve it as a local API
vllm serve ./kimi --trust-remote-code --served-model-name kimi4.Ask it something
curl localhost:8000/v1/chat/completions -H 'Content-Type: application/json' -d '{"model":"kimi","messages":[{"role":"user","content":"hello"}]}'
Smaller quantized community builds of the same weights run on far less hardware with some loss of quality — search the Hugging Face hub for GGUF versions and run them with Ollama or llama.cpp.
