Open models you can run yourself
Xiaomi MiMo
Made by Xiaomi
Open reasoning weights under an MIT license — run them privately on your own hardware.
What it is
MiMo is a family of open models with published weights, released under a license that permits commercial use and fine-tuning. Download them and they run on your machine, with nothing sent to anyone's server.
They are not the equal of the biggest hosted models, and they will be slower on consumer hardware. For summarizing, drafting and code help — particularly on anything confidential — they are plenty.
The weights are free. Xiaomi also sells hosted API access, priced per million tokens, if you'd rather not run it yourself.
Published price
- WeightsMIT license, commercial use and fine-tuning allowed$0
- API — MiMo-V2.6-Pro$0.435 in / $0.87 out per million tokens
Source: Xiaomi's MiMo pricing docs · Last checked 17 September 2026
Benchmarks
The published figures for Xiaomi MiMo, copied exactly as their sources give them and checked 17 September 2026. Each one names the test, who ran it and when, and says when the number comes from the maker rather than an independent tester.
- Artificial Analysis Intelligence Index v4.3.246
combined score across ten reasoning and knowledge tests · tested: MiMo-V2.6-Pro
Artificial Analysis · 28 September 2026 · independent
Measured at 46.9 tokens per second output speed.
Figures are for MiMo-V2.6-Pro, the current version on Artificial Analysis. The card's two old variants (MiMo-V2-Flash and MiMo-7B-RL-0530) were tested against each other, so they've been replaced with the one current set.
Figures copied from the sources above, checked 17 September 2026.
Where to read today's numbers
Scores move, so these are the boards that keep them current.
- Hugging Face open model leaderboardsIndependent
Standardized test scores for open weights you can download and run yourself.
- Artificial AnalysisIndependent
Independent testing of quality, speed and price per million tokens across hosted models.
- Xiaomi's own model cardsMaker's own figures
The maker's published scores, alongside the weights. Compare against the open leaderboards.
How it performs, in my experience
Solid open reasoning models. Nowhere near the biggest hosted models, plenty for summarizing, drafting and code help.
Free value
5 / 5
How much you get without paying.
Performance
3 / 5
How well it does its main job.
Range of uses
3 / 5
How many different jobs it suits.
Free open weights you can run privately; slower than anything hosted. These scores are my own judgment, not a measurement — the benchmark links above are the independent version.
Prompting it from this site
Endpoint wired, waiting on a key
MiMo V2.5 through OpenRouter, about $0.14 per million tokens in — the cheapest of the four, though the card now tracks the newer MiMo-V2.6-Pro. Needs an OpenRouter key.
Use it for
- Running a capable model on your own hardware
- Work you don't want sent to anyone's server
Running it yourself
Runs on your own machine
MiMo's weights are published openly, and the smaller sizes are runnable on a single decent GPU — this is one of the more approachable open models here.
What you need first
- A machine with an NVIDIA GPU (check the model card for the memory needed by the size you choose)
- Python 3.10+ with a working CUDA install
Step by step
1.Install the tooling
pip install vllm transformers huggingface_hub2.Download the weights
hf download XiaomiMiMo/MiMo-7B-RL --local-dir ./mimo3.Serve it locally
vllm serve ./mimo --trust-remote-code
For a CPU-only machine, look for a community GGUF conversion and run it with Ollama or llama.cpp instead — slower, but it works without a GPU.
