Skip to content
Launchpad Library logo
All models

Open models you can run yourself

Xiaomi MiMo

Made by Xiaomi

Open reasoning weights under an MIT license — run them privately on your own hardware.

What it is

MiMo is a family of open models with published weights, released under a license that permits commercial use and fine-tuning. Download them and they run on your machine, with nothing sent to anyone's server.

They are not the equal of the biggest hosted models, and they will be slower on consumer hardware. For summarizing, drafting and code help — particularly on anything confidential — they are plenty.

The weights are free. Xiaomi also sells hosted API access, priced per million tokens, if you'd rather not run it yourself.

Published price

  • WeightsMIT license, commercial use and fine-tuning allowed$0
  • API — MiMo-V2.6-Pro$0.435 in / $0.87 out per million tokens

Source: Xiaomi's MiMo pricing docs · Last checked 17 September 2026

Benchmarks

The published figures for Xiaomi MiMo, copied exactly as their sources give them and checked 17 September 2026. Each one names the test, who ran it and when, and says when the number comes from the maker rather than an independent tester.

  • Artificial Analysis Intelligence Index v4.3.246

    combined score across ten reasoning and knowledge tests · tested: MiMo-V2.6-Pro

    Artificial Analysis · 28 September 2026 · independent

    Measured at 46.9 tokens per second output speed.

Figures are for MiMo-V2.6-Pro, the current version on Artificial Analysis. The card's two old variants (MiMo-V2-Flash and MiMo-7B-RL-0530) were tested against each other, so they've been replaced with the one current set.

Figures copied from the sources above, checked 17 September 2026.

Where to read today's numbers

Scores move, so these are the boards that keep them current.

How it performs, in my experience

Solid open reasoning models. Nowhere near the biggest hosted models, plenty for summarizing, drafting and code help.

  • Free value

    5 / 5

    How much you get without paying.

  • Performance

    3 / 5

    How well it does its main job.

  • Range of uses

    3 / 5

    How many different jobs it suits.

Free open weights you can run privately; slower than anything hosted. These scores are my own judgment, not a measurement — the benchmark links above are the independent version.

Prompting it from this site

Endpoint wired, waiting on a key

MiMo V2.5 through OpenRouter, about $0.14 per million tokens in — the cheapest of the four, though the card now tracks the newer MiMo-V2.6-Pro. Needs an OpenRouter key.

Use it for

  • Running a capable model on your own hardware
  • Work you don't want sent to anyone's server

Running it yourself

Runs on your own machine

MiMo's weights are published openly, and the smaller sizes are runnable on a single decent GPU — this is one of the more approachable open models here.

What you need first

  • A machine with an NVIDIA GPU (check the model card for the memory needed by the size you choose)
  • Python 3.10+ with a working CUDA install

Step by step

  1. 1.Install the tooling

    pip install vllm transformers huggingface_hub
  2. 2.Download the weights

    hf download XiaomiMiMo/MiMo-7B-RL --local-dir ./mimo
  3. 3.Serve it locally

    vllm serve ./mimo --trust-remote-code

For a CPU-only machine, look for a community GGUF conversion and run it with Ollama or llama.cpp instead — slower, but it works without a GPU.

Others in open models you can run yourself