Skip to content
Launchpad Library logo
All models

Open models you can run yourself

GLM-5.3

Made by Z.ai

One of the stronger open releases, particularly on reasoning and code — weights published.

What it is

GLM-5.3 is Z.ai's flagship open release, with weights published for download and a hosted API for people who don't want to run it. Its reputation is strongest on reasoning and code.

It matters here as a marker of what open models can now do: the gap between what you can download for nothing and what you can rent by the token has narrowed considerably.

The weights are free; the API is priced per million tokens, with cheaper Air and free Flash variants and monthly coding plans.

The figures in the post are the maker's own. Cross-check against an independent benchmark.

Published price

  • WeightsOpenly published for download$0
  • API — GLM-5.3Cached input $0.26$1.40 in / $4.40 out per million tokens
  • API — GLM-4.5-Air$0.20 in / $1.10 out per million tokens
  • API — Flash modelsFree
  • Coding plansLite $18, Pro $72, Max $160 per month

Source: Z.ai's pricing page · Last checked 17 September 2026

Benchmarks

The published figures for GLM-5.3, copied exactly as their sources give them and checked 17 September 2026. Each one names the test, who ran it and when, and says when the number comes from the maker rather than an independent tester.

  • Artificial Analysis Intelligence Index v4.3.245

    combined score across ten reasoning and knowledge tests · tested: GLM-5.3

    Artificial Analysis · 28 September 2026 · independent

    Measured at 88.8 tokens per second output speed.

  • Vals AI Legal Research49.04% — 5th

    real legal research tasks answered correctly · tested: GLM-5.3

    Vals AI · 19 August 2026 · independent

Figures are for GLM-5.3, the current version on Artificial Analysis. The old mix of GLM-5 and GLM-4.6 numbers has been removed so every score here belongs to the same release.

Figures copied from the sources above, checked 17 September 2026.

Where to read today's numbers

Scores move, so these are the boards that keep them current.

  • Standardized test scores for open weights you can download and run yourself.

  • Independent testing of quality, speed and price per million tokens across hosted models.

  • Measures whether an agent can fix real GitHub issues. The standard reference for coding agents.

  • Z.ai's own release postMaker's own figures

    The figures in the announcement are the maker's own — the reason to check an independent source.

How it performs, in my experience

One of the stronger open releases, particularly on reasoning and code, per the maker's own write-up.

  • Free value

    5 / 5

    How much you get without paying.

  • Performance

    4 / 5

    How well it does its main job.

  • Range of uses

    3 / 5

    How many different jobs it suits.

Strong open release, but the published figures are the maker's own. These scores are my own judgment, not a measurement — the benchmark links above are the independent version.

Prompting it from this site

Endpoint wired, waiting on a key

Served through OpenRouter at roughly $1.40 per million tokens in — the wired endpoint still points at GLM-5.2. Needs an OpenRouter key saved here first.

Use it for

  • Keeping track of what open models can now do

Running it yourself

Runs on your own machine

Z.ai publishes GLM's weights openly, so it can be self-hosted — but the full-size version is a data-center model, not a desktop one.

What you need first

  • Multiple server-class GPUs for the full model — the model card states the figure for each release
  • Python 3.10+ with CUDA
  • Substantial free disk for the weights

Step by step

  1. 1.Install a serving engine

    pip install vllm huggingface_hub
  2. 2.Download the weights

    hf download zai-org/GLM-4.6 --local-dir ./glm
  3. 3.Serve it

    vllm serve ./glm --trust-remote-code --served-model-name glm

Replace the repository name with the release you actually want — check the maker's Hugging Face page for the current one before downloading hundreds of gigabytes.

Others in open models you can run yourself