Open models you can run yourself
GLM-5.3
Made by Z.ai
One of the stronger open releases, particularly on reasoning and code — weights published.
What it is
GLM-5.3 is Z.ai's flagship open release, with weights published for download and a hosted API for people who don't want to run it. Its reputation is strongest on reasoning and code.
It matters here as a marker of what open models can now do: the gap between what you can download for nothing and what you can rent by the token has narrowed considerably.
The weights are free; the API is priced per million tokens, with cheaper Air and free Flash variants and monthly coding plans.
The figures in the post are the maker's own. Cross-check against an independent benchmark.
Published price
- WeightsOpenly published for download$0
- API — GLM-5.3Cached input $0.26$1.40 in / $4.40 out per million tokens
- API — GLM-4.5-Air$0.20 in / $1.10 out per million tokens
- API — Flash modelsFree
- Coding plansLite $18, Pro $72, Max $160 per month
Source: Z.ai's pricing page · Last checked 17 September 2026
Benchmarks
The published figures for GLM-5.3, copied exactly as their sources give them and checked 17 September 2026. Each one names the test, who ran it and when, and says when the number comes from the maker rather than an independent tester.
- Artificial Analysis Intelligence Index v4.3.245
combined score across ten reasoning and knowledge tests · tested: GLM-5.3
Artificial Analysis · 28 September 2026 · independent
Measured at 88.8 tokens per second output speed.
- Vals AI Legal Research49.04% — 5th
real legal research tasks answered correctly · tested: GLM-5.3
Vals AI · 19 August 2026 · independent
Figures are for GLM-5.3, the current version on Artificial Analysis. The old mix of GLM-5 and GLM-4.6 numbers has been removed so every score here belongs to the same release.
Figures copied from the sources above, checked 17 September 2026.
Where to read today's numbers
Scores move, so these are the boards that keep them current.
- Hugging Face open model leaderboardsIndependent
Standardized test scores for open weights you can download and run yourself.
- Artificial AnalysisIndependent
Independent testing of quality, speed and price per million tokens across hosted models.
- SWE-bench leaderboardIndependent
Measures whether an agent can fix real GitHub issues. The standard reference for coding agents.
- Z.ai's own release postMaker's own figures
The figures in the announcement are the maker's own — the reason to check an independent source.
How it performs, in my experience
One of the stronger open releases, particularly on reasoning and code, per the maker's own write-up.
Free value
5 / 5
How much you get without paying.
Performance
4 / 5
How well it does its main job.
Range of uses
3 / 5
How many different jobs it suits.
Strong open release, but the published figures are the maker's own. These scores are my own judgment, not a measurement — the benchmark links above are the independent version.
Prompting it from this site
Endpoint wired, waiting on a key
Served through OpenRouter at roughly $1.40 per million tokens in — the wired endpoint still points at GLM-5.2. Needs an OpenRouter key saved here first.
Use it for
- Keeping track of what open models can now do
Running it yourself
Runs on your own machine
Z.ai publishes GLM's weights openly, so it can be self-hosted — but the full-size version is a data-center model, not a desktop one.
What you need first
- Multiple server-class GPUs for the full model — the model card states the figure for each release
- Python 3.10+ with CUDA
- Substantial free disk for the weights
Step by step
1.Install a serving engine
pip install vllm huggingface_hub2.Download the weights
hf download zai-org/GLM-4.6 --local-dir ./glm3.Serve it
vllm serve ./glm --trust-remote-code --served-model-name glm
Replace the repository name with the release you actually want — check the maker's Hugging Face page for the current one before downloading hundreds of gigabytes.
