Analysis · 7 min · September 29, 2026
GLM-5.3 is two different models sharing a version number
If you read “GLM-5.3” in a benchmark table this month, it is worth checking which of the two it means. The repository called zai-org/GLM-5.3 holds a 753-billion-parameter model with about 1.4 million downloads. The one called GLM-5.3-Flash holds a 321-billion-parameter model with about 5.1 million. They share a version number, a tokenizer and a technical report, and not much else.
The big one is GLM-5.2 with more training on top
Z.ai says so plainly on the model card: GLM-5.3 uses the same base model as GLM-5.2, and every improvement comes from post-training. The improvements they report are large on agentic coding. Terminal Bench 3.0 goes from 4.6 to 28.3, DeepSWE from 46.2 to 66.9, and they claim the best open-weight coding results to date. They also report, with what reads as some surprise, that cyber-offence capability grew faster than they expected during post-training: on ExploitGym GLM-5.3 more than triples GLM-5.2's score. All of this is Z.ai's own table, run in Claude Code as the agent harness, and none of it has been independently reproduced yet.
The part that changed quietly is the licence. GLM-5.2 was MIT. GLM-5.3 ships under a “GLM-5.3 License” that reads like MIT with one clause added: a company that sells model access as a service, and whose group revenue exceeds ten billion dollars over twelve months, has to pass a security review by Z.ai before using the model commercially. For anyone reading this site the clause changes nothing in practice, since you are not a hyperscaler. It is still a move away from a standard licence, in the same direction Alibaba took with the 2.4T Qwen, and it is the reason our Frontier page now files GLM-5.3 as open-weight rather than permissive.
As for running it: unsloth's smallest GGUF, UD-IQ1_S, is about 217 GB, UD-Q2_K_XL is 254 GB, and UD-Q4_K_XL is 467 GB. A 256 GB workstation can technically load the 1-bit build, and there is no configuration short of a server where it is a comfortable daily tool. It replaces GLM-5.2 on our Frontier page, and that is where it belongs.
The Flash is a new model, and the more interesting one
GLM-5.3-Flash starts from a newly trained base. It is the first natively multimodal GLM, it has 321 billion parameters with about 18 billion active per token, and it mixes attention types: three linear-attention layers for every sparse-attention layer, which is how it reaches a million-token context without the memory cost of a million-token cache. It is MIT licensed, the plain version with no added clauses. Z.ai claims it beats GLM-5.2 across the board at a tenth of the API price; again, their numbers.
The download count makes the argument we made about Qwen 3.8 last month, from a slightly different angle. The Flash has almost four times the downloads of the big model, and unsloth's GGUF conversion alone has close to a million. People are not waiting for the flagship; they are pulling the model whose size and licence let them do something with it.
On memory it sits in the same class as DeepSeek V4 Flash. unsloth's builds run from about 98 GB for UD-IQ1_M to 109 GB for UD-Q2_K_XL and 200 GB for UD-Q4_K_XL, so a 128 GB Mac or a workstation with that much unified or system memory can hold the 2-bit build with some room for context.
Why it is not in the picker
The hybrid attention is new to llama.cpp. unsloth's own model card says to run the GGUF with their llama.cpp pull request or with Unsloth Desktop, which means that at the time of writing a stock llama.cpp release, Ollama and LM Studio will not load it. We list it in the model directory with the sizes above, and we will add it to the picker when support lands in a llama.cpp release. Printing a command that depends on checking out an unmerged pull request would be setting people up to fail, and a pull request can change shape several times before it merges.
If you have 128 GB and a taste for building from source, the path exists today and unsloth documents it. Everyone else should keep using what already works at their size: GLM-4.7-Flash on a 24 GB card, or DeepSeek V4 Flash on a 128 GB machine, both of which run on stock llama.cpp and are in the picker with verified files.
A naming habit worth resisting
Giving a post-trained refresh of an old base and a genuinely new architecture the same version number is Z.ai's choice, and benchmark aggregators will not always say which one they tested. When a chart says “GLM-5.3” and shows numbers that look too good for a 753B model you cannot run, or too modest for the flagship, check the repository name before you draw conclusions. On our pages the two are always named in full.
