Qwen 3.8
Alibaba (Qwen team) · China
2026- License
- Apache 2.0 on the 27B; custom Qwen licence on the 2.4T MoE
- Context
- 262k tokens
- Sizes
- 27B dense · 2.4T-A95B MoE
CodingMultimodal workflowsMultilingualLocal workstations
August 2026, and the most-liked open-weight release of the summer. The 27B is dense, Apache 2.0, natively multimodal, and had GGUF builds from unsloth and bartowski within days — which makes it the current default for a 24 GB card. Prism ML's Bonsai 2 build squeezes the same 27B into about 6 GB, provided you run it on their llama.cpp fork. The 2.4T MoE sibling is frontier-scale and carries a different licence.
DeepSeek V4 Flash
DeepSeek AI · China
2026- License
- MIT
- Context
- 1M tokens (65k native, extended 16x with YaRN)
- Sizes
- 304B-A13B MoE (284B before the DSpark decoder)
Agentic codingLong documentsHigh-memory workstations
The most-downloaded open-weight model on the Hub, and the largest entry here that a single machine can still load. 256 routed experts with 6 active per token means roughly 13B parameters fire on any given token, so it answers far faster than 304B suggests. The 0731 checkpoint supersedes the June preview and carries a DSpark speculative-decoding module, which is where the extra weights over the 284B base come from. Weights ship natively in FP8, so a Q8 build costs only about 7 GB more than Q4.
GLM-5.3-Flash
Z.ai (Zhipu) · China
2026- License
- MIT
- Context
- 1M tokens (three linear-attention layers for every sparse one)
- Sizes
- 321B MoE (18B active)
Agentic codingMultimodal workflowsHigh-memory workstations
Late August 2026, and the first natively multimodal GLM. A new base model rather than a GLM-5.2 derivative, MIT licensed, and at 5 million downloads the most pulled GLM on the Hub. unsloth's GGUF builds run from about 98 GB (UD-IQ1_M) through 109 GB (UD-Q2_K_XL) to 200 GB (UD-Q4_K_XL), which puts it in the same 128 GB class as DeepSeek V4 Flash. Not in the picker yet: the hybrid attention needs a llama.cpp pull request that has not reached a release, and we do not print commands for a build you would have to assemble by hand.
MiMo-V2.6-Flash
Xiaomi (MiMo team) · China
2026- License
- MIT
- Context
- 1M tokens
- Sizes
- 309B MoE (15B active)
Agentic codingMultimodal workflows (text, image, video, audio)High-memory workstations
September 2026. The efficiency-balanced checkpoint of MiMo V2.6: 309B total with 15B active, text, image, video and audio in one model, and a 1M context thanks to a mostly sliding-window backbone. ggml-org publishes the GGUF conversion, which means stock llama.cpp runs it: Q2_K is about 126 GB and MXFP4, which keeps the experts at their native precision, about 167 GB. That puts it just past a 128 GB machine and comfortably inside 192 GB. Benchmarks on the card are Xiaomi's own.
MiniMax M2.7
MiniMax · China
2026- License
- MiniMax non-commercial licence (commercial use needs written permission)
- Context
- 205k tokens
- Sizes
- 229B MoE (8 of 256 experts per token)
Agentic codingPersonal self-hostingHigh-memory workstations
April 2026, and still one of the most downloaded large models on the Hub. Read the licence before anything else: personal use, self-hosted deployment for your own coding and agents, and non-profit research are explicitly free, while any commercial use needs prior written authorisation from MiniMax. unsloth's GGUF builds run from about 75 GB (UD-Q2_K_XL) to 141 GB (UD-Q4_K_XL), so a 96 or 128 GB machine can hold it. Not in the picker, because the picker's licence filter cannot yet express 'free at home, not at work'.
Qwen 3.6
Alibaba (Qwen team) · China
2026- License
- Apache 2.0 on official open-weight checkpoints
- Context
- 262k native on 27B; long-context extensions available
- Sizes
- 27B dense · 35B-A3B MoE · additional large MoE variants
CodingMultimodal workflowsMultilingualLocal workstations
The previous open-weight Qwen generation, still widely deployed and the base for most community quantization work, including the first-generation 1-bit and ternary Bonsai builds. Qwen 3.8 supersedes it for new installs, and Bonsai 2 has moved to the 3.8 base.
Qwen 3.5
Alibaba (Qwen team) · China
2026- License
- Apache 2.0
- Context
- 262k native; selected variants support longer contexts
- Sizes
- 2B · 9B · 27B · 35B-A3B MoE · 122B-A10B MoE · 397B-A17B MoE
MultimodalMultilingualCodeCost-sensitive deployments
Released in 2026. The smaller checkpoints remain excellent local choices when Qwen 3.6 is too large.
GLM-4.7-Flash
Z.ai (Zhipu) · China
2026- License
- MIT
- Context
- 202k tokens
- Sizes
- 31B MoE
Bilingual EN/ZH workloadsAgentic codingWorkstation GPUs
The runnable member of the GLM family, and the counterpart to the frontier-scale GLM-5.3: MIT licensed, 200k context, mature GGUF ecosystem, and one of the strongest agentic-coding models that fits on a single workstation GPU.
MiniCPM5 (1B and 2B)
OpenBMB · China
2026- License
- Apache 2.0
- Context
- 131k tokens
- Sizes
- 1B · 2.5B dense (sold as 2B)
On-device inferenceEdge deploymentsLow-memory hardwareTool calling
Edge-first family for phones and low-memory laptops. The 2B arrived in September 2026 with the same training recipe scaled up, official GGUF files from OpenBMB, and tool calling in the chat template; its Q4_K_M file is about 1.6 GB. OpenBMB reports it competitive with 4B-class models on coding and maths, a vendor figure worth testing before relying on it.
Xing4.0-29B-A4B
China Telecom AI (XingChen-AGI, formerly TeleChat) · China
2026- License
- Apache 2.0
- Context
- 256k tokens (512k with extension)
- Sizes
- 29B MoE (4B active)
Agentic coding24 GB GPUsDomain fine-tuning
September 2026, and the first model of this size trained entirely on Huawei Ascend hardware. The official GGUF is a single 18 GB IQ4_NL file that fits a 24 GB card. Not in the picker: the architecture is custom and the official guide builds llama.cpp from a dedicated xing4_0-port branch, which we do not print commands for. vLLM, SGLang and KTransformers support it directly.