Qwen3.8-2.4T-A95B
Alibaba (Qwen team) · China · Released August 2026
The largest model Alibaba has published, and the top trending release on the Hub at launch. Same generation as the Apache-licensed 27B, under a different licence.
- Size
- 2.4T MoE (95B active)
- Context
- 262k tokens, native multimodal
- What running it yourself actually takes
- Roughly 1.2 TB at 4-bit. Multi-node territory. Community GGUF conversions exist and are mostly of academic interest unless you have a cluster.
- Realistic access
- Alibaba Cloud's API and the usual inference providers. FP8 weights are published alongside the BF16 ones if you are renting datacenter GPUs.
- Runnable sibling
- Qwen3.8-27B — dense, Apache 2.0, multimodal, and in the picker.
