RunLocal

Model directory

AI models you can download and run yourself.

Each entry below is a family of AI models you can download for free and run on your own computer. The label on each card tells you who made it, how big it is, what licence it comes under, and what it is good at. New to all this? The glossary defines every term used here, or try the picker to find one that fits your machine.

China10 models

Qwen 3.8

Alibaba (Qwen team) · China

2026
License
Apache 2.0 on the 27B; custom Qwen licence on the 2.4T MoE
Context
262k tokens
Sizes
27B dense · 2.4T-A95B MoE
CodingMultimodal workflowsMultilingualLocal workstations

August 2026, and the most-liked open-weight release of the summer. The 27B is dense, Apache 2.0, natively multimodal, and had GGUF builds from unsloth and bartowski within days — which makes it the current default for a 24 GB card. Prism ML's Bonsai 2 build squeezes the same 27B into about 6 GB, provided you run it on their llama.cpp fork. The 2.4T MoE sibling is frontier-scale and carries a different licence.

DeepSeek V4 Flash

DeepSeek AI · China

2026
License
MIT
Context
1M tokens (65k native, extended 16x with YaRN)
Sizes
304B-A13B MoE (284B before the DSpark decoder)
Agentic codingLong documentsHigh-memory workstations

The most-downloaded open-weight model on the Hub, and the largest entry here that a single machine can still load. 256 routed experts with 6 active per token means roughly 13B parameters fire on any given token, so it answers far faster than 304B suggests. The 0731 checkpoint supersedes the June preview and carries a DSpark speculative-decoding module, which is where the extra weights over the 284B base come from. Weights ship natively in FP8, so a Q8 build costs only about 7 GB more than Q4.

GLM-5.3-Flash

Z.ai (Zhipu) · China

2026
License
MIT
Context
1M tokens (three linear-attention layers for every sparse one)
Sizes
321B MoE (18B active)
Agentic codingMultimodal workflowsHigh-memory workstations

Late August 2026, and the first natively multimodal GLM. A new base model rather than a GLM-5.2 derivative, MIT licensed, and at 5 million downloads the most pulled GLM on the Hub. unsloth's GGUF builds run from about 98 GB (UD-IQ1_M) through 109 GB (UD-Q2_K_XL) to 200 GB (UD-Q4_K_XL), which puts it in the same 128 GB class as DeepSeek V4 Flash. Not in the picker yet: the hybrid attention needs a llama.cpp pull request that has not reached a release, and we do not print commands for a build you would have to assemble by hand.

MiMo-V2.6-Flash

Xiaomi (MiMo team) · China

2026
License
MIT
Context
1M tokens
Sizes
309B MoE (15B active)
Agentic codingMultimodal workflows (text, image, video, audio)High-memory workstations

September 2026. The efficiency-balanced checkpoint of MiMo V2.6: 309B total with 15B active, text, image, video and audio in one model, and a 1M context thanks to a mostly sliding-window backbone. ggml-org publishes the GGUF conversion, which means stock llama.cpp runs it: Q2_K is about 126 GB and MXFP4, which keeps the experts at their native precision, about 167 GB. That puts it just past a 128 GB machine and comfortably inside 192 GB. Benchmarks on the card are Xiaomi's own.

MiniMax M2.7

MiniMax · China

2026
License
MiniMax non-commercial licence (commercial use needs written permission)
Context
205k tokens
Sizes
229B MoE (8 of 256 experts per token)
Agentic codingPersonal self-hostingHigh-memory workstations

April 2026, and still one of the most downloaded large models on the Hub. Read the licence before anything else: personal use, self-hosted deployment for your own coding and agents, and non-profit research are explicitly free, while any commercial use needs prior written authorisation from MiniMax. unsloth's GGUF builds run from about 75 GB (UD-Q2_K_XL) to 141 GB (UD-Q4_K_XL), so a 96 or 128 GB machine can hold it. Not in the picker, because the picker's licence filter cannot yet express 'free at home, not at work'.

Qwen 3.6

Alibaba (Qwen team) · China

2026
License
Apache 2.0 on official open-weight checkpoints
Context
262k native on 27B; long-context extensions available
Sizes
27B dense · 35B-A3B MoE · additional large MoE variants
CodingMultimodal workflowsMultilingualLocal workstations

The previous open-weight Qwen generation, still widely deployed and the base for most community quantization work, including the first-generation 1-bit and ternary Bonsai builds. Qwen 3.8 supersedes it for new installs, and Bonsai 2 has moved to the 3.8 base.

Qwen 3.5

Alibaba (Qwen team) · China

2026
License
Apache 2.0
Context
262k native; selected variants support longer contexts
Sizes
2B · 9B · 27B · 35B-A3B MoE · 122B-A10B MoE · 397B-A17B MoE
MultimodalMultilingualCodeCost-sensitive deployments

Released in 2026. The smaller checkpoints remain excellent local choices when Qwen 3.6 is too large.

GLM-4.7-Flash

Z.ai (Zhipu) · China

2026
License
MIT
Context
202k tokens
Sizes
31B MoE
Bilingual EN/ZH workloadsAgentic codingWorkstation GPUs

The runnable member of the GLM family, and the counterpart to the frontier-scale GLM-5.3: MIT licensed, 200k context, mature GGUF ecosystem, and one of the strongest agentic-coding models that fits on a single workstation GPU.

MiniCPM5 (1B and 2B)

OpenBMB · China

2026
License
Apache 2.0
Context
131k tokens
Sizes
1B · 2.5B dense (sold as 2B)
On-device inferenceEdge deploymentsLow-memory hardwareTool calling

Edge-first family for phones and low-memory laptops. The 2B arrived in September 2026 with the same training recipe scaled up, official GGUF files from OpenBMB, and tool calling in the chat template; its Q4_K_M file is about 1.6 GB. OpenBMB reports it competitive with 4B-class models on coding and maths, a vendor figure worth testing before relying on it.

Xing4.0-29B-A4B

China Telecom AI (XingChen-AGI, formerly TeleChat) · China

2026
License
Apache 2.0
Context
256k tokens (512k with extension)
Sizes
29B MoE (4B active)
Agentic coding24 GB GPUsDomain fine-tuning

September 2026, and the first model of this size trained entirely on Huawei Ascend hardware. The official GGUF is a single 18 GB IQ4_NL file that fits a 24 GB card. Not in the picker: the architecture is custom and the official guide builds llama.cpp from a dedicated xing4_0-port branch, which we do not print commands for. vLLM, SGLang and KTransformers support it directly.

Not stated on the model card2 models

Ornith 1.5

ornith-ai · Not stated on the model card

2026
License
MIT
Context
262k tokens
Sizes
9B dense · 35B-A3B MoE · 397B MoE
Agentic codingMultimodal workflowsConsumer GPUs

August 2026. Unlike Ornith 1.0, which shipped with no model card at all, 1.5 documents its lineage (continued pretraining and reinforcement learning on top of Qwen 3.5 and Gemma 4) and publishes benchmarks, all vendor-run. The 9B and 35B-A3B keep the Qwen 3.5 architecture, so the official GGUF files run on stock llama.cpp. Download counts are very high and like counts unusually thin for the traffic, roughly one like per 7,900 downloads on the 35B GGUF; we list the model because its licence and lineage now check out, and we still cannot explain that ratio. The 397B is a single 244 GB file at Q4_K_M.

Spark-X2.5

XHToken · Not stated on the model card

2026
License
Apache 2.0
Context
1M tokens
Sizes
1.7B · 4B
Edge deploymentsMultilingual (200+ languages)Low-memory laptopsTool calling

September 2026. Two small dense models with a hybrid attention design (three sliding-window layers for every full-attention layer) that keeps a 1M context affordable. Official GGUF files from XHToken; the architecture is new, so you need llama.cpp build b10828 or later, Ollama 0.34.1 or later, or LM Studio runtime 2.34.0 or later. Quality claims on the card are the vendor's.

United States6 models

Gemma 4

Google DeepMind · United States

2026
License
Apache 2.0
Context
128k–256k depending on checkpoint
Sizes
E2B · E4B · 12B · 26B-A4B MoE · 31B dense
On-device inferenceMultimodalConsumer GPUsWorkstations

Multimodal family spanning edge through workstation hardware. The 26B-A4B model activates about 4B parameters per token while retaining a 26B memory footprint.

Phi-4 family

Microsoft Research · United States

2025
License
MIT
Context
Varies by checkpoint
Sizes
Phi-4 Mini 3.8B · Phi-4 14B · reasoning variants
Edge devicesReasoning per parameterCost-sensitive inference

Verified Microsoft Phi family. RunLocal does not list an unverified Phi-5 family.

Llama 4 (Scout & Maverick)

Meta AI · United States

2025
License
Llama Community License (custom)
Context
Up to 10M tokens (Scout)
Sizes
Scout 109B MoE · Maverick 400B MoE
Long-context retrievalCodebase-scale RAGGeneral reasoning

Meta's open-weight family. Custom licensing applies; Scout and Maverick are MoE models with much smaller active parameter counts than total weights.

gpt-oss (20B & 120B)

OpenAI · United States

2025
License
Apache 2.0
Context
131k tokens
Sizes
20B MoE (MXFP4 weights) · 120B MoE
General assistant workReasoningConsumer GPUs

OpenAI's open-weight release and, by download count, the most used open model of the past year. Ships natively in MXFP4, so the 20B fits a 16 GB card at the precision it was trained for. Apache 2.0, no usage caps, no licence acceptance step.

Nemotron 3.5 Lightning

NVIDIA · United States

2026
License
NVIDIA Open Model License (custom — read before commercial use)
Context
262k tokens
Sizes
30B-A3B hybrid Mamba/MoE (~3B active)
Fast local inferenceLong documents24 GB GPUs

August 2026. A hybrid Mamba-attention MoE: 31B of weights, roughly 3B active per token, so it answers at small-model speed. GGUF builds come from ggml-org and unsloth. The licence is custom, not OSI-approved.

LFM2.5

Liquid AI · United States

2026
License
Custom Liquid AI licence (check the model card)
Context
131k tokens
Sizes
2.6B
Edge deploymentsPhonesLow-memory laptops

July 2026. A convolution-attention hybrid built for edge hardware, covering 16 languages at 2.6B parameters. Liquid publishes its own GGUF files, which is why it reached the llama.cpp crowd faster than most small models.

France (EU)2 models

Mistral Small 4

Mistral AI · France (EU)

2026
License
Apache 2.0
Context
256k tokens
Sizes
119B MoE (6.5B active)
ReasoningCoding agentsMultimodal workflowsEU-friendly deployments

Combines instruct, reasoning and coding modes in one multimodal MoE model. Low active parameters help compute speed, but the full weights still determine memory use.

Mistral Medium 3.5

Mistral AI · France (EU)

2026
License
Modified MIT / repository terms
Context
256k tokens
Sizes
128B dense
EU-friendly deploymentsCodingReasoningLong context

Dense 128B multimodal model. Local deployment is memory-heavy and generally workstation-class.

European Union1 model

EuroLLM-22B

EuroLLM Consortium · European Union

2025
License
Apache 2.0
Context
32k tokens
Sizes
1.7B · 9B · 22B
EU language coveragePublic-sector AIResearch

Transparent European family covering all 24 EU official languages plus additional languages.

United States (non-profit)1 model

Olmo 3 / 3.1

Allen Institute for AI · United States (non-profit)

2025
License
Apache 2.0
Context
65k tokens
Sizes
7B · 32B (Instruct and Think variants)
Reproducible researchAuditable trainingEducation

Weights, data, training code and intermediate checkpoints are public; transparency is the differentiator. Olmo 3.1 (December 2025) added reasoning-tuned Think variants at 7B and 32B, so the transparency argument costs far less capability than it did with OLMo 2.

Frontier open weights

The giants you (probably) can't run at home →

Kimi K3, GLM-5.3, Llama 4 Maverick, DeepSeek V4 Pro: what running them actually takes, and the runnable sibling from each family.

Trending · refreshed weekly

What the community is downloading this week →

The live top 16 from Hugging Face, ranked by downloads, likes and recency. Auto-updated every Monday.