RunLocal

Directory dei modelli

Modelli AI che puoi scaricare ed eseguire tu stesso.

Ogni scheda qui sotto è una famiglia di modelli AI scaricabile gratuitamente ed eseguibile sul tuo computer. La scheda ti dice chi lo produce, quanto è grande, con quale licenza arriva e in cosa è bravo. Le note delle schede restano in inglese (dati del catalogo). Se sei alle prime armi, il glossario definisce ogni termine, oppure prova il picker per trovare quello adatto alla tua macchina.

China6 modelli

Qwen 3.8

Alibaba (Qwen team) · China

2026
License
Apache 2.0 on the 27B; custom Qwen licence on the 2.4T MoE
Context
262k tokens
Sizes
27B dense · 2.4T-A95B MoE
CodingMultimodal workflowsMultilingualLocal workstations

August 2026, and the most-liked open-weight release of the summer. The 27B is dense, Apache 2.0, natively multimodal, and had GGUF builds from unsloth and bartowski within days — which makes it the current default for a 24 GB card. The 2.4T MoE sibling is frontier-scale and carries a different licence.

DeepSeek V4 Flash

DeepSeek AI · China

2026
License
MIT
Context
1M tokens (65k native, extended 16x with YaRN)
Sizes
304B-A13B MoE (284B before the DSpark decoder)
Agentic codingLong documentsHigh-memory workstations

The most-downloaded open-weight model on the Hub, and the largest entry here that a single machine can still load. 256 routed experts with 6 active per token means roughly 13B parameters fire on any given token, so it answers far faster than 304B suggests. The 0731 checkpoint supersedes the June preview and carries a DSpark speculative-decoding module, which is where the extra weights over the 284B base come from. Weights ship natively in FP8, so a Q8 build costs only about 7 GB more than Q4.

Qwen 3.6

Alibaba (Qwen team) · China

2026
License
Apache 2.0 on official open-weight checkpoints
Context
262k native on 27B; long-context extensions available
Sizes
27B dense · 35B-A3B MoE · additional large MoE variants
CodingMultimodal workflowsMultilingualLocal workstations

The previous open-weight Qwen generation, still widely deployed and the base for most community quantization work, including the 1-bit and ternary Bonsai builds. Qwen 3.8 supersedes it for new installs.

Qwen 3.5

Alibaba (Qwen team) · China

2026
License
Apache 2.0
Context
262k native; selected variants support longer contexts
Sizes
2B · 9B · 27B · 35B-A3B MoE · 122B-A10B MoE · 397B-A17B MoE
MultimodalMultilingualCodeCost-sensitive deployments

Released in 2026. The smaller checkpoints remain excellent local choices when Qwen 3.6 is too large.

GLM-4.7-Flash

Z.ai (Zhipu) · China

2026
License
MIT
Context
202k tokens
Sizes
31B MoE
Bilingual EN/ZH workloadsAgentic codingWorkstation GPUs

The runnable member of the GLM family, and the counterpart to the frontier-scale GLM-5.2: MIT licensed, 200k context, mature GGUF ecosystem, and one of the strongest agentic-coding models that fits on a single workstation GPU.

MiniCPM5-1B

OpenBMB · China

2026
License
Apache 2.0
Context
128k tokens
Sizes
1B
On-device inferenceEdge deploymentsLow-memory hardware

Edge-first model for phones and low-memory laptops.

United States6 modelli

Gemma 4

Google DeepMind · United States

2026
License
Apache 2.0
Context
128k–256k depending on checkpoint
Sizes
E2B · E4B · 12B · 26B-A4B MoE · 31B dense
On-device inferenceMultimodalConsumer GPUsWorkstations

Multimodal family spanning edge through workstation hardware. The 26B-A4B model activates about 4B parameters per token while retaining a 26B memory footprint.

Phi-4 family

Microsoft Research · United States

2025
License
MIT
Context
Varies by checkpoint
Sizes
Phi-4 Mini 3.8B · Phi-4 14B · reasoning variants
Edge devicesReasoning per parameterCost-sensitive inference

Verified Microsoft Phi family. RunLocal does not list an unverified Phi-5 family.

Llama 4 (Scout & Maverick)

Meta AI · United States

2025
License
Llama Community License (custom)
Context
Up to 10M tokens (Scout)
Sizes
Scout 109B MoE · Maverick 400B MoE
Long-context retrievalCodebase-scale RAGGeneral reasoning

Meta's open-weight family. Custom licensing applies; Scout and Maverick are MoE models with much smaller active parameter counts than total weights.

gpt-oss (20B & 120B)

OpenAI · United States

2025
License
Apache 2.0
Context
131k tokens
Sizes
20B MoE (MXFP4 weights) · 120B MoE
General assistant workReasoningConsumer GPUs

OpenAI's open-weight release and, by download count, the most used open model of the past year. Ships natively in MXFP4, so the 20B fits a 16 GB card at the precision it was trained for. Apache 2.0, no usage caps, no licence acceptance step.

Nemotron 3.5 Lightning

NVIDIA · United States

2026
License
NVIDIA Open Model License (custom — read before commercial use)
Context
262k tokens
Sizes
30B-A3B hybrid Mamba/MoE (~3B active)
Fast local inferenceLong documents24 GB GPUs

August 2026. A hybrid Mamba-attention MoE: 31B of weights, roughly 3B active per token, so it answers at small-model speed. GGUF builds come from ggml-org and unsloth. The licence is custom, not OSI-approved.

LFM2.5

Liquid AI · United States

2026
License
Custom Liquid AI licence (check the model card)
Context
131k tokens
Sizes
2.6B
Edge deploymentsPhonesLow-memory laptops

July 2026. A convolution-attention hybrid built for edge hardware, covering 16 languages at 2.6B parameters. Liquid publishes its own GGUF files, which is why it reached the llama.cpp crowd faster than most small models.

France (EU)2 modelli

Mistral Small 4

Mistral AI · France (EU)

2026
License
Apache 2.0
Context
256k tokens
Sizes
119B MoE (6.5B active)
ReasoningCoding agentsMultimodal workflowsEU-friendly deployments

Combines instruct, reasoning and coding modes in one multimodal MoE model. Low active parameters help compute speed, but the full weights still determine memory use.

Mistral Medium 3.5

Mistral AI · France (EU)

2026
License
Modified MIT / repository terms
Context
256k tokens
Sizes
128B dense
EU-friendly deploymentsCodingReasoningLong context

Dense 128B multimodal model. Local deployment is memory-heavy and generally workstation-class.

European Union1 modello

EuroLLM-22B

EuroLLM Consortium · European Union

2025
License
Apache 2.0
Context
32k tokens
Sizes
1.7B · 9B · 22B
EU language coveragePublic-sector AIResearch

Transparent European family covering all 24 EU official languages plus additional languages.

United States (non-profit)1 modello

Olmo 3 / 3.1

Allen Institute for AI · United States (non-profit)

2025
License
Apache 2.0
Context
65k tokens
Sizes
7B · 32B (Instruct and Think variants)
Reproducible researchAuditable trainingEducation

Weights, data, training code and intermediate checkpoints are public; transparency is the differentiator. Olmo 3.1 (December 2025) added reasoning-tuned Think variants at 7B and 32B, so the transparency argument costs far less capability than it did with OLMo 2.

Frontier open weights

I giganti che (probabilmente) non puoi eseguire a casa →

Kimi K3, GLM-5.2, Llama 4 Maverick, DeepSeek V4 Pro: cosa serve davvero per eseguirli, e il fratello minore eseguibile di ogni famiglia.

Tendenze · aggiornato ogni settimana

Cosa sta scaricando la community questa settimana →

La top 16 live da Hugging Face, ordinata per download, likes e freschezza. Si aggiorna da sola ogni lunedì.