RunLocal

Picker · Hardware-aware recommender

Which AI model can your computer actually run?

Different AI models need different amounts of memory. A small model fits on a phone, a frontier-grade model needs a workstation. Tell us what hardware you have and what you want to do, and the tool will suggest the best options that will actually run on your machine. The form updates as you type. Nothing is sent to a server.

How do I find my specs?

On a Mac: click the Apple menu → About This Mac. The number next to "Memory" is your unified memory. Pick "Apple Silicon" in the form below.

On Windows with an NVIDIA GPU: open Task Manager → Performance → GPU. The number next to "Dedicated GPU memory" is your VRAM. Pick "NVIDIA GPU" in the form below.

On Windows or Linux without a discrete GPU: pick "CPU only" and enter your system RAM. AI will run slowly, but it will run.

Not sure which terms apply? Open the glossary in a new tab.

Find it under Apple menu → About This Mac → Memory. Common configurations: 8, 16, 24, 32, 48, 64, 96, 128, 192 GB.

Available memory for the model: 26.0 GB

Computed from your specs minus a reasonable system overhead. Models that exceed this with a 15% safety margin are excluded from the recommendations.

Recommended models

6 options ranked by use-case fit and headroom.

#1Alibaba · China

Qwen 3.8 27B

Apache 2.0 · 2026

Score
94/100
Quantization
Q5_K_M
Memory fit
20.0 GB / 26.0 GB
Context
262k tokens
Moderate (~20–50 tok/s)

Dense, Apache 2.0 and natively multimodal. The current default choice for a 24 GB card; bartowski publishes an imatrix build too.

ollama run qwen3.8:27b
#2Google DeepMind · United States

Gemma 4 26B-A4B MoE

Apache 2.0 · 2026

Score
93/100
Quantization
Q4_0
Memory fit
15.0 GB / 26.0 GB
Context
256k tokens
Fast (~50+ tok/s on a single user)

Quantization-aware-trained q4_0, the only build Google publishes for this size.

ollama run gemma4:26b
#3Google DeepMind · United States

Gemma 4 31B dense

Apache 2.0 · 2026

Score
93/100
Quantization
Q4_0
Memory fit
18.0 GB / 26.0 GB
Context
256k tokens
Moderate (~20–50 tok/s)

Quantization-aware-trained q4_0, the only build Google publishes for this size.

ollama run gemma4:31b
#4Alibaba · China

Qwen 3.5 9B

Apache 2.0 · 2026

Score
90/100
Quantization
Q8_0
Memory fit
10.0 GB / 26.0 GB
Context
262k tokens
Fast (~50+ tok/s on a single user)
ollama run qwen3.5:9b
#5Alibaba · China

Qwen 3.6 27B (Ternary Bonsai, 1.71-bit)

Apache 2.0 · 2026

Score
90/100
Quantization
Q2_g64
Memory fit
8.5 GB / 26.0 GB
Context
262k tokens
Moderate (~20–50 tok/s)

Prism ML's ternary requantization of Qwen3.6-27B: a 27B that fits where a 9B used to. Apache 2.0, and the whole language stack is genuinely ternary rather than low-bit with high-precision escape hatches. Vendor-reported: 95% of FP16 intelligence retained, 80.49 average across 15 thinking-mode benchmarks. Unverified by us — test it on your own workload before trusting it. The Q2_0 pack needs Prism ML's llama.cpp fork; Q2_g64 is the file for upstream llama.cpp.

# No verified Ollama tag in the RunLocal catalog yet; check ollama.com/library or the model card.
#6Alibaba · China

Qwen 3.8 27B (Ternary Bonsai 2, 1.72-bit)

Apache 2.0 · 2026

Score
90/100
Quantization
PQ2_0
Memory fit
8.5 GB / 26.0 GB
Context
262k tokens
Moderate (~20–50 tok/s)

Prism ML's ternary rebuild of Qwen3.8-27B, September 2026: 5.95 GB (PTQ1_0) or 7.21 GB (PQ2_0) on disk, Apache 2.0. Vendor-reported 84.78 average across 14 thinking-mode benchmarks against 86.32 for FP16, unverified by us. Stock llama.cpp will not run either file: use Prism ML's fork. It is a reasoning model that thinks at length by default, so give it a large output limit (-n 16384) or the answer gets cut off mid-thought.

# No verified Ollama tag in the RunLocal catalog yet; check ollama.com/library or the model card.
4 models excluded
  • DeepSeek V4 Flash 0731 (304B-A13B MoE) — Smallest listed quant (~91 GB before safety margin) exceeds available inference memory (26.0 GB).
  • MiMo V2.6 Flash Flash-RL (309B-A15B MoE) — Smallest listed quant (~128 GB before safety margin) exceeds available inference memory (26.0 GB).
  • Mistral Small 4 119B-A6B MoE — Smallest listed quant (~68 GB before safety margin) exceeds available inference memory (26.0 GB).
  • Mistral Medium 3.5 128B dense — Smallest listed quant (~75 GB before safety margin) exceeds available inference memory (26.0 GB).

How the tool decides

Every model in the catalog is paired with realistic memory estimates per quantization (Q4_K_M, Q5_K_M, Q8_0) at a moderate 8k context. The recommender computes your usable memory by subtracting a small system overhead (six gigabytes on Apple Silicon, two gigabytes on a discrete GPU, four gigabytes on CPU-only setups), then requires the chosen model to fit with a fifteen percent safety margin. Anything that does not fit lands in the excluded list below the results, with the reason printed out. The ranking that follows weights use-case fit most heavily, then quantization quality, then recency of the release, with a modest bonus for models that leave breathing room rather than filling the memory to the brim.

Memory estimates are rounded for clarity. Actual usage depends on context length, batch size, and which inference engine you run. If a model is on the edge of fitting, give it a try at a smaller context first.