Qwen 3.8 27B
Apache 2.0 · 2026
Dense, Apache 2.0 and natively multimodal. The current default choice for a 24 GB card; bartowski publishes an imatrix build too.
ollama run qwen3.8:27bPicker · Hardware-aware recommender
Different AI models need different amounts of memory. A small model fits on a phone, a frontier-grade model needs a workstation. Tell us what hardware you have and what you want to do, and the tool will suggest the best options that will actually run on your machine. The form updates as you type. Nothing is sent to a server.
On a Mac: click the Apple menu → About This Mac. The number next to "Memory" is your unified memory. Pick "Apple Silicon" in the form below.
On Windows with an NVIDIA GPU: open Task Manager → Performance → GPU. The number next to "Dedicated GPU memory" is your VRAM. Pick "NVIDIA GPU" in the form below.
On Windows or Linux without a discrete GPU: pick "CPU only" and enter your system RAM. AI will run slowly, but it will run.
Not sure which terms apply? Open the glossary in a new tab.
Find it under Apple menu → About This Mac → Memory. Common configurations: 8, 16, 24, 32, 48, 64, 96, 128, 192 GB.
Computed from your specs minus a reasonable system overhead. Models that exceed this with a 15% safety margin are excluded from the recommendations.
6 options ranked by use-case fit and headroom.
Apache 2.0 · 2026
Dense, Apache 2.0 and natively multimodal. The current default choice for a 24 GB card; bartowski publishes an imatrix build too.
ollama run qwen3.8:27bApache 2.0 · 2026
Quantization-aware-trained q4_0, the only build Google publishes for this size.
ollama run gemma4:26bApache 2.0 · 2026
Quantization-aware-trained q4_0, the only build Google publishes for this size.
ollama run gemma4:31bApache 2.0 · 2026
ollama run qwen3.5:9bApache 2.0 · 2026
Prism ML's ternary requantization of Qwen3.6-27B: a 27B that fits where a 9B used to. Apache 2.0, and the whole language stack is genuinely ternary rather than low-bit with high-precision escape hatches. Vendor-reported: 95% of FP16 intelligence retained, 80.49 average across 15 thinking-mode benchmarks. Unverified by us — test it on your own workload before trusting it. The Q2_0 pack needs Prism ML's llama.cpp fork; Q2_g64 is the file for upstream llama.cpp.
# No verified Ollama tag in the RunLocal catalog yet; check ollama.com/library or the model card.Apache 2.0 · 2026
Prism ML's ternary rebuild of Qwen3.8-27B, September 2026: 5.95 GB (PTQ1_0) or 7.21 GB (PQ2_0) on disk, Apache 2.0. Vendor-reported 84.78 average across 14 thinking-mode benchmarks against 86.32 for FP16, unverified by us. Stock llama.cpp will not run either file: use Prism ML's fork. It is a reasoning model that thinks at length by default, so give it a large output limit (-n 16384) or the answer gets cut off mid-thought.
# No verified Ollama tag in the RunLocal catalog yet; check ollama.com/library or the model card.Every model in the catalog is paired with realistic memory estimates per quantization (Q4_K_M, Q5_K_M, Q8_0) at a moderate 8k context. The recommender computes your usable memory by subtracting a small system overhead (six gigabytes on Apple Silicon, two gigabytes on a discrete GPU, four gigabytes on CPU-only setups), then requires the chosen model to fit with a fifteen percent safety margin. Anything that does not fit lands in the excluded list below the results, with the reason printed out. The ranking that follows weights use-case fit most heavily, then quantization quality, then recency of the release, with a modest bonus for models that leave breathing room rather than filling the memory to the brim.
Memory estimates are rounded for clarity. Actual usage depends on context length, batch size, and which inference engine you run. If a model is on the edge of fitting, give it a try at a smaller context first.