RunLocal

Quantization · 8 min · September 29, 2026

Bonsai 2 fits a 27B in six gigabytes, on Prism ML's terms

Ternary-Bonsai-2-27B-gguf was second in our trending snapshot a week ago and first in yesterday's, and as of this morning the repository shows about 3.6 million downloads, most of them in September. It is a fair measure of what people wanted from Qwen3.8-27B all along: the same model, in a file that fits the machine they already have.

What is in the file

Every weight in the language model is one of three values, minus one, zero or plus one, with a shared 16-bit scale for each group of 128. Prism ML counts that as 1.72 bits per weight across the whole model, and unlike most “2-bit” builds the label is close to the truth: embeddings, attention, MLP and output head are all ternary, and only about 26 million parameters (the recurrent state of the linear-attention layers and the norms) stay at higher precision. The weights are also stored after a Hadamard rotation, a trick that spreads outliers across a block so that three values are enough to describe it. The runtime has to apply the matching rotation to the activations, and that detail is where the practical trouble starts.

Two packings ship. PTQ1_0 packs the trits densely and weighs 5.95 GB. PQ2_0 stores each trit in a 2-bit slot and weighs 7.21 GB, trading a little memory for cheaper unpacking. The vision tower is a separate 0.63 GB file you only need for images.

How much it costs in quality, according to the people selling it

Prism ML reports a 14-benchmark average of 84.78 in thinking mode, against 86.32 for the full-precision Qwen3.8-27B and 85.18 for unsloth's UD-Q4_K_XL at 17.6 GB. The comparison that matters more is with IQ2_XXS, the conventional 2-bit build, which lands at 7.27 GB, almost the same size, and averages 72.59. The gap is concentrated exactly where you would worry about it: on AIME26 the conventional build falls to 57.5 and Bonsai 2 holds 95.83, on LiveCodeBench it is 56.4 against 90.07. Meanwhile IQ2_XXS still scores about 89 on MMLU-Redux, which is why a quick chat with a heavily quantized model can feel fine while it quietly fails at the long reasoning chains you downloaded it for.

All of these are vendor figures, run by Prism ML on their own harness. We have not reproduced them and neither, as far as we can find, has anyone independent, so the right way to read them is as a reason to spend an evening testing the model on your own prompts rather than as a settled result.

The runtime is the catch

The model card is blunt about it: stock llama.cpp will not run these files. It rejects PQ2_0 and PTQ1_0 as unknown types, and a file in the older Q2_0 layout would load without complaint and generate garbage, because upstream llama.cpp has no idea the weights are rotated. You need the PrismML-Eng/llama.cpp fork, either a prebuilt binary from its releases page or a build from source, and for anything beyond a plain chat Prism ML points you to their Bonsai-demo repository as the reference setup.

That rules out, for now, the tools most people start with. Ollama and LM Studio are built on upstream llama.cpp, so neither will load Bonsai 2 until the kernels are merged upstream or the vendors bundle the fork. We have added the model to the picker anyway, because on a 12 GB card or a 16 GB Mac nothing else in the catalogue comes close, but the command it prints now starts with cloning and building the fork. While doing that we found that our entry for the first Ternary Bonsai had the same problem: its Q2_0 file needs the fork too, and our command assumed upstream. That is fixed, and the note on that entry now says that Q2_g64 is the file for upstream llama.cpp.

Things that will trip you up

It is a reasoning model and it thinks at length by default, at an effort level the chat template calls xhigh. With a small output limit it spends the whole budget thinking and returns an empty answer, which is the most common complaint in Prism ML's own known-issues page. Give it -n 16384 or more and a context of at least 65k. If you want shorter answers, send reasoning_effort: "medium"; the template accepts low, medium and xhigh, and asking for high currently gets you an HTTP 500 from the server.

On speed, Prism ML measures about 28 tokens a second on an M5 Pro laptop with PQ2_0, and about 91 on an RTX 4090 with PTQ1_0. Which packing is faster depends on the card: PTQ1_0 wins on Ada-generation GPUs and the L4, where memory bandwidth is the limit, and PQ2_0 wins on Blackwell, Hopper and Ampere. On a Mac, PQ2_0 is the one they have measured.

Who should bother

12 GB card or 16 GB Mac: this is the first time a 27B-class model has been a realistic option at this size, and the reasoning scores are close enough to the 4-bit build that the trade is worth trying. Budget an evening for building the fork. On an 8 GB card the file fits but the context and runtime buffers do not leave enough room, which is why the picker stops offering it there.

24 GB card or 32 GB Mac: the ordinary Q4_K_M of Qwen3.8-27B at about 17 GB still fits, runs on everything, and gives up nothing to a custom runtime. Bonsai 2 makes sense here only if you need the freed memory for a very long context or a second model alongside.

Anything running Ollama or LM Studio for a household or a team: wait. A fork is fine for one person who likes building things, and a poor foundation for a setup other people depend on.

At the other end of the month

September also brought MiniCPM5-2B from OpenBMB, which goes the other way: a dense 2.5B model trained small from the start rather than compressed afterwards, Apache 2.0, with a 131k context and official GGUF files that run on stock llama.cpp. Its Q4_K_M is about 1.6 GB. OpenBMB says it competes with 4B-class models on coding and maths, another vendor claim to check yourself. It is in the picker now as the step up from MiniCPM5-1B for phones and old laptops, and it is the model we would try first on a machine where even Bonsai 2's six gigabytes are too many.