Editorial
Notes on an ecosystem that moves faster than the press releases.
Long-form coverage of open source AI, written for readers who already know the basics and want the part the announcement omitted.
Bonsai 2 fits a 27B in six gigabytes, on Prism ML's terms
The most downloaded model on the Hub this month is a ternary rebuild of Qwen3.8-27B that fits a laptop. The numbers hold up better than the last round, and the catch is the runtime.
September 29, 20268 min readQuantization · Qwen · ToolsGLM-5.3 is two different models sharing a version number
Z.ai shipped a 753B post-trained GLM-5.2 under a new licence and a 321B model built from scratch under MIT, and gave them the same name. The smaller one is the one to watch, once llama.cpp catches up.
September 29, 20267 min readGLM · Licensing · AnalysisDwarfStar, and the case for an engine that runs one model
The author of Redis wrote an inference engine in C that runs essentially a single model, and it does things the general-purpose runtimes cannot. What ds4 gets right, and what it costs in hardware.
August 27, 202610 min readTools · Inference · QuantizationQwen 3.8 27B is the release that actually landed
Alibaba published two models in the same generation this month. One has four million downloads, the other twenty-seven thousand. The difference is not capability.
August 21, 20269 min readQwen · Licensing · Analysis3.5 million downloads, almost no footprint: the Ornith-1.0-35B question
A 35B MIT-licensed model from an account we'd never heard of just outran everything else on Hugging Face's trending list except the evergreen giants. Here is what we checked, what we couldn't verify, and why it isn't in the directory.
August 20, 20268 min readVerification · Directory · TrendingDeepSeek V4 Flash-0731, and the case for dated checkpoints
A quiet mid-cycle refresh went straight to the top of our weekly trending list. What a dated checkpoint tag means if you already run the model it's replacing.
August 6, 20268 min readDeepSeek · Trending · AnalysisKimi K3 changes the definition of 'open'
A 2.8-trillion-parameter model with public weights that almost nobody can run. What the largest open weight ever released actually means for people who run AI locally.
July 18, 20269 min readFrontier · Kimi K3 · AnalysisChoosing a GGUF quantization without lying to yourself
Q4, Q5, Q8 and the rest of the GGUF zoo, with a practical decision rule that holds up across hardware classes.
May 15, 202610 min readQuantization · llama.cpp · GuideApple Silicon or NVIDIA for local LLMs in 2026
Unified memory, raw VRAM, and the workloads where each wins. A practical comparison that goes beyond benchmark snippets.
May 14, 202612 min readHardware · Comparison · Apple SiliconWhy openSUSE is a serious option for running AI locally
Rolling releases, immutable variants, and an honest line between community Linux and paid enterprise. A practical look at when openSUSE earns its place in an AI stack.
May 13, 20269 min readopenSUSE · Linux · InfrastructureThe state of open weights in May 2026
Five frontier-class releases in the last thirty days, three of them from Chinese labs. A short tour of where the field actually is.
May 12, 20269 min readModels · Ecosystem · AnalysisWhich local inference engine should you actually use
Ollama, llama.cpp, LM Studio and vLLM solve different problems. A practical map of when to reach for which, and why it matters more than benchmark numbers.
May 5, 202611 min readTools · Inference · Guide
