Paul Visciano

Sci-Fi Labs · Field notes

Bonsai: the 27B Model That Fits in 16 GB

A 27B model answers on this 16 GB laptop. That model is Bonsai. Memory you own, models you run.

Bonsai local models on a 16 GB laptop
27B-class reasoning. Weights that fit. Offline is the proof.

A 27B model is answering on this 16 GB laptop. Not a demo tab. Not a remote API. The weights are on disk. The reply streams here.

That model is Bonsai. Open-weight builds from PrismML. Full-precision 27B wants on the order of 54 GB just for weights. Ordinary quant lands in the mid-teens. Neither leaves room for a browser, an editor, and a spatial canvas. Bonsai is the path that fits.

Bonsai is what makes the constraint honest instead of aspirational.

What Bonsai is

Bonsai uses end-to-end 1-bit ({−1, +1}) or ternary ({−1, 0, +1}) weights across embeddings, attention, MLPs, and the LM head. Same architecture class. Roughly 14× smaller than FP16.

The flagship Bonsai 27B is based on Qwen3.6 27B. Multi-step reasoning, tool calling, vision input, 262K-token context. A footprint a consumer laptop can hold:

VariantWeightsBest for
1-bit Bonsai 27B~3.9 GBTight memory, phones, lean laptops
Ternary Bonsai 27B~5.9 GBEveryday laptops, higher quality retention

PrismML reports the ternary build retaining roughly 95% of the full-precision baseline across their suite, and the 1-bit build around 90%. Both ship under Apache 2.0. Useful for real agent loops. Not a toy.

A remote vault of weights versus a 16 GB laptop on a desk
54 GB on paper. 3.9 GB on this desk.

Why a 16 GB laptop suddenly works

On this M2 Pro the budget is unforgiving. Right now, writing with everything open: 12.6 GB spoken for, 2.7 GB free. Bonsai has the weights. Brave and VS Code are the real tax. There is no spare room for a second large model.

ProcessRAMNotes
Bonsai-27B (llama-server)~4 GBMetal GPU; live footprint beside the rest of the desk
Brave Browser3.6 GBThe real memory tax
VS Code3.3 GBEditor open while the model answers
System (wired + active + compressed)12.6 GB2.7 GB available

That is the difference between “local AI” as a slide and local AI as something you open every day.

How it fits Knowledge Graph

Knowledge Graph is a private spatial canvas: photos, conversations, and time on a machine you own. The model under it is not a brochure claim. It is a running stack.

Right now: llama.cpp (PrismML fork) serves Bonsai. Open WebUI is the temporary conversation surface. Whisper hears. Kokoro JS speaks in the browser. Next swap is not another engine. It is the Knowledge Graph UI already built. Same models. Canvas instead of chat bubbles.

Published sites are demos. The product is the copy on your machine. Offline is the proof. Memory you own, models you run.

Getting started

More from Sci-Fi Labs

Running Local AI on 16 GB → Knowledge Graph → Reclaim your data → Where is Paul? →