A Giant Model on Hardware You Already Own
Running a frontier-scale open model locally usually means a rack of GPUs. Colibrì takes a different route: it runs GLM-5.2 — a model the guide puts at 744 billion parameters — on an ordinary laptop with no graphics card at all, fully offline.
How It Fits
GLM-5.2 is a mixture-of-experts model, so it never uses all of its parameters at once — per the guide, only around 40 billion fire for any given token. Colibrì exploits that. It keeps the always-on part of the model resident in RAM (roughly 10 GB) and streams the specific experts each token needs from your SSD on demand, the way a video streams rather than downloading first.
It's written in pure C with no dependencies, and treats VRAM, RAM and disk as nothing more than speed tiers for the same weights.
Private, Local, Unbilled
Once the model is downloaded, nothing leaves your machine. coli serve exposes an OpenAI-compatible endpoint, so your apps can call a local model exactly as they would a cloud one — with no per-request bill.
The Catch, Up Front
This trades speed for freedom, and the guide is honest about it. It's disk-bound: on a laptop, expect well under a few tokens per second, with a fast NVMe drive helping a lot. And you need roughly 370 GB of free space for the model. It's for private, offline, unhurried work — not a drop-in replacement for a fast hosted model.
What's in the Guide
How mixture-of-experts streaming makes this possible, where each part of the model lives, the build-and-serve steps, the realistic speed and disk requirements, and a direct link to the repo.
The full field guide is in the PDF.
Get the Guide
Drop your email below and we'll send it straight to your inbox.