All Guides
GuideAI DevelopmentOctober 5, 2026

Run a 744B AI Model on a Laptop — No GPU

Take this guide with you.

Drop your email and we'll send the Run a 744B AI Model on a Laptop — No GPU straight to your inbox.

No spam. Just the guide. Unsubscribe anytime.

A field guide to Colibrì, which runs the mixture-of-experts GLM-5.2 on a laptop with no GPU: the always-on stack stays in RAM while routed experts stream from SSD, served through a local OpenAI-compatible endpoint — slow, but private and unbilled.

What's inside
  • Mixture-of-experts: only a fraction of parameters fire per token
  • The always-on stack in RAM, experts streamed from SSD
  • Pure C, no dependencies, no graphics card
  • Fully offline, with a local OpenAI-compatible endpoint
  • The catch: disk-bound speed and roughly 370 GB free

A Giant Model on Hardware You Already Own

Running a frontier-scale open model locally usually means a rack of GPUs. Colibrì takes a different route: it runs GLM-5.2 — a model the guide puts at 744 billion parameters — on an ordinary laptop with no graphics card at all, fully offline.


How It Fits

GLM-5.2 is a mixture-of-experts model, so it never uses all of its parameters at once — per the guide, only around 40 billion fire for any given token. Colibrì exploits that. It keeps the always-on part of the model resident in RAM (roughly 10 GB) and streams the specific experts each token needs from your SSD on demand, the way a video streams rather than downloading first.

It's written in pure C with no dependencies, and treats VRAM, RAM and disk as nothing more than speed tiers for the same weights.


Private, Local, Unbilled

Once the model is downloaded, nothing leaves your machine. coli serve exposes an OpenAI-compatible endpoint, so your apps can call a local model exactly as they would a cloud one — with no per-request bill.


The Catch, Up Front

This trades speed for freedom, and the guide is honest about it. It's disk-bound: on a laptop, expect well under a few tokens per second, with a fast NVMe drive helping a lot. And you need roughly 370 GB of free space for the model. It's for private, offline, unhurried work — not a drop-in replacement for a fast hosted model.


What's in the Guide

How mixture-of-experts streaming makes this possible, where each part of the model lives, the build-and-serve steps, the realistic speed and disk requirements, and a direct link to the repo.

The full field guide is in the PDF.


Get the Guide

Drop your email below and we'll send it straight to your inbox.

Don't leave without it.

Drop your email and get the Run a 744B AI Model on a Laptop — No GPU in your inbox.

No spam. Just the guide. Unsubscribe anytime.

Build Something Real

If you can describe it, you can build it.