English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Running a 744B-Parameter Model in 25GB RAM: How colibrì Does It with 1,300 Lines of C Code

Forum topic · ✨步子哥 · 2026-08-03

Summary

colibrì is a zero-dependency, ~1,300-line C inference engine that runs the GLM-5.2 mixture-of-experts model (744 billion parameters) on a laptop with only 25GB of RAM and no GPU. It exploits MoE sparsity: only ~5.4% of parameters (about 40B) are activated per token, so only ~20GB of active weights are needed at any moment. colibrì streams 21,504 routed experts (~19MB each, int4) on demand from NVMe SSD, bypassing the OS page cache with expert-level per-layer LRU caching, async readahead, and router-lookahead prefetching. It implements MLA attention (57x KV-cache compression), persistent KV caches for restartable sessions, and MTP speculative decoding reaching 2.2-2.8 tokens per forward pass. Benchmarks show 0.05-0.1 tok/s on a 25GB machine (cold start) up to 5.8-6.8 tok/s on 6x RTX 5090, with token-exact correctness verified against reference implementations. The project demonstrates that MoE decouples total parameter count from memory requirements, making frontier open-weight models runnable on consumer hardware.

Running a 744B-Parameter Model in 25GB RAM: How colibrì Does It with 1,300 Lines of C Code

The Impossible Equation

GLM-5.2 has 744 billion parameters. With int4 quantization (4 bits per parameter), that's 372 GB. A typical laptop has 25 GB of RAM — a 15x gap. Before July 10, 2026, every LLM inference engine answered this equation the same way: impossible. Either buy an H100 GPU cluster or pay for cloud API access.

Then a developer called JustVugg released colibrì — 1,300 lines of C code, zero dependencies — running GLM-5.2 on a 25 GB RAM, GPU-less laptop. It hit 453 points on Hacker News in three hours. This isn't magic; it's exploiting a loophole.

The Loophole: MoE Sparsity

GLM-5.2 is a Mixture-of-Experts (MoE) model with 256 expert subnetworks, activating 8+1 per layer. Each token only uses ~40 billion parameters; the other 704 billion sleep. 40B × 0.5 bytes = 20 GB — mathematically feasible.

The engineering challenge: the router dynamically decides which experts each token needs, and those experts may live on SSD. colibrì's core problem was making disk reads fast enough to not stall inference.

Three-Tier Storage Hierarchy

  • VRAM (GPU): optional, fastest
  • RAM (~9.9 GB resident): dense layers (attention, shared experts), embeddings, and an LRU cache of hot experts
  • NVMe SSD (~370 GB): 21,504 routed experts, ~19 MB each (int4), streamed on demand
  • Only ~11 GB of parameters change between tokens — the routed experts. The other ~34 GB are dense layers permanently in memory.

    Why llama.cpp's mmap Falls Behind

    llama.cpp uses mmap, which works perfectly for dense models with sequential access patterns. But MoE access is random — 8 experts per token scattered across 370 GB. Three fatal mmap problems:

    1. The OS page cache doesn't know expert boundaries. A 19 MB expert spans ~4,750 4KB pages; OS LRU can evict half an expert. 2. OS LRU ignores expert popularity. 71.6% of routing is predictable, but the OS can't make expert-level cache decisions. 3. No async prefetching. mmap readahead is sequential; MoE access is random, so every token waits on disk I/O.

    colibrì's solution — bypass OS page cache policy and manage experts itself:

  • Per-layer LRU caches: eviction in whole-expert (19 MB) units
  • OS page cache as a free L2: still works underneath as a second layer
  • Async expert readahead: disk I/O for likely-next experts issued during current-token compute
  • Router-lookahead prefetch (experimental): exploiting 71.6% routing predictability
  • Pinned hot-store: frequently-used experts permanently resident in RAM
  • MLA Attention: 57x KV-Cache Compression

    GLM-5.2 uses Multi-head Latent Attention (MLA), compressing KV into low-dimensional latents: from 32,768 floats per token down to 576 — a 57x compression. colibrì persists the KV cache to disk (.coli_kv files), so sessions survive restarts with zero re-prefill and byte-identical state.

    MTP Speculative Decoding: 2.2-2.8 Tokens per Forward

    colibrì implements GLM-5.2's native Multi-Token Prediction (MTP) head for speculative decoding, reaching 2.2-2.8 tokens/forward. Two hard-won rules documented in the README:

    > MTP head must be int8 (int4 heads collapse to 0-4% acceptance, #8)

    And: draft and verification must compute the same function (SPEC_PIN=1 pins both to the same kernel family — see issue #163).

    Honest Benchmarks: Slow, but Correct

    | Hardware | Speed | Notes | |------|------|------| | 6× RTX 5090 (fully resident) | 5.8-6.8 tok/s | TTFT ~13s | | 128 GB CPU desktop | ~1.8 tok/s | warm cache | | M5 Max (128 GB, Metal) | 1.06-1.83 tok/s | | | Ryzen AI 9 HX 370 (128 GB) | 0.37 tok/s | | | 25 GB dev machine | 0.05-0.1 tok/s | cold start, honest baseline |

    Cloud H100 inference runs 30-50 tok/s. colibrì is 10-100x slower — but correct: forward passes are token-exact verified, 32/32 matching against the transformers reference implementation.

    > This is not fast. It is a 744B frontier-class model answering correctly on a machine that costs less than one H100 fan.

    Why This Matters

  • Technical: colibrì redefines "loading a model" — from "read all weights into RAM" to "build an on-demand streaming pipeline." The requirement drops from RAM ≥ model size to RAM ≥ activated parameters.
  • Industry: frontier open-weight models (GLM-5.2 is MIT-licensed from Zhipu AI) are no longer locked in datacenters. Privacy-sensitive deployments become feasible on local hardware.
  • Philosophy: from the README: "Not renting intelligence behind an API — holding it." The web dashboard visualizes all 19,456 experts as a living cortex — color is storage tier, brightness is routing heat, activated experts flash white.
  • Three Engineering Principles

    1. Know what the OS doesn't. Generic abstractions lose domain knowledge; when you have domain knowledge (expert boundaries, popularity), bypassing generic abstractions yields order-of-magnitude gains. 2. Measure quality, don't assume it. Every number in the README cites its source (issue numbers, experiment logs). 3. Defaults are scar tissue. Good defaults are "the configuration new users are least likely to trip over," not "the most generic configuration."

    The Deeper Insight: MoE Decouples Parameters from Memory

    Dense models have 100% parameter utilization per token, so loading requires all parameters in RAM. MoE models use ~5.4% per token — meaning 94.6% of parameters are cold data at any moment. Cold data can live on SSD (or tape, or cloud storage) as long as it's readable on demand.

    MoE decouples total parameter count from memory requirements. Parameter count sets the model's capability ceiling; memory sets deployment cost. colibrì proves these can differ by 15x — given fast enough storage and smart enough caching.

    Roadmap: From "Runs" to "Usable"

  • Better prefetching (routing predictability from 71.6% toward 90%+)
  • GPU-accelerated matmul for dense layers (10-50x speedup potential)
  • Multithreaded expert loading to overlap I/O waits
  • Expert compression via structured pruning of never-activated experts
  • Even as a proof of concept at v1.1.0, colibrì proves a key thesis: frontier models don't need frontier hardware — only frontier engineering.

    ---

  • Project: JustVugg/colibri (Apache 2.0)
  • Model: GLM-5.2 (Zhipu AI, MIT license)
  • Engine: pure C, zero dependencies, ~1,300 lines of core code
  • Video: Devsplainers original video
  • FAQ

    Q1: Who is this content for?

    Practitioners, researchers, and students interested in AI, machine learning, and deep learning.

    Q2: What are the core points?

  • The impossible equation: why 372 GB doesn't fit in 25 GB
  • The loophole: MoE architecture sparsity (only 5.4% of parameters active per token)
  • Three-tier storage: colibrì's memory architecture with expert-level caching
Q3: Is there open-source code?

Yes — see the project link above.

Tags

#colibri#moe#inference#glm#quantization#speculative-decoding#llm#c

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178503915