On July 28, the open-source research project Deltafin (gavamedia/deltafin on GitHub) did something that sounds like bragging: it got the 2.8-trillion-parameter MoE model Kimi K3 running on a first-generation M1 Max (64GB unified memory, 400GB/s memory bandwidth). The cost: a median speed of 0.0687 token/s — about 14.6 seconds per token.
This is not a productivity tool. It is an engineering existence proof that "an MoE model far exceeding local memory can indeed run on consumer hardware." What it does, and what it means, need to be examined separately.
How big is Kimi K3?
K3 is Moonshot AI's flagship open-source MoE model released in July. The numbers on its Hugging Face model card are staggering:
- Total parameters: 2.8T, with 104B activated per inference pass
- 896 experts; the router picks 16 per token, plus 2 shared experts
- 93 layers (1 dense + 69 KDA + 24 Gated MLA)
- Architecture baseline: Kimi Delta Attention (KDA) + Attention Residuals, ~2.5x scaling-efficiency improvement over K2
- Natively multimodal, 1M-token context window
- Quantization-aware training (MXFP4 weights + MXFP8 activations from the SFT stage onward)
- Full weights: ~1.56 TB across 96 safetensors shards
- Prefill first token: 28.0 s (24.9–37.9 s)
- Steady-state decode: 0.0687 token/s (0.0503–0.0779 across 6 runs)
- Token model time across 6 runs (including first step): 56.5 s
- Full wall-clock time for a new process: 64.1 s
- 07-25: Open-source engine runs Gemma 4 26B on any M-series Mac with 2GB of memory
- 07-19: 28.9M-parameter TinyStories runs on a $8 ESP32-S3
- 07-29: Deltafin runs 2.8T-parameter Kimi K3 on a 64GB Mac (0.0687 tok/s)
- 06-30: RedKnot forcibly splits the KV Cache to run Llama 3 on an M2
- Researchers: Deltafin is an engineering blueprint for "how to run an MoE model when memory is far smaller than the weights." Its architecture is not "download and go" but a complete design of "layer/expert/byte three-level caching + streaming scheduling." Read the README, and be sure to also read OPTIMIZATIONS.md.
- Mac users: Running Kimi K3 via Deltafin is mostly an experience, not productivity. But you can use it as a fully local OpenAI-compatible API server and point Cursor / Open WebUI at it. It is one of the cheapest open-source implementations of a "localized LLM backend."
- MoE framework builders: Deltafin's streaming expert-loading logic, native kernel choices (MPS / CUDA / CPU Metal paths), and speculative-decoding integration are directly reusable engineering templates.
- Products claiming "runs model X locally": Deltafin is a reminder that between "can run" and "can run usefully" lies a gap of 30x+ token/s. This line of work is still chasing ceilings; don't market it as production-grade capability.
- Deltafin GitHub repository (gavamedia / deltafin): https://github.com/gavamedia/deltafin
- Deltafin optimization details (OPTIMIZATIONS.md, same repo): https://github.com/gavamedia/deltafin/blob/main/OPTIMIZATIONS.md
- Kimi K3 deep technical analysis (M1 Max deployment, 0.0687 tok/s measurement): https://funian.blog.csdn.net/article/details/163314317
- Deltafin full bilingual report (local deployment of a 2.8T model): http://freeai.help/blog/28-wan-yi-can-shu-sai-jin-yi_zh
- Deltafin Toutiao coverage 07-29 (full deployment vs streaming deployment): https://www.toutiao.com/w/1872023784777740/
The machine has 64GB of RAM. The model is 1.56 TB. More than an order of magnitude apart.
Deltafin's three-layer approach
Deltafin's README describes not a single trick but a three-layer combination:
Layer 1: MXFP4 weights + streaming expert loading. Each MoE token only activates 16 experts + 2 shared experts, not all 896. Deltafin keeps the 1.45 TB expert library on local NVMe and streams it to the GPU on demand in MXFP4-quantized form. This structure gives MoE models a physically real "read-on-demand" paradigm — previously, single-machine MoE inference required all experts resident in VRAM or host memory, which small machines simply cannot hold.
Layer 2: int8 spine + speculative decoding verification. Deltafin lets a small model act as a "draft," but unlike a typical Draft Model, the author explicitly states: "In testing, accepted results were exactly identical to the token-by-token greedy sequence, so this is a lossless optimization." In other words, the small model can only suggest work; confidence and model choice only change "what K3 verifies," never "what is allowed to be output." Under this engineering ethic, the small model functions more like "prefetch + cache warming" than "answering on K3's behalf."
Layer 3: OpenAI-compatible API server. Once running, Deltafin exposes an OpenAI-compatible API, so any client that can call the ChatGPT API (Cursor, Continue, Claude Code, local scripts) can point at Deltafin as "their own Kimi K3 backend."
Measured median reference values:
That means 100 tokens takes roughly 24 minutes, and a full K3 chain of thought (hundreds to thousands of tokens) runs tens of minutes to hours. Clearly unsuitable for production chat, but usable for: researching routing logic, testing streaming inference, stress-testing K3's internal structure, and benchmarking.
Its place in the "compute downsizing" landscape
Over the past three months, the "small machine runs big model" frontier has kept moving:
This trend shows that in 2026 H2, it is not just model scale that is growing — the capability ceiling of local inference is rising in parallel. The ceiling for "given my machine, how large a model can I run" has moved from "70B-scale dense" to "2.8T-scale MoE."
The key distinction between Deltafin and other work: it does not care about "fast," it cares about "runs at all." As a research project, its significance lies not in practical usability on any single Mac, but in proving that the MoE + quantization + streaming-loading combination is a viable engineering route on consumer hardware. Everyone working on "trillion-scale model localization" can take a baseline from it — this engineering problem has a solution.
Concrete advice for local-AI players and researchers
One-sentence summary
Deltafin is not a killer app for Kimi K3; it is an engineering existence proof that "large MoE models can run on small machines" — 2.8T parameters, 64GB of memory, 14.6 seconds per token.
Its real value is establishing an engineering baseline for "localizing very large MoE models": with suitable hardware (especially unified-memory Macs after M3/M4), anyone doing similar work can continue along this path.
The core of 2026 H2 compute downsizing is not "whose model is smaller" but "who can push hardware costs down to consumer level while keeping model scale." Deltafin proving that a 2.8T MoE model is a "reachable boundary" on a 64GB Mac is itself a concrete node in that story.
---
References