What does the person who built Redis do next?
Salvatore Sanfilippo (antirez)'s answer: DwarfStar (ds4), a deliberately narrow local inference engine. It's trending on GitHub at +385 stars/day — unassuming at first glance, but it's the Redis creator's new answer twenty years later.
The Redis Philosophy, Continued
Redis succeeded on the philosophy of "do one thing extremely well" — not a database that does everything, but the fastest in-memory cache. Twenty years later, antirez brings that philosophy to inference engines:
> DwarfStar is self-contained and deliberately narrow, not a general GGUF runner.
Deliberately narrow — not a general GGUF runner. It supports only three models: DeepSeek V4 Flash, GLM 5.2, and DeepSeek V4 PRO. But it runs those three to the extreme.
Why Not Just Use vLLM or llama.cpp?
The local inference engine space is crowded. antirez's answer: general GGUF runners make too many compromises to support every model. vLLM has even dropped support for older GPUs (Ada Lovelace architecture), turning many enterprises' L40S cards into scrap metal.
ds4's opportunistic strategy: follow the best open weights, run only those models, but run them fastest and best. Models are replaceable — when better open weights appear, old ones can be swapped out.
Technical Highlights
Backends
- Metal: primary target for Macs with 96GB+ RAM; smaller machines use SSD streaming
- NVIDIA CUDA: including multi-GPU systems and DGX Spark
- ROCm: Strix Halo systems (e.g., Framework Desktop)
- 120 t/s aggregate generation throughput
- 2000 t/s prefill
- SSD streaming: stream from SSD when RAM is insufficient — even a MacBook can run it
- Tensor parallelism: two MacBook M5 Max / M3 Ultra machines running 4-bit DeepSeek Flash or GLM 5.2 over RDMA
- Pipeline parallelism: chain multiple systems to pool RAM for larger models
- Micro batching: separate batching for decode and prefill
- colibrì (744-billion-parameter model in 1,300 lines of C): small, precise C code still has a place in AI infrastructure
- Euclid-MCP (offloaded reasoning): don't make one general system do everything; let specialized tools do specialized things
- vLLM dropping old GPUs: the compromises of general systems are the opportunities of specialized ones
- General GGUF runner = Swiss Army knife: does everything, excels at nothing
- ds4 = scalpel: makes only one kind of cut, but the most precise
- Redis back then = fastest in-memory cache, never touching OLTP
- ds4 now = fastest for DeepSeek Flash, no other models
Measured Performance
Multi-user session test on 8x NVIDIA L40S cards:
This turns "old cards vLLM no longer supports" into a multi-user LLM server. For enterprises, this may be ds4's most practical selling point — no need to throw away existing GPUs.
Other Features
Self-Contained
Model loading, prompt rendering, tool calling, KV state, HTTP server, coding agent — all built and tested together. Not an assembly of components, but a unified engineering effort.
A Candid Statement on AI-Assisted Development
ds4's README contains a rare, candid disclosure:
> This software is developed with strong assistance from GPT 5.5, 5.6, Claude Fable and with humans leading the ideas, testing, and debugging. We say this openly because it shaped how the project was built. If you are not happy with AI-developed code, this software is not for you.
antirez doesn't hide the AI assistance — he openly acknowledges that it shaped the project. Humans lead ideas, testing, and debugging; AI leads code generation. It's a new engineering paradigm: human architect + AI programmer.
Credits to llama.cpp
ds4 doesn't link GGML, but the README states clearly:
> This project would not exist without llama.cpp and GGML.
Certain source-level snippets (GGUF quantization layout, CPU quant/dot logic, specific kernels) are retained or adapted under the MIT license. antirez even keeps the GGML author's copyright notice in the LICENSE file.
This is open-source community interaction at its ideal — standing on predecessors' shoulders, acknowledging them openly, preserving copyright. Not a fork, but an homage-style independent implementation.
Engineering Insight: The Victory of Deliberate Narrowness
ds4's design philosophy resonates with a series of recent projects:
Together these point to one theme: in the "bigger and stronger" AI narrative, "deliberately narrow" is an underrated strategy.
An Analogy
Conceptual Lineage: Division of Labor Beats Unification
ds4 belongs to the "division of labor beats unification" lineage — rather than one inference engine for all models, specialized engines for specialized models. It shares a family resemblance to octopus-style DNA pretraining + RNA inference-time compute, Euclid-MCP's LLM + Prolog split, and Rebucca's small-model + large-model verification approach.
Conclusion
antirez chose a path opposite to the mainstream. In the inference engine arms race of "support more models, run bigger parameters," ds4's answer is "support fewer models, run them faster." Redis proved that "deliberate narrowness" is a winning strategy in databases; ds4 is testing whether it holds for inference engines.
If it succeeds, it won't just add another inference engine to the world — it will offer a counterintuitive template for the AI infrastructure space: in an era of generalization, deliberate narrowness may be an underrated competitive edge.
---
Project: https://github.com/antirez/ds4 Author: Salvatore Sanfilippo (antirez, creator of Redis) License: MIT Backends: Metal / CUDA / ROCm