The creator of Redis thinks your old GPU is not obsolete. Salvatore Sanfilippo, better known as antirez, published a project called DwarfStar (the repo is antirez/ds4) that runs DeepSeek V4 Flash, GLM 5.x, and DeepSeek V4 PRO entirely on consumer hardware: Macs, DGX Spark boxes, Strix Halo desktops, and older NVIDIA cards like Ada Lovelace and L40S. It hit the front page of Hacker News this week with hundreds of points, and the repo has already passed 21,000 stars.
The interesting part is not just the speed numbers. It is the design decision behind them. Instead of being a general GGUF runner like llama.cpp, DwarfStar is deliberately narrow: it supports a short list of models, ships its own quantized weights, and tests the whole stack together, from tensor loading to tool calls to the HTTP server. One model, made to feel finished.
Full disclosure before we go further: I have not run DwarfStar on my own machine. Everything below comes from the project's README and its official documentation on GitHub, all linked inline. This is a walkthrough of what it does, what hardware it actually needs, and how you would set it up, not a personal benchmark report. I will flag where the docs are honest about the limits.
What DwarfStar actually is
A few facts from the README that frame everything else:
- It is a self-contained native inference engine in C (
ds4.c), MIT licensed, with no dependency on GGML even though the docs openly credit llama.cpp's kernels and quantization work as the foundation. - Supported models: DeepSeek V4 Flash (including a vision checkpoint), DeepSeek V4.1 Flash, DeepSeek V4 PRO, GLM 5.2, GLM 5.3, GLM 5.3 Flash, and Qwen3.8 Flash Next.
- Backends: Metal on Apple Silicon (the primary target), CUDA including multi-GPU setups, and ROCm on Strix Halo systems like the Framework Desktop.
- It is explicitly beta quality. The docs say so in the Status section, and model support is "intentionally opportunistic": a model can be removed when a better replacement arrives.
That last point matters if you are comparing it to llama.cpp or Ollama. Those tools try to run everything. DwarfStar tries to run a few excellent models perfectly on machines people actually own. You cannot throw an arbitrary GGUF at it; the docs state plainly that other files will fail because the tensor layout, quantization mix, and metadata are specific to the weights the project publishes.
Why this design works: the MoE trick
DeepSeek V4 Flash and the GLM Flash models are mixture-of-experts models. Most of their parameters are routed experts, and only a fraction activate per token. Two consequences make local inference practical:
- Aggressive quantization survives. The docs explain that the Q2 recipe spends most of its compression on routed experts (IQ2_XXS gate/up tensors, Q2_K down) while keeping attention projections, shared experts, and the output layer at much higher precision. The result is that Flash Q2 lands at roughly 81 GiB, small enough for a 96 or 128 GB machine.
- Big contexts stay cheap. DeepSeek V4's KV cache design compresses context aggressively, which is why the server can be started with 100K-token contexts on a single machine.
This is why the same quantized model that would be hopeless in a dense architecture feels "quasi-frontier" here, in antirez's own words from the README.
The hardware tiers, from the official docs
This is the decision most people get wrong, so here is the matrix, straight from the documentation.
-
96 to 128 GB Apple Silicon Mac. The default path. Download
ds4f-q2, run./ds4. This is the setup the project treats as the baseline. - Smaller Mac, less RAM than the model. SSD streaming. The engine keeps a bounded cache of routed experts in RAM and reads missing ones from the GGUF on your SSD. It trades speed for capacity, and the docs are blunt that a model can start successfully and still be too slow for interactive work, so test with a short generation first.
-
DGX Spark.
make cuda-spark. This is the hardware antirez calls the main CUDA goal. -
Older NVIDIA cards, including Ada Lovelace and L40S.
make cuda-generic. This is the part that other backends often do not support. The docs report an eight-L40S setup running Flash Q2 at about 126 tokens per second aggregate across 16 concurrent sessions, essentially a small multi-user LLM server built from cards that are a GPU generation or two old. - Two 128 GB Macs. Tensor parallelism over RDMA, enough to run 4-bit DeepSeek Flash or GLM 5.3 Flash.
- 512 GB workstations. Full-resident DeepSeek V4 PRO Q4, tested on an M3 Ultra.
The SSD streaming numbers are worth seeing because they change what "my laptop cannot fit that model" means. On a 128 GB M5 Max with automatic cache sizing, the docs' September measurements show GLM 5.3 Flash Q4 (a 177.77 GiB file) doing 121 tokens per second initial prefill and around 12 to 15 tokens per second generation, with a three-run median. DeepSeek Flash Vision MXFP4 (145.26 GiB) hit 300 tokens per second prefill and about 12 to 19 tokens per second generation. Those are reading speeds from a fast internal SSD, not a marketing claim, and the docs repeat that your workload will vary.
How to set it up
The happy path from the README takes four commands:
git clone https://github.com/antirez/ds4.git
cd ds4
make # Metal on Apple Silicon
./download_model.sh ds4f-q2
Then either the interactive CLI, the built-in agent, or the server:
./ds4 -p "Explain Redis streams in one paragraph."
./ds4-agent
./ds4-server --ctx 32768
A few details worth knowing before you start:
-
Downloads are resumable. Rerun
./download_model.sh ds4f-q2after an interrupted download and it continues withcurl -C -. That matters when the file is 81 GiB. - Budget memory beyond the model. The docs warn to leave room for context, runtime buffers, and the OS. The model size is the floor, not the requirement.
- Use a fast local SSD. Every streaming mode depends on it, and the V4.1 docs say the same thing in bold-equivalent terms: keep the GGUF on a fast local SSD because Engram tables are read directly from disk.
The part I find most interesting: it is a coding agent server
ds4-server exposes an OpenAI-compatible endpoint on port 8000, and the included ds4-agent is a native coding agent that keeps token history and live model state together in local KV snapshots, so resumed sessions do not rebuild the whole prompt.
The docs include ready-made configs for wiring local DeepSeek into real agent tools:
-
Claude Code. A shell wrapper sets
ANTHROPIC_BASE_URLto your local server and points every model role, including subagents, atdeepseek-v4-flash. -
Codex CLI. A TOML provider block using the Responses API against
127.0.0.1:8000. - OpenCode and Pi. JSON provider entries with the compatibility flags already worked out, including DeepSeek's thinking format.
That last mile is what separates this from a benchmark toy. With cost fields literally set to zero in the Pi config, the pitch is simple: unlimited local tokens for your coding agent, with the privacy of a machine that never phones home. The tradeoff is capability. DeepSeek V4 Flash is strong, but it is not the same as frontier hosted models on the hardest tasks, and the docs' own evaluation section is careful to call ds4-eval "integration checks, not official leaderboard scores."
Two honest caveats from the same docs:
- Beta quality, fast-changing. Instabilities and regressions are explicitly possible between releases.
-
One flag you should know:
--power. For DeepSeek, it trades throughput for lower sustained GPU load, defaulting to 100. If you are running a laptop hot for hours, that flag exists for you. GLM currently requires it left at 100.
One more thing worth noticing
The README has a section called "AI full disclosure" where antirez states the software is developed with strong assistance from AI coding agents, with humans leading ideas, testing, and debugging, and says openly that if you are unhappy with AI-developed code, the project is not for you. He also argues that software should now ship as a working template for the biggest use cases, with users asking coding agents to adapt it to their specific hardware.
Whether or not you agree with that philosophy, it is a notable stance from someone who has written foundational infrastructure by hand for two decades, and the acknowledgements section draws the line clearly: this project leans on AI, while llama.cpp and GGML, which made it possible, were largely written by hand.
Would I run this?
Here is my honest read of who should try it, as a checklist you can save:
- You have 96 GB+ of unified memory and want a private, free coding model: yes, this is probably the most polished path to DeepSeek V4 Flash today.
- You have an older NVIDIA card collecting dust: yes, the Ada Lovelace and L40S multi-GPU story is unusual and worth testing.
- You have a 32 GB laptop: no. The Flash Q2 file alone is 81 GiB, and streaming needs both a big SSD and realistic speed expectations.
- You want a general local LLM runner for many models: no, use llama.cpp or Ollama. DwarfStar will fight you by design.
-
You want to prototype against it first: the OpenAI-compatible server means you can point any existing tool at
127.0.0.1:8000and benchmark the quality yourself before committing disk space.
The bigger story is what it says about the next year of local AI. Specialized engines per model family, quantization recipes built around MoE sparsity, and SSDs acting as an extension of RAM are all already here in one project. The gap between "consumer hardware" and "runs a frontier-class model" keeps shrinking, and this time the person shrinking it built Redis.
I write about AI tools, developer infrastructure, and practical engineering every week. Subscribe, it is free, and it helps me keep doing the hands-on breakdowns.
Have you tried DwarfStar or any local inference setup for coding agents? What hardware are you running, and what tokens per second are you seeing? I am collecting real-world numbers for a follow-up.
Top comments (0)