On October 2, 2026, Hacker News put DwarfStar 4 back on the front page. The story, item 49936575, is titled "From the creator of Redis; run LLM locally with ds4" and points at dwarfstar.sh. Algolia showed 150 points and 39 comments on October 3. Salvatore Sanfilippo (antirez) already shipped a different C engine this summer, h3.c for MiniMax H3 video. ds4 is the text-and-vision sibling: a narrow inference stack for a few large open-weight families, not a catalog of every GGUF on disk.
explainx.ai mentioned the project only in passing. The June Qwen 3.6 local guide cited an aggressive V4 Flash quant at about 33 tok/s and 103 GB of RAM. That line is stale against the files in the repo today. Checked on October 3, 2026, against the README, docs/MODELS.md, docs/PERFORMANCE.md, the community site, and the thread above.
The repo is MIT-licensed C. GitHub listed 22,951 stars, 2,215 forks, and 753 open issues. It was created May 6, 2026, and the default branch last pushed on September 20, 2026. The community site footer says Built 2026-09-17. The README calls the software beta: fast-changing, with regressions still possible. It also says the code was written with strong help from coding agents, and that people who do not want AI-written systems code should skip it.
Can I run ds4 this week?
| Question | Answer, checked October 3, 2026 |
|---|---|
| Whose project is this? | Salvatore Sanfilippo (antirez). Engine repo antirez/ds4. Community site dwarfstar.sh. |
| What does it run? | DeepSeek V4 Flash (including an experimental vision model), V4.1 Flash, V4 PRO, GLM 5.2, GLM 5.3, GLM 5.3 Flash, and Qwen3.8 Flash Next. Text and vision, with the vision encoder passed separately. |
| What is the first download? | ./download_model.sh ds4f-q2 on a 96 GB or 128 GB machine. That file is about 81 GiB. |
| 64 GB Mac? | Qwen3.8 Flash Next Q2. Resident weights are 41.73 GiB. The download is 137.10 GiB because BF16 n-grams stay on disk. |
| 128 GB Mac, "runs well"? | The site's fit line is right for V4 Flash Q2, GLM 5.3 Flash Q2 (about 90 GiB), and Qwen Q4 (about 69.74 GiB resident). Full GLM 5.3 Q2 is about 197 GiB and is a streaming or bigger-machine download. |
| V4.1 on 128 GB? | Stream from SSD. Q2 main weights are 152 GiB inside a 341 GiB file. Q4 main weights are 294 GiB; full residency wants a 512 GB Mac. |
| Speed on M5 Max, 128 GB, Flash Q2? | 34.36 tok/s generation and 557.04 tok/s prefill at 32,768 context. At 2,048 context: 39.35 gen, 790.18 prefill. |
| DGX Spark, same quant? | Faster prefill, slower generation: 18.05 tok/s gen at 2,048 context, 13.84 at 65,536. |
| Claude Code or Codex? | Yes, via ./ds4-server on 127.0.0.1:8000. The binary is ./ds4, not a chat subcommand. |
| llama.cpp instead? | Use llama.cpp when you need many families. Use ds4 when the model is on its download list and you want that layout tested with the server and agent. |
| License and maturity? | MIT. Beta. 753 open issues on October 3. |
The community homepage compresses hardware into "Apple Silicon 64 GB+ depending on model." The README is stricter, and that stricter line is the one to plan a purchase around. Metal is the primary target on Macs with 96 GB or more. Smaller machines can stream experts from SSD. Qwen on 64 GB is a real documented path, called out in the site banner and in the Qwen section of the model guide. It is not a promise that V4 Flash Q2 sits resident in 64 GB. An 81 GiB checkpoint plus context and runtime buffers does not.
Which quant should I download?
Start from the download script names, not from a generic quant label. ds4f-q2 is DeepSeek V4 Flash 0731, about 81 GiB, and the model guide calls it the starting point for 96 GB and 128 GB systems. The compression is asymmetric on purpose. Routed-expert gate and up projections use IQ2_XXS. Down projections use Q2_K. Other components stay at higher precision. There is also ds4f-q2-q4, which lifts the last six routed-expert layers to Q4 and needs more memory, and ds4f-q4 for larger or distributed machines.
DeepSeek V4.1 is a different checkpoint. V4 Flash weights and V4.1 weights are not interchangeable, and V4.1 does not reuse the V4 DSpark support file. ds41f-q2 is a 341 GiB file with 152 GiB of main weights, plus 189 GiB of Engram tables that stay on disk in every mode. On one 128 GB Mac the model guide says to pass --ssd-streaming. A 256 GB or larger Mac can hold the main weights. Full residency has been tested on an M3 Ultra with 512 GB. ds41f-q4 is 483 GiB on disk and 294 GiB of main weights. That is the "Mac Studio 512 GB" line on the community site: Q4 wants that class of machine for residency, or SSD streaming on something smaller. PRO Q2 (pro-q2-imatrix) is in the same bucket: a 512 GB resident target, or streaming.
GLM naming on the homepage is easy to over-read. "GLM 5.3 Q2 also fits" at 128 GB matches GLM 5.3 Flash (glm53-q2, about 90 GiB), which the model guide lists for one 128 GB Mac, a DGX Spark, or ROCm. The same page says that file sits close enough to the memory budget that other apps and context length still matter. Full GLM 5.3 Q2 is about 197 GiB (glm53-full-q2) and the guide tells you to stream it or use a larger machine. GLM 5.2 has its own larger downloads, including an Unsloth Q4 shard set the project fetches as glm-unsloth-q4. glm53-fp8 is packaged and not implemented for inference. Do not download it expecting a server.
Qwen3.8 Flash Next is the 64 GB on-ramp, and it is also a sensible 128 GB model if you would rather spend RAM on context than on DeepSeek. qwen38-q2 puts 41.73 GiB of main and MTP weights in memory and leaves 95.37 GiB of original BF16 n-grams on the SSD. The README's first command for that file is an 8,192-token context and a 1,024-token prefill chunk. qwen38-q4k is 69.74 GiB resident and 165.11 GiB on disk, which is why the site can say Qwen Q4 fits a 128 GB Mac while the download is still larger than the RAM. Vision is a second file (qwen38-vision) passed with --vision. Metal and CUDA are the documented Qwen backends.
If you already compared dense Qwen 3.6 27B with bigger MoE checkpoints in the local development guide, treat ds4 as the path for the families on its own download list, especially DeepSeek V4 Flash. The Qwen3.8-Flash-Next launch note is the model card; ds4 is one runtime that now has a first-party download for it.
What tok/s should I expect?
Read prefill and generation as different jobs. Prefill is what you pay when the prompt, the repo, or the tool definitions land. Generation is what you feel while the agent writes. The upstream table is a recorded Flash Q2 sweep: an M5 Max with 128 GB, 2,048-token continued-prefill steps, and 128 greedy generation tokens at each frontier. docs/PERFORMANCE.md says it is a baseline, not a fresh timing of every commit. The community site reprints four of those rows and rounds them. The unrounded figures are:
| Machine | Context | Prefill | Generation |
|---|---|---|---|
| M5 Max, 128 GB | 2,048 | 790.18 t/s | 39.35 t/s |
| M5 Max, 128 GB | 16,384 | 572.53 t/s | 36.14 t/s |
| M5 Max, 128 GB | 32,768 | 557.04 t/s | 34.36 t/s |
| M5 Max, 128 GB | 65,536 | 398.50 t/s | 27.64 t/s |
| DGX Spark, 128 GB | 2,048 | 825.76 t/s | 18.05 t/s |
| DGX Spark, 128 GB | 16,384 | 872.44 t/s | 15.10 t/s |
| DGX Spark, 128 GB | 32,768 | 855.94 t/s | 14.43 t/s |
| DGX Spark, 128 GB | 65,536 | 822.98 t/s | 13.84 t/s |
The site's "32K ctx: 34.4 tok/s gen, 557 tok/s prefill" line is that 32,768 row, rounded. Its 2,048 and 65,536 cells (790.2 / 39.4, 398.5 / 27.6, 825.8 / 18.1, 823.0 / 13.8) are the same sweep. Use the performance doc when a tenth of a token matters, and expect the homepage to keep the shorter decimals.
A commenter hoping for about 50 tok/s will not see that on this Flash Q2 baseline. The M5 Max is in the high 30s at short context and the high 20s at 65,536. The Spark generates slower than the Max and prefills as fast or faster. That split matters for agents: a long first prompt is a prefill problem, and a chatty tool loop is a generation problem. The README also reports a different shape of speed, aggregate throughput: an eight-L40S Flash setup reaching about 126 tok/s across 16 sessions. That is a server number, not a single-user laptop number.
Speculative decoding is off unless you ask. Qwen and GLM take --mtp. V4 Flash DSpark needs a matching support GGUF. The performance notes say a faster drafter can still lose to ordinary decode on an unpredictable prompt, so turn it on only after you time your own workload. Thinking is on by default. --nothink or /nothink asks for a direct answer. Sampling defaults are temperature 1, top-p 1, and min-p 0.05.
Will Claude Code, Codex, or OpenCode attach?
The server is the compatibility layer. ./ds4-agent runs inference in-process, keeps token history next to live model state, and uses each family's native tool format. Sessions land in ~/.ds4/kvcache. A compatible snapshot skips a full prefill. Sessions that contain images cannot be saved yet. For Pi, OpenCode, Codex CLI, or Claude Code, the README sends you to ds4-server and docs/CLIENTS.md instead of the native agent.
The client guide's server line is the long-context one, not the short README demo:
./ds4-server --ctx 100000 --kv-disk-dir /tmp/ds4-kv --kv-disk-space-mb 8192
The README's everyday example is ./ds4-server --ctx 32768. Both are upstream. The homepage quickstart also shows --ctx 100000. Pick one budget and keep the client at or under it. Output tokens spend the same window. The placeholder API key in the docs, dsv4-local, is not authentication. The process listens on 127.0.0.1:8000 unless you change it. Point a remote harness at that port only on a network you trust.
Codex uses the Responses API. The provider block sets base_url to http://127.0.0.1:8000/v1 and wire_api to responses, then:
codex --model deepseek-v4-flash -c model_provider=ds4
Claude Code uses the Anthropic-style route on the same host, without the /v1 suffix. A wrapper unsets ANTHROPIC_API_KEY, sets ANTHROPIC_BASE_URL to http://127.0.0.1:8000, sets ANTHROPIC_AUTH_TOKEN to the placeholder, and points the main model plus the Sonnet, Haiku, Opus, and subagent model variables at deepseek-v4-flash. Name the wrapper something other than claude. The first prefill on an agent prompt can sit there for a while. Disk KV is how a later session reuses a matching prefix.
OpenCode is the same idea as the local OpenCode stack: an OpenAI-compatible base URL, here http://127.0.0.1:8000/v1, model id deepseek-v4-flash, context limit copied from the server. Pi gets a longer provider entry in the client guide, including a note that the sample config is text-only until you declare image input and pass a vision encoder. If you already run OpenCode against llama.cpp or Ollama, ds4 is another provider block, not a new harness.
What will ds4 refuse to run?
The README's own line is the scope: "not a general GGUF runner: you need to use the GGUF files the project produces." A random Llama, Gemma, or Mistral file from another quantizer is outside that scope. Model support is opportunistic. The project follows whatever open weights fit 128 GB laptops and 256 GB or 512 GB workstations, and it says a family can be removed when a better one shows up.
Documented builds, from the start-here table:
| Machine | Command |
|---|---|
| Apple Silicon, Metal | make |
| NVIDIA DGX Spark | make cuda-spark |
| AMD Strix Halo, including Framework Desktop | make strix-halo |
| One or more CUDA cards, including Ada and L40S | make cuda-generic |
The homepage quickstart only prints make and make cuda-spark. The Strix and generic CUDA targets are still in the README. ROCm for Strix Halo is a mainline guide, not a rumor. Two 128 GB Macs over RDMA can hold 4-bit DeepSeek Flash or GLM 5.3 Flash with tensor parallelism. Pipeline parallelism is how the project adds RAM across machines. A Windows build is not in that table.
llama.cpp remains the general tool. ds4 does not link GGML. It does keep GGML's copyright in the license, and it says some quant layouts, CPU dot-product code, and kernels were adapted from that project. The practical difference Simon Willison described on the October thread: a runner that accepts every architecture also accepts a lot of ways to break tool calling. ds4's bet is a short list, checked end to end, including tool calls, KV state, the HTTP server, and the agent. That is also the limitation. If your model is not on the list, the engine will not grow a backend for it this afternoon.
Hardware outside those classes is a community patch, not a support promise. One commenter reported about 22 tok/s on an Intel Ultra 7 255H after a small iGPU change. That is a thread report, not a row in docs/PERFORMANCE.md. A discrete NVIDIA card that is not in the CUDA guide, or a CPU-only box with a thin GPU, is the case the older HN threads already warned about: this is not the llama.cpp split where shared layers sit on the GPU and experts sit in system RAM. For a Spark-shaped box, start from the DGX Spark local setup and then use make cuda-spark. For the Mac-versus-tower choice in general, the hardware comparison still applies; ds4 just moves the Mac bar toward 96 GB and 128 GB unified memory.
Is 2-bit actually usable?
The project thinks these particular MoE families tolerate aggressive expert quantization. That is a design claim, written in the motivations section, not a leaderboard. ds4-eval runs the repo's own capability checks against a real GGUF. The README is explicit that those checks are integration tests, not official scores.
The October thread split in public. One commenter wrote that the heavily quantized DeepSeek V4 checkpoint is not very good. Another wrote that ds4 quants beat Unsloth and linked a SlopCodeBench write-up of DeepSeek V4 Flash 0731. explainx.ai opened that file. It never says ds4, DwarfStar, or Unsloth. It compares two local runs of the same three problems: a vendor UD-Q2_K_XL quant that strictly solved 1 of 17 checkpoints, and a larger community q2q4 imatrix quant (about 91 GB) that strictly solved 5 of 17. The author says both the output-token cap and the quant changed between the runs, and that the correctness jump cannot be credited to the quant alone. A reply under the "Unsloth quants are amazing" comment asked which ds4 checkpoint was even tested. Nobody posted a clean A/B in that thread.
So the disagreement is real, and the linked file does not crown a winner. If you need a higher-precision DeepSeek path inside ds4, the downloads already include Q4 and mixed Q2/Q4 targets, at a RAM or SSD cost the tables above spell out. Hosted V4 pricing and agent benchmarks live in the DeepSeek V4 Pro write-up. Local Q2 will not reproduce a hosted full-precision score, and this post will not pretend a 17-checkpoint swing settled it.
SSD streaming versus buying more RAM
RAM decides whether the main weights can sit still. The SSD decides whether a too-large expert stack, the Engram tables, or the KV prefix can still move. Those are different bills.
V4 Flash Q2 at about 81 GiB is a resident model on 128 GB, with room left for context, if you close other heavy apps. GLM 5.3 Flash Q2 at about 90 GiB is the tight version of the same idea. V4.1 Q2 cannot be resident on 128 GB, because the main weights alone are 152 GiB, and the file on disk is 341 GiB. Streaming is the supported mode there, and the Engram rows are read from disk even when the rest of the model is resident. Qwen Q2 looks small in RAM and large on disk for the same reason: the n-grams are not loaded as a table.
KV reuse is the other SSD job. The community site's demo says the cache key is the SHA1 of the rendered prompt prefix, stored on disk, so a matching prefix reloads instead of prefilling again. The README's version of that feature: compatible local snapshots avoid rebuilding the prompt; stripped sessions and network tensor-parallel restores still prefill. The client guide's --kv-disk-dir flag is how the server keeps that cache across processes. A commenter who has been running Qwen3.8 Flash Next on an M5 Max with 128 GB described ds4-agent as append-only, so the prefix stays reusable, and said they drive it with the "oh my pi" harness. That matches the documented design. It is also why a harness that rewrites the top of the prompt will throw away the thing you bought the SSD for.
A 1,000,000-token context on a 128 GB M5 Max is not in the recorded baseline. Victor Lowther's pull request 1115, opened September 24, 2026, adds fused TurboQuant KV at integer widths from 2 to 8 and claims that an 8-bit KV cache makes a 1M context practical for Qwen3.8 Flash Next on that machine. GitHub still lists the pull request as open (merged: false on October 3). A reviewer on an M5 Max 128 GB reported kernel tests passing and a restart-cache miss that the author later called a buggy test. Until that branch is on main, the numbers to quote are the 65,536-token baseline, not 1M.
Community forks exist. neomantra maintains an FFI build and a Go wrapper called ds4go, with vision and Qwen following upstream. Useful if you want a library instead of the CLI. It is a fork, not the project this post is timing.
Why C, and what people argued on the thread
The language question showed up immediately. Sanfilippo has written C for decades. On the thread, readers restated his position: he finds Rust less ergonomic for this kind of work, he thinks coding agents write C more reliably because so much high-quality systems code in their training data is C, and he treats security-critical software as the case for Rust. A commenter pointed at his own YouTube video, L'entusiasmo di DHH è il nemico sbagliato, at about 8 minutes 46 seconds, Italian with an English dub, for that argument. The video is his. The thread is the citation for the timestamp. Other commenters disagreed, including the point that C makes safety a whole-program property and that agents are weak at whole-program reasoning. ds4 does not claim to be a memory-safe runtime. The README's bet is a small vertical engine you can point an agent at and modify, with QA runs before releases.
Tool-calling throughput, in tokens per second, is not published in the performance doc. What is published is that tool calls are part of the integration surface, and that each family has its own chat template. Willison's comment is the operational version: fewer models, fewer ways to wire the template wrong. If you need a measured tool-call rate, you will be timing your own harness. ds4-eval can tell you whether a GGUF still passes the repo's cases. It will not print a tok/s for Claude Code.
A short setup, then stop
This is the 96 GB or 128 GB Flash Q2 path from the README, not a rewrite of every guide:
git clone https://github.com/antirez/ds4.git
cd ds4
./download_model.sh ds4f-q2
make
./ds4
make is macOS Metal. Spark is make cuda-spark. Strix Halo is make strix-halo. Generic CUDA is make cuda-generic. Downloads resume if you rerun the script. Pass -m when you do not want the ds4flash.gguf link the downloader updates. For Qwen on 64 GB, swap the script target to qwen38-q2 and start at --ctx 8192. For V4.1 on one 128 GB Mac, use ds41f-q2 and --ssd-streaming.
Leave headroom. The model file size is not the process size. Context, runtime buffers, and anything else on the machine sit beside the weights. Unified-memory prices are why that headroom is a budget question as much as a kernel question; the RAM price note from August is the cost side of buying a 128 GB Mac for this.
What to do with the decision
Run ds4 this week if the model you want is on the download list and the machine matches a row you can actually hold: Flash Q2 on 96 GB or 128 GB, Qwen Q2 on 64 GB, V4.1 Q2 only with a large fast SSD or a much larger Mac, GLM 5.3 Flash Q2 when you can give it most of a 128 GB box. Attach Claude Code, Codex, or OpenCode to ds4-server when you want a harness you already know. Use ds4-agent when you want the native tool format and a KV snapshot in ~/.ds4.
Wait, or use another runtime, if you need an arbitrary GGUF, a measured guarantee that 2-bit matches a hosted model, a merged 1M-context path, or generation near 50 tok/s on the published Flash Q2 sweep. The Spark will not give you the M5 Max's generation rate. The M5 Max will not give you the Spark's long-context prefill. The homepage rounds the table. The performance file has the rest of the rows.
Figures above are from antirez/ds4 (README, docs/MODELS.md, docs/PERFORMANCE.md, docs/CLIENTS.md), dwarfstar.sh, GitHub metadata, and Hacker News item 49936575, as checked on October 3, 2026. The speed table is a recorded baseline and will drift as the beta moves.
Related reading
- antirez's h3.c: MiniMax H3 on Apple Silicon — the same author's earlier C and Metal engine
- Qwen 3.6 27B local development — the June note that mentioned DwarfStar in passing
- Run open-source models locally with OpenCode — provider blocks for a local OpenAI-compatible server
- What is llama.cpp? — the general local runner, for families outside the ds4 download list
- DGX Spark local LLM setup — the CUDA box ds4's Spark target assumes
- DeepSeek V4 Pro benchmarks and pricing — hosted V4, for when local Q2 is the wrong comparison
- MacBook vs dedicated GPU for local LLMs — memory tiers before you buy for a 128 GB quant
- Qwen3.8-Flash-Next — the Qwen family ds4 now downloads
Official docs: dwarfstar.sh, antirez/ds4, performance, models, clients.
