Redis Creator's DwarfStar: Specialized Local Inference Engine for DeepSeek V4 & GLM 5

DwarfStar is a specialized local inference engine by Redis creator antirez, optimized exclusively for DeepSeek V4, GLM 5, and Qwen3.8 Flash Next models across Metal, CUDA, and ROCm platforms, achieving 790 tokens/s prefill on M5 Max and 126 aggregated tokens/s on 8×L40S, with built-in coding agent and HTTP server.

Architecture Digest
Architecture Digest
Architecture Digest
Redis Creator's DwarfStar: Specialized Local Inference Engine for DeepSeek V4 & GLM 5

Introduction

Running open-source large models locally often means wrestling with generic runners like Ollama that support everything but optimize nothing, or vLLM that demands complex environment setup. Redis creator antirez took a different approach: build a dedicated inference engine for only the strongest open-weight models, pushing performance to the limit.

The project, DwarfStar (repository: github.com/antirez/ds4), reached 23,000 stars in under five months under the MIT license. It targets DeepSeek V4 Flash (including experimental vision), DeepSeek V4.1 Flash, DeepSeek V4 PRO, GLM 5.2/5.3/5.3 Flash, and Qwen3.8 Flash Next — and only these models. Users must use the project's own GGUF files, trading breadth for extreme memory efficiency and throughput.

DwarfStar overview diagram
DwarfStar overview diagram

Core Highlights

Highlight 1: Purpose-Built for a Few Models, Not a Universal Runner

Antirez explicitly states the project deliberately narrows scope. The supported model list is short but each is a current top-tier open weight. The project follows an opportunistic model-support strategy: it tracks the best open weights and replaces older models when better ones appear. This philosophy mirrors Redis's own focus on speed, simplicity, and correctness.

Highlight 2: Full Hardware Spectrum — Metal, CUDA, ROCm, and Multi-Node Memory Pooling

DwarfStar covers the entire hardware spectrum by addressing both small-memory and multi-node paths:

Metal (primary target): Macs with 96 GB+ run models fully in memory; smaller machines use SSD streaming to pull weights on demand.

CUDA : DGX Spark is the main target, with support for Ada and L40S multi-GPU setups.

ROCm : Strix Halo / Framework Desktop APU platforms.

Multi-node scaling is a standout feature:

Two 128 GB Macs connected via RDMA perform tensor parallelism to run the 4-bit quantized DeepSeek V4 Flash.

Pipeline parallelism aggregates memory across multiple machines for even larger models.

Official benchmarks:

M5 Max 128 GB running Flash Q2: prefill 790 tokens/s , generation 39 tokens/s (2048 context).

DGX Spark 128 GB : prefill 825 tokens/s , generation 18 tokens/s .

8×L40S running Flash Q4: ~126 aggregated generation tokens/s (16 concurrent decode sessions, 100k context). Note: this is aggregate throughput across 16 sessions, not per-user speed.

Benchmark chart for M5 Max and DGX Spark
Benchmark chart for M5 Max and DGX Spark

PRO model on M3 Ultra shows prefill peak ~188 tokens/s, generation ~20 tokens/s — the same specialized inference path serves both Flash and PRO tiers.

PRO model performance curve on M3 Ultra
PRO model performance curve on M3 Ultra

Highlight 3: Native Coding Agent and HTTP Server

DwarfStar integrates a coding agent directly into the inference engine:

ds4-agent : Native coding agent without HTTP server, preserves token history and model state, uses the model's native tool format. The /hints on command reveals reasoning behind programming choices.

ds4-server : Standard HTTP server compatible with Pi, OpenCode, Codex CLI, and Claude Code.

Speculative decoding (MTP) : GLM and Qwen use --mtp; DeepSeek has dedicated support models, visibly boosting generation speed.

Directional steering + think level : --think-level 1–100 adjusts reasoning intensity, plus directional steering for advanced control.

Official data for Qwen3.8 Flash Next on M3 Ultra: MTP decode reaches 90.6 tokens/s vs. 54.7 tokens/s for regular decode — a clear gain for coding workloads.

MTP vs regular decode comparison for Qwen3.8 Flash Next
MTP vs regular decode comparison for Qwen3.8 Flash Next

Core Principles & Architecture Differences

Three key decisions explain why DwarfStar outperforms generic runners:

Decision 1: No generic GGUF runner — write a dedicated inference path per model. Generic runners accumulate branches for hundreds of models. DwarfStar's ds4.c does not link against GGML; it optimizes solely for the DeepSeek V4 inference path. Antirez acknowledges the foundation laid by llama.cpp and GGML (Georgi Gerganov), with kernels, quantization formats, and GGUF ecosystem knowledge derived from that work; parts of the source are retained or adapted under MIT, and the LICENSE preserves GGML's copyright notice.

Decision 2: Quantization tailored to model architecture, not one-size-fits-all. The README notes DeepSeek V4 Flash/PRO and GLM 5.2 tolerate aggressive routed-expert quantization. This allows DwarfStar to apply more aggressive quantization for these specific architectures; Qwen3.8 Q2 main weights are only 41.73 GiB, enabling a 64 GB Mac to run it.

Decision 3: Treat KV cache and Engram table design as first-class citizens. Compressed KV cache plus fast local SSD makes long contexts practical. The Engram table stays on disk, so a fast SSD is repeatedly emphasized.

Installation & Configuration

Example for 96 GB or 128 GB Mac:

git clone https://github.com/antirez/ds4.git
cd ds4
make                          # Metal platform build
./download_model.sh ds4f-q2   # Download DeepSeek V4 Flash Q2
./ds4 -p "Explain Redis streams in one paragraph."
./ds4-agent                   # Native coding agent
./ds4-server --ctx 32768      # HTTP server, default 127.0.0.1:8000

Other platforms use different build targets:

Metal (Apple Silicon): make DGX Spark: make cuda-spark Strix Halo / Framework Desktop: make strix-halo Single/Multi-GPU CUDA (Ada/L40S): make cuda-generic Small-memory Macs use SSD streaming: after ./download_model.sh ds4f-q2, add --ctx to control context; model weights stay on SSD and are read on demand, allowing models larger than RAM.

Team Deployment Solutions

Team access pattern. If your team has DGX Spark or L40S workstations, DwarfStar can serve as an internal inference service. For multi-user scenarios, use ds4-server + --batched-session with tensor parallelism across multiple GPUs, letting one machine serve the whole team.

Batch deployment. Build uniformly with make cuda-generic or make cuda-spark. Model files are fetched via ./download_model.sh into the gguf/ directory with resume support. The official 8×L40S deployment command is provided verbatim:

--gpu-devices 0,2,4,6,1,3,5,7 --ctx 100000 --batched-session 16

.

Coding client integration. ds4-server exposes an OpenAI-compatible API, directly pluggable into Pi, OpenCode, Codex CLI, and Claude Code (full config in docs/CLIENTS.md). CI regression testing for inference correctness can use ds4-eval.

Team policy customization. Standardize reasoning intensity with --think-level (1–100), balance throughput and GPU load with --power. Security note: no built-in permission system ; runs with launching user's privileges. Isolation requires Gondolin, Docker, or OpenShell sandbox.

Use Cases

Enthusiasts running DeepSeek V4 / GLM 5 / Qwen3.8 locally : Mac large-memory users, DGX Spark owners, Strix Halo hobbyists.

Teams needing multi-GPU/multi-node memory pooling for larger models : Dual Mac RDMA, 8×L40S multi-user serving.

Developers wanting a local coding agent : Keep code off cloud APIs, run ds4-agent entirely locally. Technical leads making the specialist-vs-generalist trade-off : Willing to lock into a few models for peak performance.

Pros, Cons & Pitfalls

Core strengths : Antirez's engineering taste backs the specialist approach, delivering measurable throughput gains; MIT licensed; complete hardware coverage with honest technical disclosure.

Limitations :

Beta quality, rapid iteration . Author admits possible instability and regressions; each release gets a QA pass but stability isn't guaranteed.

Only supports the project's own GGUFs . Other models won't work — the cost of specialization.

High hardware barrier . Primary targets are 96 GB+ Mac, DGX Spark, Strix Halo; ordinary PCs fall back to slow SSD streaming.

Pitfall warnings :

The 126 tokens/s is aggregate throughput across 16 sessions , not single-user speed — don't mistake it for personal performance.

No permission system; plan isolation before running the coding agent on production machines.

Model downloads are ~137 GiB; ensure SSD capacity and speed are sufficient.

Official disclosure: heavily AI-assisted development; factor that in if it matters to you.

Final Thoughts

Most people take the wide road of universal runners; antirez chose the narrow path of a specialist engine. The narrow road is harder, but once traversed it delivers an experience others can't match — M5 Max prefill at 790 tokens/s, 8×L40S aggregate at 126 tokens/s — numbers earned by obsessing over a handful of models. If you need to run DeepSeek V4, GLM 5, or Qwen3.8 Flash Next locally, DwarfStar is the purest enthusiast choice.

Closing illustration
Closing illustration

GitHub: github.com/antirez/ds4

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

CUDAMetallocal inferenceantirezROCmcoding agentGLM 5DeepSeek V4Qwen3.8DwarfStar
Architecture Digest
Written by

Architecture Digest

Focusing on Java backend development, covering application architecture from top-tier internet companies (high availability, high performance, high stability), big data, machine learning, Java architecture, and other popular fields.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.