DS4 · LOCAL FRONTIER INFERENCE
Run frontier open weights locally with ds4.
DwarfStar 4 is a narrow C inference engine for high-memory Mac, CUDA and ROCm machines. It supports DeepSeek V4 and V4.1 Flash, GLM 5.x and Qwen3.8 Flash Next, with text and vision models, local APIs, a CLI and a native agent in one stack.
SUPPORTED: DEEPSEEK V4 / V4.1 + GLM 5.x + QWEN3.8 · MIT LICENSE · C / METAL / CUDA / ROCM · QWEN ON 64GB
PRINCIPLE · LOCAL MODEL STACK
PHASE 1 · THE GIANT
A 284-billion-parameter star
DeepSeek V4 Flash is a large mixture-of-experts model. The usual path is remote serving; ds4 starts from the opposite constraint.
PHASE 2 · THE COLLAPSE
Compressed, not lobotomized
Asymmetric quantization targets the routed experts while preserving critical paths. The model becomes practical on high-memory machines.
PHASE 3 · THE DWARF STAR
Dense, resident, yours
The local engine exposes a CLI, HTTP APIs and a native agent, all sharing the same model state and cache.
How the collapse works →SCROLL ▾
CORE 01
Asymmetric 2-bit quantization
Compress the routed experts, keep critical shared paths precise. That is how the supported routed-MoE builds fit their target machines.
CORE 02
KV cache as a disk citizen
Save long prefixes to SSD and resume by prompt hash. Restarts do not have to mean full re-prefill.
CORE 03
One engine, three interfaces
Use ./ds4 for chat, ./ds4-server for local
APIs and ./ds4-agent for persistent coding sessions.
- SSD STREAMING
- TENSOR PARALLELISM
- SESSION BATCHING
- DSPARK + MTP
- VISION INPUT
- SSD STREAMING
- TENSOR PARALLELISM
- SESSION BATCHING
- DSPARK + MTP
- VISION INPUT
RUNTIME MAP · SIMPLIFIED. SEE ARCHITECTURE NOTES FOR THE FULL DRAWING.
STEP 1 · FETCH THE WEIGHTS
$ cd ds4 && ./download_model.sh ds4f-q2
STEP 2 · BUILD FOR YOUR BACKEND
$ make cuda-spark # Linux · DGX Spark
STEP 3 · TALK TO IT
$ ./ds4-server --ctx 100000 # or serve an API
✓ Runs well
V4 Flash Q2 is the baseline. At 128 GB, GLM 5.3 Q2 and Qwen Q4 also fit; V4.1 Q2 streams from SSD.
./download_model.sh ds4f-q2 && make
REF · M5 MAX 128GB · 32K CTX: 34.4 T/S GEN · 557 T/S PREFILL
Estimates from the ds4 benchmark table. Full guide in Hardware and Installation.
| Machine | Context | Prefill t/s | Generation t/s |
|---|---|---|---|
| M5 Max, 128 GB | q2 · 2,048 tok | 790.2 | 39.4 |
| M5 Max, 128 GB | q2 · 65,536 tok | 398.5 | 27.6 |
| DGX Spark, 128 GB | q2 · 2,048 tok | 825.8 | 18.1 |
| DGX Spark, 128 GB | q2 · 65,536 tok | 823.0 | 13.8 |
Own your local AI inference.
Start with the quickstart, check the hardware matrix, then connect your editor, agent or API client to the local server.