Profile
Back to NewsBack
Hacker News 3 min
Reader Mode
From the creator of Redis; run LLM locally with ds4

From the creator of Redis; run LLM locally with ds4

11 hours ago

DS4 · LOCAL FRONTIER INFERENCE

Run frontier open weights locally with ds4.

DwarfStar 4 is a narrow C inference engine for high-memory Mac, CUDA and ROCm machines. It supports DeepSeek V4 and V4.1 Flash, GLM 5.x and Qwen3.8 Flash Next, with text and vision models, local APIs, a CLI and a native agent in one stack.

SUPPORTED: DEEPSEEK V4 / V4.1 + GLM 5.x + QWEN3.8 · MIT LICENSE · C / METAL / CUDA / ROCM · QWEN ON 64GB

ds4 · local session

PRINCIPLE · LOCAL MODEL STACK

PHASE 1 · THE GIANT

A 284-billion-parameter star

DeepSeek V4 Flash is a large mixture-of-experts model. The usual path is remote serving; ds4 starts from the opposite constraint.

PHASE 2 · THE COLLAPSE

Compressed, not lobotomized

Asymmetric quantization targets the routed experts while preserving critical paths. The model becomes practical on high-memory machines.

PHASE 3 · THE DWARF STAR

Dense, resident, yours

The local engine exposes a CLI, HTTP APIs and a native agent, all sharing the same model state and cache.

How the collapse works →

SCROLL ▾

CORE 01

Asymmetric 2-bit quantization

Compress the routed experts, keep critical shared paths precise. That is how the supported routed-MoE builds fit their target machines.

CORE 02

KV cache as a disk citizen

Save long prefixes to SSD and resume by prompt hash. Restarts do not have to mean full re-prefill.

CORE 03

One engine, three interfaces

Use ./ds4 for chat, ./ds4-server for local APIs and ./ds4-agent for persistent coding sessions.

  • SSD STREAMING
  • TENSOR PARALLELISM
  • SESSION BATCHING
  • DSPARK + MTP
  • VISION INPUT
  • SSD STREAMING
  • TENSOR PARALLELISM
  • SESSION BATCHING
  • DSPARK + MTP
  • VISION INPUT

RUNTIME MAP · SIMPLIFIED. SEE ARCHITECTURE NOTES FOR THE FULL DRAWING.

STEP 1 · FETCH THE WEIGHTS

ds4 · zsh
$ git clone https://github.com/antirez/ds4
$ cd ds4 && ./download_model.sh ds4f-q2

STEP 2 · BUILD FOR YOUR BACKEND

ds4 · zsh
$ make # macOS · Metal
$ make cuda-spark # Linux · DGX Spark

STEP 3 · TALK TO IT

ds4 · zsh
$ ./ds4
$ ./ds4-server --ctx 100000 # or serve an API

✓ Runs well

V4 Flash Q2 is the baseline. At 128 GB, GLM 5.3 Q2 and Qwen Q4 also fit; V4.1 Q2 streams from SSD.

./download_model.sh ds4f-q2 && make

REF · M5 MAX 128GB · 32K CTX: 34.4 T/S GEN · 557 T/S PREFILL

Estimates from the ds4 benchmark table. Full guide in Hardware and Installation.

Machine Context Prefill t/s Generation t/s
M5 Max, 128 GB q2 · 2,048 tok 790.2 39.4
M5 Max, 128 GB q2 · 65,536 tok 398.5 27.6
DGX Spark, 128 GB q2 · 2,048 tok 825.8 18.1
DGX Spark, 128 GB q2 · 65,536 tok 823.0 13.8
All benchmarks →
OpenCode Claude Code Codex CLI Pi /v1/chat/completions · /v1/messages · /v1/responses

Own your local AI inference.

Start with the quickstart, check the hardware matrix, then connect your editor, agent or API client to the local server.

Chat with me