vLLM Docker Optimized for DGX Spark (single or multi-node)
This repository contains the Docker configuration and startup scripts to run vLLM on DGX Spark, from a single node to multi-node clusters using Ray or vLLM's native PyTorch distributed mode. It supports InfiniBand/RDMA (NCCL), custom environment configuration, and high-performance model loading through fastsafetensors and InstantTensor. Cluster setup supports direct connections between dual Sparks, QSFP/RoCE switch configurations, and 3-node mesh configurations.
While it was primarily developed to support multi-node inference, it works just as well on single-node setups.
If you point an AI agent at this repository to prepare a host or run a recipe, use the operational agent runbook. Agent-assisted repository changes use the separate development guide; AGENTS.md routes agents to the appropriate guide.
Table of Contents
- DISCLAIMER
- QUICK START
- REGULAR QUICK START
- CHANGELOG
- 1. Building the Docker Image
- 2. Launching the Cluster (Recommended)
- 3. Running the Container (Manual)
- 4. Configuration Details
- 5. Mods and Patches
- 6. Launch Scripts
- 7. Using cluster mode for inference
- 8. Model Loading
- 9. Benchmarking
- 10. Downloading Models
DISCLAIMER
This repository is not affiliated with NVIDIA or their subsidiaries. This is a community effort aimed to help DGX Spark users to set up and run the most recent versions of vLLM on Spark cluster or single nodes.
By default, build-and-copy.sh pulls the tested nightly runner image from DockerHub: eugr/spark-vllm:latest. Nightly images are built and tested on multiple models in both cluster and solo configuration before latest is advanced.
We will expand the selection of models we test in the pipeline, but since vLLM is a rapidly developing platform, some things may break.
Selecting --exp-b12x without local-build flags or customizations pulls the separately tested
eugr/spark-vllm-b12x:latest image and tags it as vllm-node-b12x.
If you want to build only the runner from precompiled vLLM and FlashInfer wheels, specify --use-wheels. This option never falls back to compiling missing wheels: if a wheel cannot be downloaded or found locally, the command stops with an error. To build the latest vLLM from the main branch, use --rebuild-vllm; to target a specific repository, release, or commit, set --vllm-repo and/or --vllm-ref. To build a private or already-available checkout without cloning it inside Docker, use --vllm-source-dir.
Similarly, --rebuild-flashinfer, --flashinfer-ref, and --apply-flashinfer-pr control the FlashInfer build and force the local build path.
Local builds include a CUDA-on-WSL memory-reporting fix in both compiled vLLM
wheels and runners built with --use-wheels. On integrated NVIDIA GPUs under
WSL, vLLM keeps CUDA's reported free memory instead of replacing it with guest
RAM availability. Native Linux UMA accounting and proactive allocator-cache
release keep their upstream behavior.
Source builds also trim unused glibc CPU heap pages after startup garbage
collection in API servers and workers. This complements the existing CUDA
allocator cleanup before KV cache sizing/allocation. The CPU trim is included
in exported vLLM wheels and runs after warmup; platforms without malloc_trim
skip it.
Runtime images also set VLLM_WSL2_ENABLE_PIN_MEMORY=1 by default. Pass
-e VLLM_WSL2_ENABLE_PIN_MEMORY=0 to launch-cluster.sh or docker run to opt
out.
Regular and B12X source builds include a targeted patch based on
vLLM PR #58028. When --api-key
is configured, the Python frontend requires authentication for every route
except /health, /ping, /load, and /version; CORS preflight requests also
remain exempt. This includes protecting /metrics, /tokenize, and the API
docs. The patch is included in exported vLLM wheels and leaves the Rust
frontend unchanged.
QUICK START (USING RECIPES)
Single Spark
Check out locally. Do it on the head node of the cluster. This will build the image, download the model and launch it in the container.
git clone https://github.com/eugr/spark-vllm-docker.git
cd spark-vllm-docker
./run-recipe.sh recipes/qwen3.8-flash-next-nvfp4-solo.yaml --solo --setup
Dual Sparks (or more)
Before you start, make sure you connect your Sparks together and enable passwordless SSH as described in our Networking Guide. You can also check out NVIDIA's Connect Two Sparks Playbook, but using our guide is the best way to get started. The guide includes instructions for 3-node Spark mesh clusters.
Check out locally. Do it on the head node of the cluster. This will build the image, download and distribute the model and launch the cluster.
git clone https://github.com/eugr/spark-vllm-docker.git
cd spark-vllm-docker
./run-recipe.sh recipes/deepseek-v4-flash-vision-exp.yaml --setup
QUICK START (USING LAUNCHER)
These examples show image preparation, model download, and launch-cluster.sh
commands separately. Qwen3.8-27B uses the regular image (vllm-node). The
Qwen3.8 Flash Next and DeepSeek Vision examples match the shortcut above and
use the B12X image (vllm-node-b12x). Choose the example for your setup.
Check out the repository
Run the commands on your Spark, or on the head node of your cluster:
git clone https://github.com/eugr/spark-vllm-docker.git
cd spark-vllm-docker
Single Spark: Qwen3.8-27B NVFP4 (regular image)
This matches the Qwen3.8-27B NVFP4 with DFlash2 recipe in solo mode. It uses FP8 KV cache, InstantTensor loading, and the DFlash2 draft model for speculative decoding, with a maximum context length of 262144 tokens.
Pull the tested regular image and download both the main and draft models:
./build-and-copy.sh
./hf-download.sh nvidia/Qwen3.8-27B-NVFP4
./hf-download.sh z-lab/Qwen3.8-27B-DFlash2
Launch the server:
./launch-cluster.sh --solo -t vllm-node \
exec vllm serve nvidia/Qwen3.8-27B-NVFP4 \
--host 0.0.0.0 \
--port 8000 \
--trust-remote-code \
--kv-cache-dtype fp8 \
--gpu-memory-utilization 0.7 \
--max-model-len 262144 \
--max-num-seqs 8 \
--max-num-batched-tokens 16384 \
--enable-chunked-prefill \
--async-scheduling \
--enable-prefix-caching \
--speculative-config '{"method":"dflash","model":"z-lab/Qwen3.8-27B-DFlash2","num_speculative_tokens":8,"draft_tensor_parallel_size":1}' \
--load-format instanttensor \
--reasoning-parser qwen3 \
--tool-call-parser qwen3_xml \
--enable-auto-tool-choice \
--tensor-parallel-size 1
Single Spark: Qwen3.8 Flash Next NVFP4
This matches the solo Qwen3.8 Flash Next recipe, including PLE tables offloaded to disk and speculative decoding. It uses the B12X model loader and a maximum context length of 262144 tokens.
Pull the tested B12X image and download the model:
./build-and-copy.sh --exp-b12x
./hf-download.sh local-inference-lab/Qwen3.8-Flash-Next-NVFP4
Launch the server:
./launch-cluster.sh --solo -t vllm-node-b12x \
-e VLLM_PLE_TABLE_MEMORY=disk \
-e CUTE_DSL_ARCH=sm_121a \
-e SAFETENSORS_FAST_GPU=1 \
-e VLLM_WORKER_MULTIPROC_METHOD=spawn \
-e VLLM_SSM_CONV_STATE_LAYOUT=DS \
-e VLLM_USE_AOT_COMPILE=1 \
-e VLLM_USE_MEGA_AOT_ARTIFACT=1 \
-e VLLM_USE_V2_MODEL_RUNNER=1 \
-e B12X_POLICY_MODE=auto \
exec vllm serve local-inference-lab/Qwen3.8-Flash-Next-NVFP4 \
--host 0.0.0.0 \
--port 8000 \
--trust-remote-code \
--tensor-parallel-size 1 \
--pipeline-parallel-size 1 \
--mamba-cache-mode align \
--enable-prefix-caching \
--enable-chunked-prefill \
--dtype bfloat16 \
--kv-cache-dtype fp8 \
--quantization modelopt_mixed \
--block-size 16 \
--load-format b12x \
--max-model-len 262144 \
--max-num-seqs 8 \
--max-num-batched-tokens 4096 \
--speculative-config '{"method":"mtp","num_speculative_tokens":4}' \
--gdn-decode-kernel b12x \
--linear-backend b12x \
--moe-backend b12x \
--no-enable-flashinfer-autotune \
--mm-encoder-tp-mode data \
--reasoning-parser qwen3 \
--tool-call-parser qwen3_xml \
--enable-auto-tool-choice \
--compilation-config '{"pass_config":{"fuse_act_quant":true}}' \
--gpu-memory-utilization 0.8
Launcher script uses host networking by default. To use Docker port publishing, add
-p 8000:8000 before exec in the command above.
Dual Sparks: DeepSeek V4 Flash Vision Exp
This matches the DeepSeek V4 Flash Vision Exp recipe. It requires two Sparks and uses B12X attention, linear, and MoE backends, FP8 KV cache, and DSpark speculative decoding. The loader mod keeps the primary model on InstantTensor and avoids a second full InstantTensor load for the embedded draft.
Connect the Sparks and configure passwordless SSH as described in the Networking Guide. Run the following on the head node to pull and distribute the B12X image, then download the model once and copy it to the other nodes. The scripts use your saved cluster configuration or autodiscovery.
./build-and-copy.sh --exp-b12x -c --copy-parallel
./hf-download.sh deepseek-ai/DeepSeek-V4-Flash-Vision-Exp -c --copy-parallel
Launch the server from the head node:
./launch-cluster.sh -t vllm-node-b12x \
--apply-mod mods/instanttensor-hybrid-draft-loader \
-e CUTE_DSL_ARCH=sm_121a \
-e VLLM_USE_AOT_COMPILE=1 \
-e VLLM_USE_BREAKABLE_CUDAGRAPH=0 \
-e VLLM_USE_MEGA_AOT_ARTIFACT=1 \
-e VLLM_MEMORY_PROFILE_INCLUDE_ATTN=1 \
-e VLLM_USE_FLASHINFER_SAMPLER=1 \
-e VLLM_USE_B12X_WO_PROJECTION=1 \
-e VLLM_USE_B12X_MHC=1 \
-e VLLM_USE_B12X_FP8_GEMM=1 \
-e VLLM_USE_B12X_MOE=1 \
-e VLLM_USE_B12X_SPARSE_INDEXER=1 \
-e VLLM_USE_V2_MODEL_RUNNER=1 \
-e VLLM_MOE_SKIP_PADDING=0 \
-e B12X_MLA_SM120_UNIFIED=1 \
-e B12X_MOE_FORCE_A8=1 \
exec vllm serve deepseek-ai/DeepSeek-V4-Flash-Vision-Exp \
--host 0.0.0.0 \
--port 8000 \
--trust-remote-code \
--tensor-parallel-size 2 \
--kv-cache-dtype fp8 \
--block-size 256 \
--max-model-len auto \
--max-num-seqs 8 \
--max-num-batched-tokens 8192 \
--gpu-memory-utilization 0.85 \
--enable-prefix-caching \
--tokenizer-mode deepseek_v4 \
--tool-call-parser deepseek_v4 \
--enable-auto-tool-choice \
--reasoning-parser deepseek_v4 \
--reasoning-config '{"reasoning_parser":"deepseek_v4","reasoning_start_str":"","reasoning_end_str":""}' \
--default-chat-template-kwargs.thinking=true \
--default-chat-template-kwargs.reasoning_effort=high \
--load-format instanttensor \
--moe-backend b12x \
--linear-backend b12x \
--attention-backend B12X \
--max-cudagraph-capture-size 48 \
--compilation-config '{"cudagraph_mode":"FULL_AND_PIECEWISE","custom_ops":["all"]}' \
--speculative-config '{"method":"dspark","num_speculative_tokens":6,"draft_sample_method":"probabilistic","attention_backend":"B12X"}'
With --tensor-parallel-size 2, the launcher uses two nodes even if more are
configured. It supplies the distributed backend, node ranks, and coordination
addresses automatically; keep those settings out of the vllm serve command.
Check the server
Wait for vLLM to finish loading and report that the API server is ready. Then, in a second terminal on the serving Spark or cluster head, run:
curl --fail http://localhost:8000/health
curl --fail http://localhost:8000/v1/models
Use http:// as the OpenAI-compatible API base URL in
your client, with the model ID from the example you launched. For solo mode,
use that Spark's IP address.
Build and launch options
The regular image command (./build-and-copy.sh) pulls
eugr/spark-vllm:latest and tags it as vllm-node. Use
./build-and-copy.sh --use-wheels to build the regular runner from precompiled
wheels, or ./build-and-copy.sh --rebuild-vllm to compile upstream vLLM.
The B12X image commands (--exp-b12x) pull eugr/spark-vllm-b12x:latest and
tag it locally as vllm-node-b12x. To compile vLLM from the experimental fork,
add --rebuild-vllm to the corresponding build-and-copy.sh command.
--use-wheels is incompatible with --exp-b12x because experimental vLLM
wheels are not published. See Building the Docker Image
for other build profiles and customizations.
You can add --earlyoom before exec in any launch command to enable the
container's low-memory monitor. See Launching the Cluster
for additional launcher options.
CHANGELOG
2026-10-02
hf-download.sh now supports model cache management locally and across the
cluster: list models and revisions with --list, remove selected models with
--delete, and reclaim space from unreferenced revisions with --cleanup.
Add -c to extend these operations to all selected cluster nodes. It also
enables cluster distribution for downloads and restores, and extends deletion
after backup (--backup --delete) to peer nodes. Backups and backup listings
use the head node's directory.
Listings include model sizes and node locations, with sorting and console,
CSV, JSON, or Markdown output. Use --backup with --backup-dir to save models
to a mounted directory on the head node, --list-backup to inspect backups,
and --restore to restore and optionally distribute them with -c and
--copy-parallel. Backup and deletion accept --revision to select a cached
commit by hash, unique prefix, branch, or tag. --backup --delete removes cached
copies only after a successful backup and coverage checks. Deletion and cleanup
ask for confirmation unless --force is supplied, and missing uvx
installations are now handled automatically on the head and peer nodes.
Backups automatically use regular snapshot files on external drives without
symlink support, with the same listing, restore, and revision selection commands.
2026-09-30
./hf-download.sh will now try to check and automatically repair cache permissions before downloading or distributing the model across the nodes.
2026-09-23
EarlyOOM in 3rd-party containers
--earlyoom flag now works with any Debian/Ubuntu-based vLLM container, such as vllm/vllm-openai or NVIDIA NGC ones.
If EarlyOOM is not installed, it will try to install it via apt.
2026-09-10
Qwen3.8 Flash Next solo PLE disk offload
The solo qwen3.8-flash-next-nvfp4-solo recipe now offloads PLE tables to disk
with VLLM_PLE_TABLE_MEMORY=disk that allows 1M+ k/v cache allocation (the model max context size is still 262144).
Default memory allocation is reduced to 0.8.
2026-09-08
Qwen3.8 Flash Next solo and dual-Spark recipes
Added two recipes for serving
local-inference-lab/Qwen3.8-Flash-Next-NVFP4 with the B12X container.
# Single DGX Spark
./run-recipe.sh qwen3.8-flash-next-nvfp4-solo --solo --earlyoom --setup
Dual DGX Spark cluster
./run-recipe.sh qwen3.8-flash-next-nvfp4-cluster --earlyoom --setup
Deepseek V4 Flash Vision Exp support
B12X container now supports deepseek-ai/DeepSeek-V4-Flash-Vision-Exp.
Run with:
./run-recipe.sh deepseek-v4-flash-vision-exp --setup
2026-09-06
Experimental b12x loader
GLM-5.3 Flash recipe is now using experimental b12x loader that is faster and more memory efficient than Instanttensor on DGX Spark. Also reduced KV-cache memory to 8GB to relax memory pressure. Please note that this recipe and b12x builds in general are still experimental, so please update the repository often to keep everything up to date.
2026-09-05
Full GitHub URLs for vLLM PR patches
All --apply-vllm-pr arguments now accept either the existing numeric
shorthand for vllm-project/vllm or a full public GitHub pull-request URL such
as https://github.com/local-inference-lab/vllm/pull/669. Launch-time patches
download the URL's .diff directly; source builds do the same before applying
the patch to the selected vLLM ref.
2026-09-04
GLM 5.3 Flash dual-Spark recipe
Added the cluster-only glm-5.3-flash recipe for serving
local-inference-lab/GLM-5.3-Flash-NVFP4-Spark on two DGX Spark nodes. It uses the B12X container with MTP and 1M context.
For a first-time cluster setup, discover the nodes and then let the recipe prepare and distribute the B12X image and model:
./run-recipe.sh --discover
./run-recipe.sh glm-5.3-flash --setup
2026-08-27
InstantTensor zero-copy loader mod
Added the opt-in instanttensor-zero-copy mod for memory-constrained model
loads. It disables vLLM's per-tensor InstantTensor ownership clone while
retaining InstantTensor's required ring buffer. Use with caution.
2026-08-25
Local vLLM source checkouts
build-and-copy.sh --vllm-source-dir now builds vLLM from a clean local
Git checkout without requiring the Docker builder to access its remote or host
credentials. An optional --vllm-ref is resolved locally; otherwise the build
uses the checkout's current HEAD. The host checkout is never modified.
2026-08-21
B12X package in regular builds
Regular local runner builds from vllm-project/vllm now build and install the
external B12X package, matching the B12X support being integrated into upstream
vLLM. The experimental --exp-b12x profile continues to install the same
package for its maintained fork.
2026-08-19
Launch-time vLLM PR application
launch-cluster.sh and run-recipe.sh now accept repeatable
--apply-vllm-pr options. The launcher fetches each upstream PR once,
validates that it only changes installed vllm/ runtime files, and applies it
to every newly created container in command-line layer order. PRs that require
native compilation, dependency changes, or packaging changes are rejected with
instructions to use the existing build-time option instead.
Example:
./launch-cluster.sh --solo --apply-vllm-pr 52816 exec vllm serve Inferact/Qwen3.8-27B-NVFP4 --host 0.0.0.0 --port 8000 --trust-remote-code --kv-cache-dtype fp8 --gpu-memory-utilization 0.7 --max-model-len 262144 --max-num-seqs 8 --max-num-batched-tokens 16384 --enable-chunked-prefill --async-scheduling --enable-prefix-caching --speculative-config '{"method":"dflash","model": "z-lab/Qwen3.8-27B-DFlash2", "num_speculative_tokens":8}' --load-format instanttensor --reasoning-parser qwen3 --tool-call-parser qwen3_xml --enable-auto-tool-choice
2026-08-14
B12X source branch update
B12X source builds now use local-inference-lab/vllm@dev/infernal-invocation.
2026-08-12
PyTorch 2.13 and CUTLASS DSL 4.7
Regular and B12X builds now use PyTorch 2.13.0, torchvision 0.28.0, torchaudio 2.11.0, and CUTLASS DSL 4.7.0.
2026-08-11
Nemotron 3.5 Lightning recipe
Added a recipe for NVIDIA's Nemotron 3.5 Lightning with DSpark and 1M context.
Quick start:
git pull
./hf-download.sh nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 # add -c for cluster
./hf-download.sh nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4-DSpark # speculator model; add -c for cluster
./run-recipe.sh recipes/nemotron-3.5-lightning.yaml --solo
2026-08-06
GLM-5.2 NVFP4 8x Spark cluster recipe
Added the cluster-only recipes/8x-spark-cluster/glm-5.2-nvfp4.yaml recipe for serving nvidia/GLM-5.2-NVFP4. Requires 8x nodes. The recipe uses the vllm-node-b12x image.
./run-recipe.sh recipes/8x-spark-cluster/glm-5.2-nvfp4.yaml
2026-08-03
New B12X image
Added --exp-b12x (alias: --experimental-b12x) as an alternative version built from a fork by Luke Alonso. This fork supports a collection of experimental high-performance B12X kernels for sm12x architecture.
Since it is built from a forked vLLM branch, it will be supported in parallel to the main ("regular") build, at least for the time being.
Specifying --exp-b12x without arguments will pull eugr/spark-vllm-b12x:latest from Dockerhub, which is now built and tested together with the main image by CI pipeline on nightly basis (if there are any updates to the source branch or B12X).
Add --rebuild-vllm to compile from the source.
The preset defaults the image tag to vllm-node-b12x; an explicit -t still takes precedence. Additional vLLM changes can be layered onto the preset with one or more --apply-vllm-pr flags. PRs are applied from the upstream vLLM repository, not from Luke's fork! If any PRs are specified, the script will build from the source, not prebuilt images.
vLLM wheels are not published for this build, but it will reuse published Flashinfer wheels if you build from the source (unless --rebuild-flashinfer is specified).
For now, we have only one recipe using this build, with more to come.
DeepSeek V4 Flash 0731 B12X cluster recipe
Added the cluster-only deepseek-v4-flash-0731 recipe for serving deepseek-ai/DeepSeek-V4-Flash-0731 on a dual DGX Spark cluster. The recipe requires B12X container (vllm-node-b12x) that can be pulled by using ./build-and-copy.sh --exp-b12x -c (or just allow the recipe system to pull it for you). See details on b12x build above.
./run-recipe.sh deepseek-v4-flash-0731
2026-07-30
Inkling Small NVFP4 support
Added the support for
thinkingmachines/Inkling-Small-NVFP4. It requires at least dual DGX Spark setup.
./run-recipe.sh inkling-small-nvfp4
Add --setup on the first run to prepare the container and download and
distribute the model.
Official vLLM image support with earlyoom, InstantTensor, and SciPy
mods/use-official-vllm now installs earlyoom, InstantTensor, and SciPy in
addition to the compatibility packages needed by other mods. This makes
launch-cluster.sh --earlyoom, --load-format instanttensor, and SciPy-based
functionality available when launching official vLLM images such as
vllm-openai. The Python package install pins the image's existing Torch
packages so the CUDA-enabled build is not replaced during dependency
resolution.
Repeatable Docker volume mappings
launch-cluster.sh and run-recipe.sh now accept repeatable -v / --volume mappings using Docker's local_path:container_path syntax. In cluster mode, each mapping is applied to every launched node.
./launch-cluster.sh --solo \
-v "$PWD/models:/models" \
exec vllm serve /models/qwen3.6-35b-a4b-nvfp4 ...
2026-07-14
Custom vLLM repositories and PyTorch versions
build-and-copy.sh can now build vLLM from a fork with --vllm-repo. Custom repositories bypass the shared upstream Git checkout cache, force a vLLM source build, and suppress the Dockerfile's upstream preset PRs unless --apply-preset-vllm-prs is explicitly requested.
--torch-version, --torchvision-version, and --torchaudio-version select the packages installed in both the source-build environment and final runner image. The defaults are PyTorch 2.13.0, torchvision 0.28.0, and torchaudio 2.11.0, matching current upstream vLLM CUDA requirements.
Builds from any ref in local-inference-lab/vllm also clone and build the master ref of lukealonso/b12x automatically. The repository produces the b12x distribution. The source layer is refreshed on every applicable runner build so a previously cached clone cannot hide newer upstream commits. Only the locally built B12X wheel is installed. Its CUTLASS DSL metadata is adjusted to the image-wide 4.7.0 pin before installation. The checked-out commit is recorded at /workspace/b12x-source-commit in the image. B12X kernels remain JIT-compiled at runtime; building its Python wheel does not add another CUDA compilation phase to the image build.
2026-07-10
Optional earlyoom monitor
Runner images now include earlyoom, and launch-cluster.sh --earlyoom can run it as the container foreground process instead of sleep infinity. This keeps the existing launcher flow, where Ray and vLLM are started with docker exec, while allowing the container to monitor low host memory continuously. run-recipe.py and run-recipe.sh pass the same --earlyoom and --earlyoom-args options through to the launcher.
The default policy is conservative and absolute-memory based: -M 524288,102400 -s 100 -r 60. That sends SIGTERM when available memory drops below 512 MiB, escalates to SIGKILL below 100 MiB, does not wait for swap to fill before acting, and prints one memory report per minute. You can override the default per launch with --earlyoom-args or set VLLM_SPARK_EARLYOOM_ARGS.
Build and runtime compatibility updates
--gpu-arch now also drives the NCCL NVCC_GENCODE build argument, so non-default local builds compile NCCL for the same target architecture as Torch and FlashInfer.
The source-build KV-cache cleanup is now embedded directly in the Dockerfile, while mods/kv-cache-prealloc-cleanup keeps only the runtime policy tweaks.
mods/gpu-mem-util-gb was refreshed against current vLLM memory profiling code so fixed-GiB GPU memory reservations continue to work with the newer startup path.
2026-07-02
Prebuilt runner image by default
build-and-copy.sh now pulls prebuilt eugr/spark-vllm:latest by default and tags it locally as vllm-node or the tag requested with -t. The latest tag points at the latest tested nightly image. The prebuilt image is updated at the same time as prebuilt wheels by the CI pipeline, so they all stay in sync.
Use --use-wheels to keep the previous wheel-based runner build path. Build customization flags such as --exp-mxfp4, non-default --gpu-arch, --vllm-ref, --flashinfer-ref, rebuild/download flags, and PR application flags also keep the local build path. --tf5 remains a tag-compatibility alias and pulls the prebuilt image as vllm-node-tf5.
Copy now checks the image ID locally and on each remote host before saving the image. Hosts that already have the same image ID are skipped, and docker save is skipped entirely when every target is already current. --no-build still skips image preparation and only copies an already-local tag when needed.
2026-07-01
No-Ray is now the default multi-node backend
launch-cluster.sh and run-recipe.sh now default to no-Ray multi-node launches. Use --ray to opt into Ray; Ray mode ensures vLLM commands include --distributed-executor-backend ray when they omit it. --no-ray remains accepted for compatibility in multi-node launches.
Build and dependency updates
We now use NCCL main branch and include new experimental vLLM Rust frontend in the builds. DeepGEMM now tracks the nv_dev branch.
Transformers 5 flag deprecation
--tf5, --pre-tf, and --pre-transformers are now deprecated compatibility aliases. They no longer override dependency resolution; they only preserve the legacy default image tag. Recipes that previously used vllm-node-tf5 now use the standard vllm-node image.
Recipe updates
Added the gemma4-26b-a4b-nvfp4 recipe, reverted the Qwen3.6-35B-A3B-NVFP4 recipes to --kv-cache-dtype fp8, and cleaned up stale TF5 build args from affected recipes.
2026-06-22
Deepseek V4 Flash support
Support and a recipe for Deepseek V4 Flash has been added based on a newly merged vLLM PR. Please note that this PR requires the DeepGEMM nv_dev branch, which is not present in default vLLM builds, but included in this community build.
To run DSV4F you will need a Spark cluster (2 or more nodes).
To run:
git pull
./build-and-copy.sh -c
./hf-download.sh deepseek-ai/DeepSeek-V4-Flash -c
./run-recipe.sh deepseek-v4-flash --no-ray
DeepGEMM support added
vLLM now uses NVIDIA branch for DeepGEMM that includes support for sm12x GPU family, including DGX Spark.
Please note that it is currently not compatible with NVRTC compiler, so DG_JIT_USE_NVRTC has been turned off for new builds.
MiniMax AWQ Weight-Shape Loader Workaround
Added a Dockerfile-level workaround for a vLLM nightly regression where compressed-tensors MoE weight_shape metadata is treated as a scalar during load, causing affected MoE quantized models to fail with shape '[]' is invalid for input of size 2.
2026-06-18
KV Cache Preallocation Cleanup Mod & Updated Qwen3.5-397B recipe for dual Sparks
Added KV cache preallocation cleanup for source-built vLLM wheels, which clears cached CUDA allocator memory before vLLM sizes and allocates KV cache blocks. The dual-node Qwen3.5-397B INT4 AutoRound recipe also applies mods/kv-cache-prealloc-cleanup after mods/gpu-mem-util-gb for its model-specific memory policy tweaks.
The same recipe now keeps the 108 GiB startup reservation, sets VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS=0, and uses a manual 2.25 GiB KV-cache allocation to bypass vLLM's conservative profiler-derived KV budget. The source build handles profiling-only graph-pool cleanup before KV-cache sizing, while the recipe mod makes the env var skip CUDA graph memory profiling entirely and allows the fixed-GiB reservation argument to coexist with --kv-cache-memory-bytes for this model to load. This may result in a bit of swap to be used to offload unused resources, so make sure swap is enabled.
Added recipes for nvidia/Qwen3.6-35B-A3B-NVFP4
Added recipes for NVIDIA's mixed-precision quant for Qwen3.6-35B model. It has a higher precision and an improved performance compared to the earlier NVFP4 quants.
Use:
qwen3.6-35b-a3b-nvfp4to run with MTP onqwen3.6-35b-a3b-nvfp4-no-mtpto run without MTP.
2026-06-10
DiffusionGemma Recipes and Mod
Added day0 support for Google DeepMind's DiffusionGemma model via mods/diffusiongemma. Check out NVIDIA blog for details!
Added four solo-only DiffusionGemma recipes:
diffusion-gemma-bf16-thinkingforgoogle/diffusiongemma-26B-A4B-itwith thinking enabled.diffusion-gemma-bf16forgoogle/diffusiongemma-26B-A4B-itwith thinking disabled.diffusion-gemma-nvfp4-thinkingfornvidia/diffusiongemma-26B-A4B-it-NVFP4with thinking enabled.diffusion-gemma-nvfp4fornvidia/diffusiongemma-26B-A4B-it-NVFP4with thinking disabled.
--reasoning-parser gemma4, since these models can emit Gemma4 channel markers even when thinking is disabled.
Example:
./hf-download.sh google/diffusiongemma-26B-A4B-it
./run-recipe.sh diffusion-gemma-bf16-thinking --solo
run-recipe.sh Launch Flag Passthrough
run-recipe.sh now passes additional launch-cluster.sh flags through when running recipes:
--apply-mod, -p / --publish, and --keep-entrypoint.
Port publishing is still solo-only, matching launch-cluster.sh behavior.
2026-06-09
Recipe Memory Defaults
Raised the default gpu_memory_utilization from 0.7 to 0.8 across the main single-node and two-node recipes to match the current vLLM memory allocation behavior.
2026-06-07
Docker Base Image Compatibility
The default CUDA base image was changed to nvidia/cuda:13.0.2-devel-ubuntu24.04 for broader host compatibility.
The Dockerfile also now passes --allow-change-held-packages when installing the custom NCCL Debian packages, avoiding apt failures when replacing held CUDA/NCCL packages during image builds.
2026-06-06
MiniMax Multi-Node Regression Workaround
Added a targeted Dockerfile patch that disables the MiniMax QK RMSNorm CUDA IPC fused path introduced by vLLM PR #43410. The fused path can fail when tensor parallelism spans DGX Spark nodes; the workaround preserves MiniMax multi-node TP while avoiding a full upstream revert.
2026-06-03
Solo Port Publishing
launch-cluster.sh now supports Docker-style -p / --publish port mappings in solo mode. When port publishing is used, the launcher switches from host networking to Docker bridge networking for that solo container.
Example:
./launch-cluster.sh --solo -p 8000:8000 exec vllm serve ...
2026-05-29
Wheel Freshness Detection
Improved build-and-copy.sh wheel freshness checks so newer locally built wheels are not overwritten just because their filenames differ from the latest release assets. The script now compares local wheel mtimes against remote release asset timestamps before deciding to download. Also switched from using GitHub API to regular HTTP checks to avoid throttling.
gpu-mem-util-gb Patch Refresh
Refreshed mods/gpu-mem-util-gb so it applies against newer vLLM CacheConfig code after upstream line/context drift.
2026-05-28
StepFun Step 3.7 Flash Support
Added support for StepFun Step 3.7 Flash multimodal model.
Requires at least 2 Sparks in a cluster. Both FP8 and NVFP4 checkpoints are supported. FP8 requires more memory, so using NVFP4 is recommended.
Update the repo and build a fresh container first:
git pull
./build-and-copy.sh --cleanup -c
To run NVFP4 version:
Download the model:
./hf-download.sh stepfun-ai/Step-3.7-Flash-NVFP4 -c
Run:
./run-recipe.sh step-3.7-flash-nvfp4 --no-ray
To run FP8 version:
Download the model:
./hf-download.sh stepfun-ai/Step-3.7-Flash-FP8 -c
Run:
./run-recipe.sh step-3.7-flash-fp8 --no-ray
Please note that --no-ray is required for FP8 to fit with full context!
use-official-vllm NCCL Workaround
Updated mods/use-official-vllm to also handle the NCCL load-order bug tracked in vllm-project/vllm#42354. When both the pip-installed nvidia/nccl/lib/libnccl.so.2 and system libnccl2 are present, the mod redirects the pip-installed NCCL path to the system /usr/lib soname, matching the manual workaround that fixes multi-node DGX Spark hangs.
Use it with official vLLM images before starting the model:
./launch-cluster.sh -t vllm/vllm-openai:latest \
--apply-mod mods/use-official-vllm \
exec vllm serve ...
Torch Pinning During Wheel Install
Pinned the already-installed CUDA torch build via uv --override when installing locally built wheels and final runtime dependencies in both Dockerfiles. This prevents transitive dependencies from re-resolving torch to a CPU wheel during image builds.
2026-05-22
New Mod: use-official-vllm
Added mods/use-official-vllm, a prerequisite mod for applying patches inside official vLLM Docker containers (e.g. vllm-openai). Official containers do not ship git, which several mods require. This mod installs git via apt-get if it is not already present.
Apply it before any other mod that requires git:
./launch-cluster.sh -t vllm/vllm-openai:latest \
--apply-mod mods/use-official-vllm \
--apply-mod mods/gpu-mem-util-gb \
exec vllm serve ...
gpu-mem-util-gb Updated for Latest vLLM Main
Updated mods/gpu-mem-util-gb patch to apply cleanly against the current vLLM main branch. The mod now checks for git at startup and prints a hint to apply mods/use-official-vllm first if git is missing (relevant when using official vLLM containers).
2026-05-18
NCCL Updated to NVIDIA v2.30u1
The Dockerfile now builds NCCL from NVIDIA's v2.30u1 branch instead of the custom NCCL fork. This branch incorporates all features of the custom fork and is based on the latest NCCL release. The networking guide's NCCL test commands have been updated to use the same branch for 3-node mesh clusters.
2026-05-14
Default Entrypoint Clearing
launch-cluster.sh now clears the Docker image entrypoint by default when starting idle cluster containers. This allows images with server-style entrypoints, such as vllm-openai, to work with the same cluster launcher flow. Use --keep-entrypoint to preserve the image entrypoint.
2026-05-10
Qwen3.5-397B Recipe Memory Updates
Updated Qwen3.5-397B AutoRound recipes to reduce OOM risk. The dual-node recipe now uses standard fractional --gpu-memory-utilization, and the 3-node pipeline-parallel recipe uses InstantTensor loading with lower memory pressure.
2026-05-06
Qwen3.6-35B-A3B-FP8 Recipes
Added qwen3.6-35b-a3b-fp8 and qwen3.6-35b-a3b-fp8-dflash recipes plus a dedicated Qwen3.6 chat-template mod. The DFlash recipe has prefix caching disabled because it caused accuracy issues.
2026-04-29
Gemma4 Recipe Fixes and Experimental b12x Mod
The Gemma4-26B-A4B recipe now uses safetensors loading and no longer applies the obsolete tool parser mod by default.
2026-04-25
MiniMax-M2.7-AWQ Recipe
Added minimax-m2.7-awq, a cluster-only MiniMax M2.7 AWQ recipe using cyankiwi/MiniMax-M2.7-AWQ-4bit.
2026-04-14
Added --load-format instanttensor support to vLLM - thanks @SeraphimSerapis.
An experimental option for now, but allows for faster loading than the current fastsafetensors default. You need to rebuild the container to start using the option, but you don't have to trigger the source build.
2026-04-12
Drop-caches mod for Qwen3.5-397B
Updated Qwen3.5-397B recipe (for dual node configuration) to use the new mod mods/drop-caches which clears filesystem caches every minute while the container is running, resolving fastsafetensors getting stuck during loading and a few other bugs when operating close to max memory limit.
2026-04-11
Pinned PyTorch Version
Pinned PyTorch to version 2.11.0 (previously using nightly builds) to fix incompatibility with transformers 5.x and avoid torch version mismatch in builds.
2026-04-02
A new recipe for Gemma4-26B-A4B in "on-the-fly" FP8 quantization:
Single Spark:
./run-recipe.sh gemma4-26b-a4b --solo
Dual Sparks:
./run-recipe.sh gemma4-26b-a4b --no-ray
2026-03-31
Flags to specify Flashinfer ref and apply PRs
build-and-copy.sh gains two new flags that mirror the existing vLLM equivalents:
--flashinfer-ref— build FlashInfer from a specific commit SHA, branch, or tag instead ofmain. Forces a local FlashInfer build (skips prebuilt wheel download).--apply-flashinfer-pr— fetch and apply a FlashInfer GitHub PR patch before building. Can be specified multiple times. Forces a local FlashInfer build.
--exp-mxfp4.
Default image tag in build-and-copy.sh
build-and-copy.sh now automatically sets a sensible default image tag when -t is not specified:
--tf5/--pre-tf- deprecated compatibility flag; normal build, tag defaults tovllm-node-tf5--exp-mxfp4- tag defaults tovllm-node-mxfp4- in all other cases - tag defaults to
vllm-node(no change)
-t always takes precedence.
Support for 3-node mesh setups
Added initial support for setups where 3 Sparks are connected in a ring-like mesh without an additional switch. See Networking Guide for instructions on how to connect and set up networking in such cluster.
Autodiscover function in both launch-cluster.sh and run-recipe.sh now can detect mesh setups and configure parameters accordingly.
You can try running a model on all 3 nodes in pipeline-parallel configuration using the following recipe:
./run-recipe.sh --discover # force mesh discovery
./run-recipe.sh recipes/3x-spark-cluster/qwen3.5-397b-int4-autoround.yaml --setup --no-ray --force-build # you can drop --setup and --force-build on subsequent calls
Please note that --tensor-parallel-size 3 or -tp 3 is not supported by any commonly used model, so the only two viable options to utilize all three nodes for a single model are:
--pipeline-parallel 3will let you run a model that can't fit on dual Sparks, but without additional speed improvements (total throughtput may improve though).--data-parallel 3(possibly with--enable-expert-parallel) will let you run a model that can fit on a single Spark, but allow for better concurrency.
--tensor-parallel 2 in a 3-node configuration - in this case only first two nodes (from autodiscovery/.env or from the CLI parameters) will be utilized.
GB10 Verification During Node Discovery
Node discovery now confirms each SSH-reachable peer is a GB10 system before adding it to the cluster:
Only hosts reporting NVIDIA GB10 are included. This prevents accidentally adding non-Spark machines that happen to be on the same subnet.
Separate COPY_HOSTS Discovery
Autodiscover now determines the host list used for image and model distribution separately from CLUSTER_NODES:
- Non-mesh:
COPY_HOSTSmirrorsCLUSTER_NODES(no change in behaviour). - Mesh: scans the direct IB-attached
enp1s0f0np0andenp1s0f1np1interfaces (not the OOB ETH interface), so large file transfers use the faster direct InfiniBand path.
COPY_HOSTS is saved to .env and respected by build-and-copy.sh, hf-download.sh, and run-recipe.py.
Interactive Configuration Save in autodiscover.sh
autodiscover.sh now handles .env creation with a guided interactive flow, replacing the previous logic in run-recipe.py:
- Runs automatically when
.envis absent. - Asks per-node confirmation for both
CLUSTER_NODESandCOPY_HOSTS. - Skips if
.envalready exists (use--setupto force).
run-recipe.py no longer contains its own .env-save prompt — it delegates entirely to autodiscover.sh.
--setup Flag in launch-cluster.sh and build-and-copy.sh
Both scripts now accept --setup to force a full autodiscovery run and overwrite the existing .env:
./launch-cluster.sh --setup exec vllm serve ...
./build-and-copy.sh --setup -c
This is equivalent to the existing --setup in run-recipe.sh.
--config Flag
hf-download.sh, build-and-copy.sh and launch-cluster.sh now accept --config to load a custom .env configuration file. COPY_HOSTS from the config is used for model distribution:
./hf-download.sh QuantTrio/MiniMax-M2-AWQ --config /path/to/cluster.env -c --copy-parallel
Parallelism-Aware Node Trimming
launch-cluster.sh now parses -tp / --tensor-parallel-size, -pp / --pipeline-parallel-size, and -dp / --data-parallel-size from the exec command or launch script and adjusts the active node count accordingly — for both Ray and no-Ray modes.
- If fewer nodes are needed than configured, only the required nodes get containers started (excess nodes are left idle).
- If more nodes are needed than available, an error is raised before anything starts.
Note: Command requires 2 node(s) (tp=2 pp=1 dp=1); using 2 of 3 configured node(s).
Error: Command requires 4 nodes (tp=4 pp=1 dp=1) but only 3 node(s) are configured.
No flags required — the check is automatic whenever parallelism arguments are present in the command.
2026-03-18
--master-port / --head-port Parameter
Added --master-port (synonym: --head-port) to both launch-cluster.sh and run-recipe.sh to configure the port used for cluster coordination:
- In Ray mode: sets the Ray head node port (previously hardcoded to 6379)
- In No-Ray mode: sets the PyTorch distributed
--master-portpassed to vLLM
29501.
./launch-cluster.sh --master-port 29501 --no-ray exec vllm serve ...
./run-recipe.sh qwen3.5-122b-fp8 --no-ray --master-port 29501
--network Parameter in Build Arguments
Added --network to build-and-copy.sh to allow using host networking during builds.
Thanks @apairmont for the PR.
2026-03-17
EXPERIMENTAL Intel/Qwen3.5-397B-A17B-int4-AutoRound Recipe
You can run full 397B Qwen3.5 model on just two Sparks with vision and full context, however you need to make sure your Sparks don't run anything extra that can take a lot of RAM. That means that you don't want to log into the graphical interface or use remote desktop. Connect to the head node via ssh.
Alternatively, you can run in non-graphical mode (runlevel 3) by using sudo systemctl isolate multi-user.target to switch (you can use sudo systemctl set-default graphical.target to switch back to graphical mode), however this is known to reduce performance a bit.
You can run the model with the following command on the head node:
./run-recipe.sh qwen3.5-397b-int4-autoround.yaml --no-ray
Please, note --no-ray is necessary to fit full context. It also improves inference speed by ~1 t/s.
By default it will try to allocate 108 GiB for vLLM on each node. You can change this by changing gpu_memory_utilization in the recipe or passing --gpu-mem; this recipe maps that value to --gpu-memory-utilization-gb, so it is GiB rather than a percentage.
KNOWN ISSUES:
- The current firmware may cause sudden shutdown event on one or both Sparks during heavy inference. If you have this issue, you will need to lower GPU clock frequency on the affected unit(s), e.g.
sudo nvidia-smi -lgc 200,2150. This command will reduce max GPU frequency to 2150 MHz. You can play with higher values to see what works for you (default is 2411 MHz, but can boost to 3000 MHz). Please note that this setting only survives until the next reboot, but can be applied at any time. - You will need to use the new
--no-rayargument to fit full context. - If the model gets stuck loading weights, clearing the cache on both nodes can "unstuck" it. Use
sudo sh -c 'sync; echo 3 > /proc/sys/vm/drop_caches'to clear the cache.
Major Cluster Orchestration Refactoring
Significantly refactored the internal cluster startup logic in launch-cluster.sh:
- Removed the standalone
run-cluster-node.shscript; its logic is now fully integrated intolaunch-cluster.sh. - Ray head/worker startup, environment variable injection, and launch script distribution are now handled by
launch-cluster.shdirectly. - Worker containers are started with proper per-node environment variables (
VLLM_HOST_IP,NCCL_SOCKET_IFNAME, etc.) injected viadocker run/docker execinstead of relying on.bashrc. - You will now be able to run other vLLM containers without applying
use-ngc-vllmmod (current version is just an empty stub).
No-Ray Multi-Node Mode
Added --no-ray flag to launch-cluster.sh to run multi-node vLLM clusters without Ray, using PyTorch's native distributed backend instead. It slightly improves inference performance for most models and reduces memory requirements.
./launch-cluster.sh --no-ray exec vllm serve ...
--no-ray is incompatible with --solo (which already runs without Ray).
run-recipe.sh No-Ray Mode and Extended Flag Passthrough
run-recipe.sh now supports --no-ray flag for running multi-node inference without Ray (uses PyTorch distributed backend instead):
./run-recipe.sh qwen3.5-122b-fp8 --no-ray
The following launch-cluster.sh flags are now also passed through from run-recipe.sh:
--master-port, --name, --eth-if, --ib-if, -j, --no-cache-dirs, --non-privileged, --mem-limit-gb, --mem-swap-limit-gb, --pids-limit, --shm-size-gb.
Nemotron-3-Nano-NVFP4 Switched to Marlin Backend
The nemotron-3-nano-nvfp4 recipe has been updated to use the Marlin backend for better performance and reliability (until Flashinfer fully supports NVFP4 on sm121).
2026-03-12
Experimental --gpu-memory-utilization-gb Mod
Added a new mod mods/gpu-mem-util-gb that adds a --gpu-memory-utilization-gb flag to vLLM, allowing you to specify GPU memory reservation in GiB instead of as a fraction. This is particularly useful on DGX Spark's unified memory architecture where available memory changes dynamically.
./launch-cluster.sh --apply-mod mods/gpu-mem-util-gb exec vllm serve ... \
--gpu-memory-utilization-gb 110
Cannot be used simultaneously with --kv-cache-memory-bytes.
Qwen3.5-397B INT4-AutoRound TP=4 Recipe (4× Spark Cluster)
Added recipes/4x-spark-cluster/qwen3.5-397b-int4-autoround.yaml for running Intel/Qwen3.5-397B-A17B-int4-AutoRound across 4 DGX Spark nodes with tensor parallelism (TP=4).
Benchmarked at ~37 tok/s single-user, ~103 tok/s aggregate (4 concurrent users).
Includes a new mod mods/fix-qwen35-tp4-marlin that resolves a Marlin kernel constraint (MIN_THREAD_N=64) that breaks certain projection layers at TP=4.
Note: Requires NVIDIA driver 580.x. Driver 590.x has a CUDAGraph capture deadlock on GB10 unified memory.
./run-recipe.sh 4x-spark-cluster/qwen3.5-397b-int4-autoround
Thanks @sonusflow for the contribution.
Nemotron-3-Super-120B NVFP4 Recipe
Added a new recipe nemotron-3-super-nvfp4 for running nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 with Marlin kernels. Supports both solo and cluster modes. Includes a custom reasoning parser (super_v3_reasoning_parser.py) fetched from the model reposi
... (README truncated for length)