Profile
Back to NewsBack
GitHub Trending 34 min
Reader Mode
eugr/spark-vllm-docker: Docker configuration for running VLLM on dual DGX Sparks

eugr/spark-vllm-docker: Docker configuration for running VLLM on dual DGX Sparks

7 hours ago

vLLM Docker Optimized for DGX Spark (single or multi-node)

This repository contains the Docker configuration and startup scripts to run vLLM on DGX Spark, from a single node to multi-node clusters using Ray or vLLM's native PyTorch distributed mode. It supports InfiniBand/RDMA (NCCL), custom environment configuration, and high-performance model loading through fastsafetensors and InstantTensor. Cluster setup supports direct connections between dual Sparks, QSFP/RoCE switch configurations, and 3-node mesh configurations.

While it was primarily developed to support multi-node inference, it works just as well on single-node setups.

If you point an AI agent at this repository to prepare a host or run a recipe, use the operational agent runbook. Agent-assisted repository changes use the separate development guide; AGENTS.md routes agents to the appropriate guide.

Table of Contents

DISCLAIMER

This repository is not affiliated with NVIDIA or their subsidiaries. This is a community effort aimed to help DGX Spark users to set up and run the most recent versions of vLLM on Spark cluster or single nodes.

By default, build-and-copy.sh pulls the tested nightly runner image from DockerHub: eugr/spark-vllm:latest. Nightly images are built and tested on multiple models in both cluster and solo configuration before latest is advanced. We will expand the selection of models we test in the pipeline, but since vLLM is a rapidly developing platform, some things may break.

Selecting --exp-b12x without local-build flags or customizations pulls the separately tested eugr/spark-vllm-b12x:latest image and tags it as vllm-node-b12x.

If you want to build only the runner from precompiled vLLM and FlashInfer wheels, specify --use-wheels. This option never falls back to compiling missing wheels: if a wheel cannot be downloaded or found locally, the command stops with an error. To build the latest vLLM from the main branch, use --rebuild-vllm; to target a specific repository, release, or commit, set --vllm-repo and/or --vllm-ref. To build a private or already-available checkout without cloning it inside Docker, use --vllm-source-dir.

Similarly, --rebuild-flashinfer, --flashinfer-ref, and --apply-flashinfer-pr control the FlashInfer build and force the local build path.

Local builds include a CUDA-on-WSL memory-reporting fix in both compiled vLLM wheels and runners built with --use-wheels. On integrated NVIDIA GPUs under WSL, vLLM keeps CUDA's reported free memory instead of replacing it with guest RAM availability. Native Linux UMA accounting and proactive allocator-cache release keep their upstream behavior.

Source builds also trim unused glibc CPU heap pages after startup garbage collection in API servers and workers. This complements the existing CUDA allocator cleanup before KV cache sizing/allocation. The CPU trim is included in exported vLLM wheels and runs after warmup; platforms without malloc_trim skip it.

Runtime images also set VLLM_WSL2_ENABLE_PIN_MEMORY=1 by default. Pass -e VLLM_WSL2_ENABLE_PIN_MEMORY=0 to launch-cluster.sh or docker run to opt out.

Regular and B12X source builds include a targeted patch based on vLLM PR #58028. When --api-key is configured, the Python frontend requires authentication for every route except /health, /ping, /load, and /version; CORS preflight requests also remain exempt. This includes protecting /metrics, /tokenize, and the API docs. The patch is included in exported vLLM wheels and leaves the Rust frontend unchanged.

QUICK START (USING RECIPES)

Single Spark

Check out locally. Do it on the head node of the cluster. This will build the image, download the model and launch it in the container.

git clone https://github.com/eugr/spark-vllm-docker.git
cd spark-vllm-docker
./run-recipe.sh recipes/qwen3.8-flash-next-nvfp4-solo.yaml --solo --setup

Dual Sparks (or more)

Before you start, make sure you connect your Sparks together and enable passwordless SSH as described in our Networking Guide. You can also check out NVIDIA's Connect Two Sparks Playbook, but using our guide is the best way to get started. The guide includes instructions for 3-node Spark mesh clusters.

Check out locally. Do it on the head node of the cluster. This will build the image, download and distribute the model and launch the cluster.

git clone https://github.com/eugr/spark-vllm-docker.git
cd spark-vllm-docker
./run-recipe.sh recipes/deepseek-v4-flash-vision-exp.yaml --setup

QUICK START (USING LAUNCHER)

These examples show image preparation, model download, and launch-cluster.sh commands separately. Qwen3.8-27B uses the regular image (vllm-node). The Qwen3.8 Flash Next and DeepSeek Vision examples match the shortcut above and use the B12X image (vllm-node-b12x). Choose the example for your setup.

Check out the repository

Run the commands on your Spark, or on the head node of your cluster:

git clone https://github.com/eugr/spark-vllm-docker.git
cd spark-vllm-docker

Single Spark: Qwen3.8-27B NVFP4 (regular image)

This matches the Qwen3.8-27B NVFP4 with DFlash2 recipe in solo mode. It uses FP8 KV cache, InstantTensor loading, and the DFlash2 draft model for speculative decoding, with a maximum context length of 262144 tokens.

Pull the tested regular image and download both the main and draft models:

./build-and-copy.sh
./hf-download.sh nvidia/Qwen3.8-27B-NVFP4
./hf-download.sh z-lab/Qwen3.8-27B-DFlash2

Launch the server:

./launch-cluster.sh --solo -t vllm-node \
  exec vllm serve nvidia/Qwen3.8-27B-NVFP4 \
    --host 0.0.0.0 \
    --port 8000 \
    --trust-remote-code \
    --kv-cache-dtype fp8 \
    --gpu-memory-utilization 0.7 \
    --max-model-len 262144 \
    --max-num-seqs 8 \
    --max-num-batched-tokens 16384 \
    --enable-chunked-prefill \
    --async-scheduling \
    --enable-prefix-caching \
    --speculative-config '{"method":"dflash","model":"z-lab/Qwen3.8-27B-DFlash2","num_speculative_tokens":8,"draft_tensor_parallel_size":1}' \
    --load-format instanttensor \
    --reasoning-parser qwen3 \
    --tool-call-parser qwen3_xml \
    --enable-auto-tool-choice \
    --tensor-parallel-size 1

Single Spark: Qwen3.8 Flash Next NVFP4

This matches the solo Qwen3.8 Flash Next recipe, including PLE tables offloaded to disk and speculative decoding. It uses the B12X model loader and a maximum context length of 262144 tokens.

Pull the tested B12X image and download the model:

./build-and-copy.sh --exp-b12x
./hf-download.sh local-inference-lab/Qwen3.8-Flash-Next-NVFP4

Launch the server:

./launch-cluster.sh --solo -t vllm-node-b12x \
  -e VLLM_PLE_TABLE_MEMORY=disk \
  -e CUTE_DSL_ARCH=sm_121a \
  -e SAFETENSORS_FAST_GPU=1 \
  -e VLLM_WORKER_MULTIPROC_METHOD=spawn \
  -e VLLM_SSM_CONV_STATE_LAYOUT=DS \
  -e VLLM_USE_AOT_COMPILE=1 \
  -e VLLM_USE_MEGA_AOT_ARTIFACT=1 \
  -e VLLM_USE_V2_MODEL_RUNNER=1 \
  -e B12X_POLICY_MODE=auto \
  exec vllm serve local-inference-lab/Qwen3.8-Flash-Next-NVFP4 \
    --host 0.0.0.0 \
    --port 8000 \
    --trust-remote-code \
    --tensor-parallel-size 1 \
    --pipeline-parallel-size 1 \
    --mamba-cache-mode align \
    --enable-prefix-caching \
    --enable-chunked-prefill \
    --dtype bfloat16 \
    --kv-cache-dtype fp8 \
    --quantization modelopt_mixed \
    --block-size 16 \
    --load-format b12x \
    --max-model-len 262144 \
    --max-num-seqs 8 \
    --max-num-batched-tokens 4096 \
    --speculative-config '{"method":"mtp","num_speculative_tokens":4}' \
    --gdn-decode-kernel b12x \
    --linear-backend b12x \
    --moe-backend b12x \
    --no-enable-flashinfer-autotune \
    --mm-encoder-tp-mode data \
    --reasoning-parser qwen3 \
    --tool-call-parser qwen3_xml \
    --enable-auto-tool-choice \
    --compilation-config '{"pass_config":{"fuse_act_quant":true}}' \
    --gpu-memory-utilization 0.8

Launcher script uses host networking by default. To use Docker port publishing, add -p 8000:8000 before exec in the command above.

Dual Sparks: DeepSeek V4 Flash Vision Exp

This matches the DeepSeek V4 Flash Vision Exp recipe. It requires two Sparks and uses B12X attention, linear, and MoE backends, FP8 KV cache, and DSpark speculative decoding. The loader mod keeps the primary model on InstantTensor and avoids a second full InstantTensor load for the embedded draft.

Connect the Sparks and configure passwordless SSH as described in the Networking Guide. Run the following on the head node to pull and distribute the B12X image, then download the model once and copy it to the other nodes. The scripts use your saved cluster configuration or autodiscovery.

./build-and-copy.sh --exp-b12x -c --copy-parallel
./hf-download.sh deepseek-ai/DeepSeek-V4-Flash-Vision-Exp -c --copy-parallel

Launch the server from the head node:

./launch-cluster.sh -t vllm-node-b12x \
  --apply-mod mods/instanttensor-hybrid-draft-loader \
  -e CUTE_DSL_ARCH=sm_121a \
  -e VLLM_USE_AOT_COMPILE=1 \
  -e VLLM_USE_BREAKABLE_CUDAGRAPH=0 \
  -e VLLM_USE_MEGA_AOT_ARTIFACT=1 \
  -e VLLM_MEMORY_PROFILE_INCLUDE_ATTN=1 \
  -e VLLM_USE_FLASHINFER_SAMPLER=1 \
  -e VLLM_USE_B12X_WO_PROJECTION=1 \
  -e VLLM_USE_B12X_MHC=1 \
  -e VLLM_USE_B12X_FP8_GEMM=1 \
  -e VLLM_USE_B12X_MOE=1 \
  -e VLLM_USE_B12X_SPARSE_INDEXER=1 \
  -e VLLM_USE_V2_MODEL_RUNNER=1 \
  -e VLLM_MOE_SKIP_PADDING=0 \
  -e B12X_MLA_SM120_UNIFIED=1 \
  -e B12X_MOE_FORCE_A8=1 \
  exec vllm serve deepseek-ai/DeepSeek-V4-Flash-Vision-Exp \
    --host 0.0.0.0 \
    --port 8000 \
    --trust-remote-code \
    --tensor-parallel-size 2 \
    --kv-cache-dtype fp8 \
    --block-size 256 \
    --max-model-len auto \
    --max-num-seqs 8 \
    --max-num-batched-tokens 8192 \
    --gpu-memory-utilization 0.85 \
    --enable-prefix-caching \
    --tokenizer-mode deepseek_v4 \
    --tool-call-parser deepseek_v4 \
    --enable-auto-tool-choice \
    --reasoning-parser deepseek_v4 \
    --reasoning-config '{"reasoning_parser":"deepseek_v4","reasoning_start_str":"","reasoning_end_str":""}' \
    --default-chat-template-kwargs.thinking=true \
    --default-chat-template-kwargs.reasoning_effort=high \
    --load-format instanttensor \
    --moe-backend b12x \
    --linear-backend b12x \
    --attention-backend B12X \
    --max-cudagraph-capture-size 48 \
    --compilation-config '{"cudagraph_mode":"FULL_AND_PIECEWISE","custom_ops":["all"]}' \
    --speculative-config '{"method":"dspark","num_speculative_tokens":6,"draft_sample_method":"probabilistic","attention_backend":"B12X"}'

With --tensor-parallel-size 2, the launcher uses two nodes even if more are configured. It supplies the distributed backend, node ranks, and coordination addresses automatically; keep those settings out of the vllm serve command.

Check the server

Wait for vLLM to finish loading and report that the API server is ready. Then, in a second terminal on the serving Spark or cluster head, run:

curl --fail http://localhost:8000/health
curl --fail http://localhost:8000/v1/models

Use http://:8000/v1 as the OpenAI-compatible API base URL in your client, with the model ID from the example you launched. For solo mode, use that Spark's IP address.

Build and launch options

The regular image command (./build-and-copy.sh) pulls eugr/spark-vllm:latest and tags it as vllm-node. Use ./build-and-copy.sh --use-wheels to build the regular runner from precompiled wheels, or ./build-and-copy.sh --rebuild-vllm to compile upstream vLLM.

The B12X image commands (--exp-b12x) pull eugr/spark-vllm-b12x:latest and tag it locally as vllm-node-b12x. To compile vLLM from the experimental fork, add --rebuild-vllm to the corresponding build-and-copy.sh command. --use-wheels is incompatible with --exp-b12x because experimental vLLM wheels are not published. See Building the Docker Image for other build profiles and customizations.

You can add --earlyoom before exec in any launch command to enable the container's low-memory monitor. See Launching the Cluster for additional launcher options.

CHANGELOG

2026-10-02

hf-download.sh now supports model cache management locally and across the cluster: list models and revisions with --list, remove selected models with --delete, and reclaim space from unreferenced revisions with --cleanup. Add -c to extend these operations to all selected cluster nodes. It also enables cluster distribution for downloads and restores, and extends deletion after backup (--backup --delete) to peer nodes. Backups and backup listings use the head node's directory.

Listings include model sizes and node locations, with sorting and console, CSV, JSON, or Markdown output. Use --backup with --backup-dir to save models to a mounted directory on the head node, --list-backup to inspect backups, and --restore to restore and optionally distribute them with -c and --copy-parallel. Backup and deletion accept --revision to select a cached commit by hash, unique prefix, branch, or tag. --backup --delete removes cached copies only after a successful backup and coverage checks. Deletion and cleanup ask for confirmation unless --force is supplied, and missing uvx installations are now handled automatically on the head and peer nodes. Backups automatically use regular snapshot files on external drives without symlink support, with the same listing, restore, and revision selection commands.

2026-09-30

./hf-download.sh will now try to check and automatically repair cache permissions before downloading or distributing the model across the nodes.

2026-09-23

EarlyOOM in 3rd-party containers

--earlyoom flag now works with any Debian/Ubuntu-based vLLM container, such as vllm/vllm-openai or NVIDIA NGC ones. If EarlyOOM is not installed, it will try to install it via apt.

2026-09-10

Qwen3.8 Flash Next solo PLE disk offload

The solo qwen3.8-flash-next-nvfp4-solo recipe now offloads PLE tables to disk with VLLM_PLE_TABLE_MEMORY=disk that allows 1M+ k/v cache allocation (the model max context size is still 262144). Default memory allocation is reduced to 0.8.

2026-09-08

Qwen3.8 Flash Next solo and dual-Spark recipes

Added two recipes for serving local-inference-lab/Qwen3.8-Flash-Next-NVFP4 with the B12X container.

# Single DGX Spark
./run-recipe.sh qwen3.8-flash-next-nvfp4-solo --solo --earlyoom --setup

Dual DGX Spark cluster

./run-recipe.sh qwen3.8-flash-next-nvfp4-cluster --earlyoom --setup

Deepseek V4 Flash Vision Exp support

B12X container now supports deepseek-ai/DeepSeek-V4-Flash-Vision-Exp.

Run with:

./run-recipe.sh deepseek-v4-flash-vision-exp --setup

2026-09-06

Experimental b12x loader

GLM-5.3 Flash recipe is now using experimental b12x loader that is faster and more memory efficient than Instanttensor on DGX Spark. Also reduced KV-cache memory to 8GB to relax memory pressure. Please note that this recipe and b12x builds in general are still experimental, so please update the repository often to keep everything up to date.

2026-09-05

Full GitHub URLs for vLLM PR patches

All --apply-vllm-pr arguments now accept either the existing numeric shorthand for vllm-project/vllm or a full public GitHub pull-request URL such as https://github.com/local-inference-lab/vllm/pull/669. Launch-time patches download the URL's .diff directly; source builds do the same before applying the patch to the selected vLLM ref.

2026-09-04

GLM 5.3 Flash dual-Spark recipe

Added the cluster-only glm-5.3-flash recipe for serving local-inference-lab/GLM-5.3-Flash-NVFP4-Spark on two DGX Spark nodes. It uses the B12X container with MTP and 1M context.

For a first-time cluster setup, discover the nodes and then let the recipe prepare and distribute the B12X image and model:

./run-recipe.sh --discover
./run-recipe.sh glm-5.3-flash --setup

2026-08-27

InstantTensor zero-copy loader mod

Added the opt-in instanttensor-zero-copy mod for memory-constrained model loads. It disables vLLM's per-tensor InstantTensor ownership clone while retaining InstantTensor's required ring buffer. Use with caution.

2026-08-25

Local vLLM source checkouts

build-and-copy.sh --vllm-source-dir now builds vLLM from a clean local Git checkout without requiring the Docker builder to access its remote or host credentials. An optional --vllm-ref is resolved locally; otherwise the build uses the checkout's current HEAD. The host checkout is never modified.

2026-08-21

B12X package in regular builds

Regular local runner builds from vllm-project/vllm now build and install the external B12X package, matching the B12X support being integrated into upstream vLLM. The experimental --exp-b12x profile continues to install the same package for its maintained fork.

2026-08-19

Launch-time vLLM PR application

launch-cluster.sh and run-recipe.sh now accept repeatable --apply-vllm-pr options. The launcher fetches each upstream PR once, validates that it only changes installed vllm/ runtime files, and applies it to every newly created container in command-line layer order. PRs that require native compilation, dependency changes, or packaging changes are rejected with instructions to use the existing build-time option instead.

Example:

./launch-cluster.sh --solo --apply-vllm-pr 52816 exec vllm serve Inferact/Qwen3.8-27B-NVFP4   --host 0.0.0.0   --port 8000   --trust-remote-code   --kv-cache-dtype fp8   --gpu-memory-utilization 0.7   --max-model-len 262144   --max-num-seqs 8   --max-num-batched-tokens 16384   --enable-chunked-prefill   --async-scheduling   --enable-prefix-caching   --speculative-config '{"method":"dflash","model": "z-lab/Qwen3.8-27B-DFlash2", "num_speculative_tokens":8}'   --load-format instanttensor   --reasoning-parser qwen3   --tool-call-parser qwen3_xml   --enable-auto-tool-choice

2026-08-14

B12X source branch update

B12X source builds now use local-inference-lab/vllm@dev/infernal-invocation.

2026-08-12

PyTorch 2.13 and CUTLASS DSL 4.7

Regular and B12X builds now use PyTorch 2.13.0, torchvision 0.28.0, torchaudio 2.11.0, and CUTLASS DSL 4.7.0.

2026-08-11

Nemotron 3.5 Lightning recipe

Added a recipe for NVIDIA's Nemotron 3.5 Lightning with DSpark and 1M context.

Quick start:

git pull
./hf-download.sh nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 # add -c for cluster
./hf-download.sh nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4-DSpark # speculator model; add -c for cluster
./run-recipe.sh recipes/nemotron-3.5-lightning.yaml --solo

2026-08-06

GLM-5.2 NVFP4 8x Spark cluster recipe

Added the cluster-only recipes/8x-spark-cluster/glm-5.2-nvfp4.yaml recipe for serving nvidia/GLM-5.2-NVFP4. Requires 8x nodes. The recipe uses the vllm-node-b12x image.

./run-recipe.sh recipes/8x-spark-cluster/glm-5.2-nvfp4.yaml

2026-08-03

New B12X image

Added --exp-b12x (alias: --experimental-b12x) as an alternative version built from a fork by Luke Alonso. This fork supports a collection of experimental high-performance B12X kernels for sm12x architecture.

Since it is built from a forked vLLM branch, it will be supported in parallel to the main ("regular") build, at least for the time being.

Specifying --exp-b12x without arguments will pull eugr/spark-vllm-b12x:latest from Dockerhub, which is now built and tested together with the main image by CI pipeline on nightly basis (if there are any updates to the source branch or B12X).

Add --rebuild-vllm to compile from the source.

The preset defaults the image tag to vllm-node-b12x; an explicit -t still takes precedence. Additional vLLM changes can be layered onto the preset with one or more --apply-vllm-pr flags. PRs are applied from the upstream vLLM repository, not from Luke's fork! If any PRs are specified, the script will build from the source, not prebuilt images.

vLLM wheels are not published for this build, but it will reuse published Flashinfer wheels if you build from the source (unless --rebuild-flashinfer is specified).

For now, we have only one recipe using this build, with more to come.

DeepSeek V4 Flash 0731 B12X cluster recipe

Added the cluster-only deepseek-v4-flash-0731 recipe for serving deepseek-ai/DeepSeek-V4-Flash-0731 on a dual DGX Spark cluster. The recipe requires B12X container (vllm-node-b12x) that can be pulled by using ./build-and-copy.sh --exp-b12x -c (or just allow the recipe system to pull it for you). See details on b12x build above.

./run-recipe.sh deepseek-v4-flash-0731

2026-07-30

Inkling Small NVFP4 support

Added the support for thinkingmachines/Inkling-Small-NVFP4. It requires at least dual DGX Spark setup.

./run-recipe.sh inkling-small-nvfp4

Add --setup on the first run to prepare the container and download and distribute the model.

Official vLLM image support with earlyoom, InstantTensor, and SciPy

mods/use-official-vllm now installs earlyoom, InstantTensor, and SciPy in addition to the compatibility packages needed by other mods. This makes launch-cluster.sh --earlyoom, --load-format instanttensor, and SciPy-based functionality available when launching official vLLM images such as vllm-openai. The Python package install pins the image's existing Torch packages so the CUDA-enabled build is not replaced during dependency resolution.

Repeatable Docker volume mappings

launch-cluster.sh and run-recipe.sh now accept repeatable -v / --volume mappings using Docker's local_path:container_path syntax. In cluster mode, each mapping is applied to every launched node.

./launch-cluster.sh --solo \
  -v "$PWD/models:/models" \
  exec vllm serve /models/qwen3.6-35b-a4b-nvfp4 ...

2026-07-14

Custom vLLM repositories and PyTorch versions

build-and-copy.sh can now build vLLM from a fork with --vllm-repo. Custom repositories bypass the shared upstream Git checkout cache, force a vLLM source build, and suppress the Dockerfile's upstream preset PRs unless --apply-preset-vllm-prs is explicitly requested.

--torch-version, --torchvision-version, and --torchaudio-version select the packages installed in both the source-build environment and final runner image. The defaults are PyTorch 2.13.0, torchvision 0.28.0, and torchaudio 2.11.0, matching current upstream vLLM CUDA requirements.

Builds from any ref in local-inference-lab/vllm also clone and build the master ref of lukealonso/b12x automatically. The repository produces the b12x distribution. The source layer is refreshed on every applicable runner build so a previously cached clone cannot hide newer upstream commits. Only the locally built B12X wheel is installed. Its CUTLASS DSL metadata is adjusted to the image-wide 4.7.0 pin before installation. The checked-out commit is recorded at /workspace/b12x-source-commit in the image. B12X kernels remain JIT-compiled at runtime; building its Python wheel does not add another CUDA compilation phase to the image build.

2026-07-10

Optional earlyoom monitor

Runner images now include earlyoom, and launch-cluster.sh --earlyoom can run it as the container foreground process instead of sleep infinity. This keeps the existing launcher flow, where Ray and vLLM are started with docker exec, while allowing the container to monitor low host memory continuously. run-recipe.py and run-recipe.sh pass the same --earlyoom and --earlyoom-args options through to the launcher.

The default policy is conservative and absolute-memory based: -M 524288,102400 -s 100 -r 60. That sends SIGTERM when available memory drops below 512 MiB, escalates to SIGKILL below 100 MiB, does not wait for swap to fill before acting, and prints one memory report per minute. You can override the default per launch with --earlyoom-args or set VLLM_SPARK_EARLYOOM_ARGS.

Build and runtime compatibility updates

--gpu-arch now also drives the NCCL NVCC_GENCODE build argument, so non-default local builds compile NCCL for the same target architecture as Torch and FlashInfer.

The source-build KV-cache cleanup is now embedded directly in the Dockerfile, while mods/kv-cache-prealloc-cleanup keeps only the runtime policy tweaks.

mods/gpu-mem-util-gb was refreshed against current vLLM memory profiling code so fixed-GiB GPU memory reservations continue to work with the newer startup path.

2026-07-02

Prebuilt runner image by default

build-and-copy.sh now pulls prebuilt eugr/spark-vllm:latest by default and tags it locally as vllm-node or the tag requested with -t. The latest tag points at the latest tested nightly image. The prebuilt image is updated at the same time as prebuilt wheels by the CI pipeline, so they all stay in sync.

Use --use-wheels to keep the previous wheel-based runner build path. Build customization flags such as --exp-mxfp4, non-default --gpu-arch, --vllm-ref, --flashinfer-ref, rebuild/download flags, and PR application flags also keep the local build path. --tf5 remains a tag-compatibility alias and pulls the prebuilt image as vllm-node-tf5.

Copy now checks the image ID locally and on each remote host before saving the image. Hosts that already have the same image ID are skipped, and docker save is skipped entirely when every target is already current. --no-build still skips image preparation and only copies an already-local tag when needed.

2026-07-01

No-Ray is now the default multi-node backend

launch-cluster.sh and run-recipe.sh now default to no-Ray multi-node launches. Use --ray to opt into Ray; Ray mode ensures vLLM commands include --distributed-executor-backend ray when they omit it. --no-ray remains accepted for compatibility in multi-node launches.

Build and dependency updates

We now use NCCL main branch and include new experimental vLLM Rust frontend in the builds. DeepGEMM now tracks the nv_dev branch.

Transformers 5 flag deprecation

--tf5, --pre-tf, and --pre-transformers are now deprecated compatibility aliases. They no longer override dependency resolution; they only preserve the legacy default image tag. Recipes that previously used vllm-node-tf5 now use the standard vllm-node image.

Recipe updates

Added the gemma4-26b-a4b-nvfp4 recipe, reverted the Qwen3.6-35B-A3B-NVFP4 recipes to --kv-cache-dtype fp8, and cleaned up stale TF5 build args from affected recipes.

2026-06-22

Deepseek V4 Flash support

Support and a recipe for Deepseek V4 Flash has been added based on a newly merged vLLM PR. Please note that this PR requires the DeepGEMM nv_dev branch, which is not present in default vLLM builds, but included in this community build.

To run DSV4F you will need a Spark cluster (2 or more nodes).

To run:

git pull
./build-and-copy.sh -c
./hf-download.sh deepseek-ai/DeepSeek-V4-Flash -c
./run-recipe.sh deepseek-v4-flash --no-ray

DeepGEMM support added

vLLM now uses NVIDIA branch for DeepGEMM that includes support for sm12x GPU family, including DGX Spark. Please note that it is currently not compatible with NVRTC compiler, so DG_JIT_USE_NVRTC has been turned off for new builds.

MiniMax AWQ Weight-Shape Loader Workaround

Added a Dockerfile-level workaround for a vLLM nightly regression where compressed-tensors MoE weight_shape metadata is treated as a scalar during load, causing affected MoE quantized models to fail with shape '[]' is invalid for input of size 2.

2026-06-18

KV Cache Preallocation Cleanup Mod & Updated Qwen3.5-397B recipe for dual Sparks

Added KV cache preallocation cleanup for source-built vLLM wheels, which clears cached CUDA allocator memory before vLLM sizes and allocates KV cache blocks. The dual-node Qwen3.5-397B INT4 AutoRound recipe also applies mods/kv-cache-prealloc-cleanup after mods/gpu-mem-util-gb for its model-specific memory policy tweaks.

The same recipe now keeps the 108 GiB startup reservation, sets VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS=0, and uses a manual 2.25 GiB KV-cache allocation to bypass vLLM's conservative profiler-derived KV budget. The source build handles profiling-only graph-pool cleanup before KV-cache sizing, while the recipe mod makes the env var skip CUDA graph memory profiling entirely and allows the fixed-GiB reservation argument to coexist with --kv-cache-memory-bytes for this model to load. This may result in a bit of swap to be used to offload unused resources, so make sure swap is enabled.

Added recipes for nvidia/Qwen3.6-35B-A3B-NVFP4

Added recipes for NVIDIA's mixed-precision quant for Qwen3.6-35B model. It has a higher precision and an improved performance compared to the earlier NVFP4 quants.

Use:

  • qwen3.6-35b-a3b-nvfp4 to run with MTP on
  • qwen3.6-35b-a3b-nvfp4-no-mtp to run without MTP.

2026-06-10

DiffusionGemma Recipes and Mod

Added day0 support for Google DeepMind's DiffusionGemma model via mods/diffusiongemma. Check out NVIDIA blog for details!

Added four solo-only DiffusionGemma recipes:

  • diffusion-gemma-bf16-thinking for google/diffusiongemma-26B-A4B-it with thinking enabled.
  • diffusion-gemma-bf16 for google/diffusiongemma-26B-A4B-it with thinking disabled.
  • diffusion-gemma-nvfp4-thinking for nvidia/diffusiongemma-26B-A4B-it-NVFP4 with thinking enabled.
  • diffusion-gemma-nvfp4 for nvidia/diffusiongemma-26B-A4B-it-NVFP4 with thinking disabled.
The non-thinking variants still keep --reasoning-parser gemma4, since these models can emit Gemma4 channel markers even when thinking is disabled.

Example:

./hf-download.sh google/diffusiongemma-26B-A4B-it
./run-recipe.sh diffusion-gemma-bf16-thinking --solo

run-recipe.sh Launch Flag Passthrough

run-recipe.sh now passes additional launch-cluster.sh flags through when running recipes: --apply-mod, -p / --publish, and --keep-entrypoint.

Port publishing is still solo-only, matching launch-cluster.sh behavior.

2026-06-09

Recipe Memory Defaults

Raised the default gpu_memory_utilization from 0.7 to 0.8 across the main single-node and two-node recipes to match the current vLLM memory allocation behavior.

2026-06-07

Docker Base Image Compatibility

The default CUDA base image was changed to nvidia/cuda:13.0.2-devel-ubuntu24.04 for broader host compatibility.

The Dockerfile also now passes --allow-change-held-packages when installing the custom NCCL Debian packages, avoiding apt failures when replacing held CUDA/NCCL packages during image builds.

2026-06-06

MiniMax Multi-Node Regression Workaround

Added a targeted Dockerfile patch that disables the MiniMax QK RMSNorm CUDA IPC fused path introduced by vLLM PR #43410. The fused path can fail when tensor parallelism spans DGX Spark nodes; the workaround preserves MiniMax multi-node TP while avoiding a full upstream revert.

2026-06-03

Solo Port Publishing

launch-cluster.sh now supports Docker-style -p / --publish port mappings in solo mode. When port publishing is used, the launcher switches from host networking to Docker bridge networking for that solo container.

Example:

./launch-cluster.sh --solo -p 8000:8000 exec vllm serve ...

2026-05-29

Wheel Freshness Detection

Improved build-and-copy.sh wheel freshness checks so newer locally built wheels are not overwritten just because their filenames differ from the latest release assets. The script now compares local wheel mtimes against remote release asset timestamps before deciding to download. Also switched from using GitHub API to regular HTTP checks to avoid throttling.

gpu-mem-util-gb Patch Refresh

Refreshed mods/gpu-mem-util-gb so it applies against newer vLLM CacheConfig code after upstream line/context drift.

2026-05-28

StepFun Step 3.7 Flash Support

Added support for StepFun Step 3.7 Flash multimodal model.

Requires at least 2 Sparks in a cluster. Both FP8 and NVFP4 checkpoints are supported. FP8 requires more memory, so using NVFP4 is recommended.

Update the repo and build a fresh container first:

git pull
./build-and-copy.sh --cleanup -c

To run NVFP4 version:

Download the model:

./hf-download.sh stepfun-ai/Step-3.7-Flash-NVFP4 -c

Run:

./run-recipe.sh step-3.7-flash-nvfp4 --no-ray

To run FP8 version:

Download the model:

./hf-download.sh stepfun-ai/Step-3.7-Flash-FP8 -c

Run:

./run-recipe.sh step-3.7-flash-fp8 --no-ray

Please note that --no-ray is required for FP8 to fit with full context!

use-official-vllm NCCL Workaround

Updated mods/use-official-vllm to also handle the NCCL load-order bug tracked in vllm-project/vllm#42354. When both the pip-installed nvidia/nccl/lib/libnccl.so.2 and system libnccl2 are present, the mod redirects the pip-installed NCCL path to the system /usr/lib soname, matching the manual workaround that fixes multi-node DGX Spark hangs.

Use it with official vLLM images before starting the model:

./launch-cluster.sh -t vllm/vllm-openai:latest \
  --apply-mod mods/use-official-vllm \
  exec vllm serve ...

Torch Pinning During Wheel Install

Pinned the already-installed CUDA torch build via uv --override when installing locally built wheels and final runtime dependencies in both Dockerfiles. This prevents transitive dependencies from re-resolving torch to a CPU wheel during image builds.

2026-05-22

New Mod: use-official-vllm

Added mods/use-official-vllm, a prerequisite mod for applying patches inside official vLLM Docker containers (e.g. vllm-openai). Official containers do not ship git, which several mods require. This mod installs git via apt-get if it is not already present.

Apply it before any other mod that requires git:

./launch-cluster.sh -t vllm/vllm-openai:latest \
  --apply-mod mods/use-official-vllm \
  --apply-mod mods/gpu-mem-util-gb \
  exec vllm serve ...

gpu-mem-util-gb Updated for Latest vLLM Main

Updated mods/gpu-mem-util-gb patch to apply cleanly against the current vLLM main branch. The mod now checks for git at startup and prints a hint to apply mods/use-official-vllm first if git is missing (relevant when using official vLLM containers).

2026-05-18

NCCL Updated to NVIDIA v2.30u1

The Dockerfile now builds NCCL from NVIDIA's v2.30u1 branch instead of the custom NCCL fork. This branch incorporates all features of the custom fork and is based on the latest NCCL release. The networking guide's NCCL test commands have been updated to use the same branch for 3-node mesh clusters.

2026-05-14

Default Entrypoint Clearing

launch-cluster.sh now clears the Docker image entrypoint by default when starting idle cluster containers. This allows images with server-style entrypoints, such as vllm-openai, to work with the same cluster launcher flow. Use --keep-entrypoint to preserve the image entrypoint.

2026-05-10

Qwen3.5-397B Recipe Memory Updates

Updated Qwen3.5-397B AutoRound recipes to reduce OOM risk. The dual-node recipe now uses standard fractional --gpu-memory-utilization, and the 3-node pipeline-parallel recipe uses InstantTensor loading with lower memory pressure.

2026-05-06

Qwen3.6-35B-A3B-FP8 Recipes

Added qwen3.6-35b-a3b-fp8 and qwen3.6-35b-a3b-fp8-dflash recipes plus a dedicated Qwen3.6 chat-template mod. The DFlash recipe has prefix caching disabled because it caused accuracy issues.

2026-04-29

Gemma4 Recipe Fixes and Experimental b12x Mod

The Gemma4-26B-A4B recipe now uses safetensors loading and no longer applies the obsolete tool parser mod by default.

2026-04-25

MiniMax-M2.7-AWQ Recipe

Added minimax-m2.7-awq, a cluster-only MiniMax M2.7 AWQ recipe using cyankiwi/MiniMax-M2.7-AWQ-4bit.

2026-04-14

Added --load-format instanttensor support to vLLM - thanks @SeraphimSerapis. An experimental option for now, but allows for faster loading than the current fastsafetensors default. You need to rebuild the container to start using the option, but you don't have to trigger the source build.

2026-04-12

Drop-caches mod for Qwen3.5-397B

Updated Qwen3.5-397B recipe (for dual node configuration) to use the new mod mods/drop-caches which clears filesystem caches every minute while the container is running, resolving fastsafetensors getting stuck during loading and a few other bugs when operating close to max memory limit.

2026-04-11

Pinned PyTorch Version

Pinned PyTorch to version 2.11.0 (previously using nightly builds) to fix incompatibility with transformers 5.x and avoid torch version mismatch in builds.

2026-04-02

A new recipe for Gemma4-26B-A4B in "on-the-fly" FP8 quantization:

Single Spark:

./run-recipe.sh gemma4-26b-a4b --solo

Dual Sparks:

./run-recipe.sh gemma4-26b-a4b --no-ray

2026-03-31

Flags to specify Flashinfer ref and apply PRs

build-and-copy.sh gains two new flags that mirror the existing vLLM equivalents:

  • --flashinfer-ref — build FlashInfer from a specific commit SHA, branch, or tag instead of main. Forces a local FlashInfer build (skips prebuilt wheel download).
  • --apply-flashinfer-pr — fetch and apply a FlashInfer GitHub PR patch before building. Can be specified multiple times. Forces a local FlashInfer build.
Both flags are incompatible with --exp-mxfp4.

Default image tag in build-and-copy.sh

build-and-copy.sh now automatically sets a sensible default image tag when -t is not specified:

  • --tf5 / --pre-tf - deprecated compatibility flag; normal build, tag defaults to vllm-node-tf5
  • --exp-mxfp4 - tag defaults to vllm-node-mxfp4
  • in all other cases - tag defaults to vllm-node (no change)
An explicit -t always takes precedence.

Support for 3-node mesh setups

Added initial support for setups where 3 Sparks are connected in a ring-like mesh without an additional switch. See Networking Guide for instructions on how to connect and set up networking in such cluster.

Autodiscover function in both launch-cluster.sh and run-recipe.sh now can detect mesh setups and configure parameters accordingly.

You can try running a model on all 3 nodes in pipeline-parallel configuration using the following recipe:

./run-recipe.sh --discover # force mesh discovery
./run-recipe.sh recipes/3x-spark-cluster/qwen3.5-397b-int4-autoround.yaml --setup --no-ray --force-build # you can drop --setup and --force-build on subsequent calls

Please note that --tensor-parallel-size 3 or -tp 3 is not supported by any commonly used model, so the only two viable options to utilize all three nodes for a single model are:

  • --pipeline-parallel 3 will let you run a model that can't fit on dual Sparks, but without additional speed improvements (total throughtput may improve though).
  • --data-parallel 3 (possibly with --enable-expert-parallel) will let you run a model that can fit on a single Spark, but allow for better concurrency.
You can also run models with --tensor-parallel 2 in a 3-node configuration - in this case only first two nodes (from autodiscovery/.env or from the CLI parameters) will be utilized.

GB10 Verification During Node Discovery

Node discovery now confirms each SSH-reachable peer is a GB10 system before adding it to the cluster: Only hosts reporting NVIDIA GB10 are included. This prevents accidentally adding non-Spark machines that happen to be on the same subnet.

Separate COPY_HOSTS Discovery

Autodiscover now determines the host list used for image and model distribution separately from CLUSTER_NODES:

  • Non-mesh: COPY_HOSTS mirrors CLUSTER_NODES (no change in behaviour).
  • Mesh: scans the direct IB-attached enp1s0f0np0 and enp1s0f1np1 interfaces (not the OOB ETH interface), so large file transfers use the faster direct InfiniBand path.
COPY_HOSTS is saved to .env and respected by build-and-copy.sh, hf-download.sh, and run-recipe.py.

Interactive Configuration Save in autodiscover.sh

autodiscover.sh now handles .env creation with a guided interactive flow, replacing the previous logic in run-recipe.py:

  • Runs automatically when .env is absent.
  • Asks per-node confirmation for both CLUSTER_NODES and COPY_HOSTS.
  • Skips if .env already exists (use --setup to force).
run-recipe.py no longer contains its own .env-save prompt — it delegates entirely to autodiscover.sh.

--setup Flag in launch-cluster.sh and build-and-copy.sh

Both scripts now accept --setup to force a full autodiscovery run and overwrite the existing .env:

./launch-cluster.sh --setup exec vllm serve ...
./build-and-copy.sh --setup -c

This is equivalent to the existing --setup in run-recipe.sh.

--config Flag

hf-download.sh, build-and-copy.sh and launch-cluster.sh now accept --config to load a custom .env configuration file. COPY_HOSTS from the config is used for model distribution:

./hf-download.sh QuantTrio/MiniMax-M2-AWQ --config /path/to/cluster.env -c --copy-parallel

Parallelism-Aware Node Trimming

launch-cluster.sh now parses -tp / --tensor-parallel-size, -pp / --pipeline-parallel-size, and -dp / --data-parallel-size from the exec command or launch script and adjusts the active node count accordingly — for both Ray and no-Ray modes.

  • If fewer nodes are needed than configured, only the required nodes get containers started (excess nodes are left idle).
  • If more nodes are needed than available, an error is raised before anything starts.
Note: Command requires 2 node(s) (tp=2  pp=1  dp=1); using 2 of 3 configured node(s).
Error: Command requires 4 nodes (tp=4  pp=1  dp=1) but only 3 node(s) are configured.

No flags required — the check is automatic whenever parallelism arguments are present in the command.

2026-03-18

--master-port / --head-port Parameter

Added --master-port (synonym: --head-port) to both launch-cluster.sh and run-recipe.sh to configure the port used for cluster coordination:

  • In Ray mode: sets the Ray head node port (previously hardcoded to 6379)
  • In No-Ray mode: sets the PyTorch distributed --master-port passed to vLLM
Default is 29501.
./launch-cluster.sh --master-port 29501 --no-ray exec vllm serve ...
./run-recipe.sh qwen3.5-122b-fp8 --no-ray --master-port 29501

--network Parameter in Build Arguments

Added --network to build-and-copy.sh to allow using host networking during builds. Thanks @apairmont for the PR.

2026-03-17

EXPERIMENTAL Intel/Qwen3.5-397B-A17B-int4-AutoRound Recipe

You can run full 397B Qwen3.5 model on just two Sparks with vision and full context, however you need to make sure your Sparks don't run anything extra that can take a lot of RAM. That means that you don't want to log into the graphical interface or use remote desktop. Connect to the head node via ssh.

Alternatively, you can run in non-graphical mode (runlevel 3) by using sudo systemctl isolate multi-user.target to switch (you can use sudo systemctl set-default graphical.target to switch back to graphical mode), however this is known to reduce performance a bit.

You can run the model with the following command on the head node:

./run-recipe.sh qwen3.5-397b-int4-autoround.yaml --no-ray

Please, note --no-ray is necessary to fit full context. It also improves inference speed by ~1 t/s. By default it will try to allocate 108 GiB for vLLM on each node. You can change this by changing gpu_memory_utilization in the recipe or passing --gpu-mem; this recipe maps that value to --gpu-memory-utilization-gb, so it is GiB rather than a percentage.

KNOWN ISSUES:

  1. The current firmware may cause sudden shutdown event on one or both Sparks during heavy inference. If you have this issue, you will need to lower GPU clock frequency on the affected unit(s), e.g. sudo nvidia-smi -lgc 200,2150. This command will reduce max GPU frequency to 2150 MHz. You can play with higher values to see what works for you (default is 2411 MHz, but can boost to 3000 MHz). Please note that this setting only survives until the next reboot, but can be applied at any time.
  2. You will need to use the new --no-ray argument to fit full context.
  3. If the model gets stuck loading weights, clearing the cache on both nodes can "unstuck" it. Use sudo sh -c 'sync; echo 3 > /proc/sys/vm/drop_caches' to clear the cache.

Major Cluster Orchestration Refactoring

Significantly refactored the internal cluster startup logic in launch-cluster.sh:

  • Removed the standalone run-cluster-node.sh script; its logic is now fully integrated into launch-cluster.sh.
  • Ray head/worker startup, environment variable injection, and launch script distribution are now handled by launch-cluster.sh directly.
  • Worker containers are started with proper per-node environment variables (VLLM_HOST_IP, NCCL_SOCKET_IFNAME, etc.) injected via docker run/docker exec instead of relying on .bashrc.
  • You will now be able to run other vLLM containers without applying use-ngc-vllm mod (current version is just an empty stub).

No-Ray Multi-Node Mode

Added --no-ray flag to launch-cluster.sh to run multi-node vLLM clusters without Ray, using PyTorch's native distributed backend instead. It slightly improves inference performance for most models and reduces memory requirements.

./launch-cluster.sh --no-ray exec vllm serve ...

--no-ray is incompatible with --solo (which already runs without Ray).

run-recipe.sh No-Ray Mode and Extended Flag Passthrough

run-recipe.sh now supports --no-ray flag for running multi-node inference without Ray (uses PyTorch distributed backend instead):

./run-recipe.sh qwen3.5-122b-fp8 --no-ray

The following launch-cluster.sh flags are now also passed through from run-recipe.sh: --master-port, --name, --eth-if, --ib-if, -j, --no-cache-dirs, --non-privileged, --mem-limit-gb, --mem-swap-limit-gb, --pids-limit, --shm-size-gb.

Nemotron-3-Nano-NVFP4 Switched to Marlin Backend

The nemotron-3-nano-nvfp4 recipe has been updated to use the Marlin backend for better performance and reliability (until Flashinfer fully supports NVFP4 on sm121).

2026-03-12

Experimental --gpu-memory-utilization-gb Mod

Added a new mod mods/gpu-mem-util-gb that adds a --gpu-memory-utilization-gb flag to vLLM, allowing you to specify GPU memory reservation in GiB instead of as a fraction. This is particularly useful on DGX Spark's unified memory architecture where available memory changes dynamically.

./launch-cluster.sh --apply-mod mods/gpu-mem-util-gb exec vllm serve ... \
  --gpu-memory-utilization-gb 110

Cannot be used simultaneously with --kv-cache-memory-bytes.

Qwen3.5-397B INT4-AutoRound TP=4 Recipe (4× Spark Cluster)

Added recipes/4x-spark-cluster/qwen3.5-397b-int4-autoround.yaml for running Intel/Qwen3.5-397B-A17B-int4-AutoRound across 4 DGX Spark nodes with tensor parallelism (TP=4).

Benchmarked at ~37 tok/s single-user, ~103 tok/s aggregate (4 concurrent users).

Includes a new mod mods/fix-qwen35-tp4-marlin that resolves a Marlin kernel constraint (MIN_THREAD_N=64) that breaks certain projection layers at TP=4.

Note: Requires NVIDIA driver 580.x. Driver 590.x has a CUDAGraph capture deadlock on GB10 unified memory.

./run-recipe.sh 4x-spark-cluster/qwen3.5-397b-int4-autoround
Thanks @sonusflow for the contribution.

Nemotron-3-Super-120B NVFP4 Recipe

Added a new recipe nemotron-3-super-nvfp4 for running nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 with Marlin kernels. Supports both solo and cluster modes. Includes a custom reasoning parser (super_v3_reasoning_parser.py) fetched from the model reposi

... (README truncated for length)

Chat with me