How Worker Nodes Expose the HuggingFace Cache via NFSv4 in DeepSeek v4 Flash

The worker node exposes the HuggingFace cache via NFSv4 by mounting a Docker volume backed by a head-node NFS export, allowing read-only access to model checkpoints without local storage.

The MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark repository orchestrates distributed inference across DGX systems. To avoid duplicating multi-gigabyte model weights on every worker, the launch scripts expose the head node's HuggingFace cache via NFSv4, creating a centralized, read-only storage layer accessible to all containerized workers.

NFS Export Architecture Overview

The implementation relies on files/nfs-share.sh to coordinate server-side exports and worker-side mounts, while docker-compose.dspark-nfs.override.yml handles the container-level volume binding. The head node acts as the NFSv4 server, exporting its local HuggingFace cache directory ($HF_CACHE_DIR), while workers consume this export through a Docker-managed NFS volume named dspark-hf.

Step 1: Detect the NFS Server IP

The function nfs_detect_server_ip in files/nfs-share.sh automatically identifies the network interface for the NFS export. It reads the IPv4 address of the interface defined by the IFACE environment variable—typically the ConnectX-7 NIC—and stores the result in NFS_SERVER_IP.


# From files/nfs-share.sh

nfs_detect_server_ip() {
    NFS_SERVER_IP=$(ip -4 addr show "${IFACE}" | grep -oP '(?<=inet\s)\d+(\.\d+){3}')
    export NFS_SERVER_IP
}

This IP is subsequently used for all NFS client mount operations on the worker nodes.

Step 2: Launch the NFSv4 Server on the Head Node

The nfs_ensure_server function checks for an existing NFSv4 service before starting a new exporter. It queries the detected IP using rpcinfo -t "$NFS_SERVER_IP" nfs 4 to verify if a server is already listening. If not, it builds the dspark-nfs:local image from files/nfs-server/ and runs a privileged container that exports the host's HuggingFace cache as /export.

The export enforces NFSv4.2 with read-only permissions:


# nfs_ensure_server excerpt from files/nfs-share.sh

if ! rpcinfo -t "$NFS_SERVER_IP" nfs 4 >/dev/null 2>&1; then
    docker build -t dspark-nfs:local files/nfs-server/
    docker run -d --privileged \
        -v "${HF_CACHE_DIR}:/export:ro" \
        -p "${NFS_SERVER_IP}:2049:2049" \
        dspark-nfs:local \
        /export ${NFS_SERVER_IP}(ro,sync,no_subtree_check,fsid=0,no_root_squash,nohide)
fi

The container binds to port 2049 on the ConnectX-7 interface, ensuring high-bandwidth transport between the head node and workers.

Step 3: Provision the Worker NFS Volume

On the worker side, nfs_ensure_worker_volume executes via SSH to create a Docker volume using the local NFS driver. This volume points to the server IP discovered in Step 1 and uses explicit NFSv4.2 mount options for performance optimization.


# Executed on worker via SSH

ssh $WORKER_HOST "
  docker volume create --driver local \
    --opt type=nfs \
    --opt o=addr=${NFS_SERVER_IP},nfsvers=4.2,ro,nconnect=8,rsize=1048576,wsize=1048576,hard,timeo=600 \
    --opt device=:/ \
    dspark-hf
"

The nconnect=8 parameter enables multiple TCP connections per NFS session, while rsize and wsize are tuned to 1MB for efficient large-model checkpoint transfers.

Step 4: Prepare Runtime Cache Subdirectories

While the HuggingFace model cache is read-only, workers require writable directories for JIT compilation artifacts. The nfs_jit_subdirs function defines these subdirectories: triton-cache, tilelang-cache, vllm-cache, flashinfer, b12x-cute-cache, and nccl-fr.

The nfs_ensure_worker_jit_dirs function creates these directories on the worker's local host filesystem (not via NFS), ensuring that runtime compilers can write temporary files without traversing the network:


# From files/nfs-share.sh

nfs_jit_subdirs() {
    echo "triton-cache tilelang-cache vllm-cache flashinfer b12x-cute-cache nccl-fr"
}

nfs_ensure_worker_jit_dirs() {
    local worker=$1
    for dir in $(nfs_jit_subdirs); do
        ssh "$worker" "mkdir -p /var/cache/dspark/${dir}"
    done
}

Step 5: Mount Inside Worker Containers

The docker-compose.dspark-nfs.override.yml file overrides the standard compose configuration to bind the NFS volume into every DSpark container at /cache/huggingface with read-only semantics:

services:
  dspark:
    volumes:
      - dspark-hf:/cache/huggingface:ro
    environment:
      - HF_HOME=/cache/huggingface

volumes:
  dspark-hf:
    external: true

When the launcher sets HF_HOME=/cache/huggingface, frameworks like VLLM, FlashInfer, and TileLang automatically resolve model paths against the NFS-mounted cache. Workers stream model weights directly from the head node without maintaining local copies.

Launching with NFS Support

To enable the NFS workflow, invoke the launcher with the --nfs flag. This automatically sources files/nfs-share.sh and executes the server detection, volume creation, and directory provisioning sequences:

./start-deepseek-v4-flash-dspark.sh --nfs

The script internally calls nfs_ensure_server, nfs_ensure_worker_volume, and nfs_ensure_worker_jit_dirs before bringing up the container stack.

Verifying the Mount

Confirm that workers can access the shared cache by inspecting the volume contents inside a temporary container:

ssh $WORKER_HOST "docker run --rm -v 'dspark-hf:/hf:ro' alpine:latest ls /hf/hub/models--deepseek-ai--DeepSeek-V4-Flash-0731"

Successful execution indicates the NFSv4 export is functioning correctly and the worker node can resolve model artifacts from the head node's HuggingFace cache.

Summary

  • Head-node detection: The nfs_detect_server_ip function identifies the ConnectX-7 interface IP to use as the NFS server address.
  • Server initialization: nfs_ensure_server launches a privileged Docker container exporting $HF_CACHE_DIR via NFSv4.2 with read-only permissions.
  • Worker volume creation: nfs_ensure_worker_volume provisions a Docker NFS volume dspark-hf pointing to the server IP with optimized mount parameters.
  • JIT directory preparation: Writable subdirectories for compilation caches are created locally on workers via nfs_ensure_worker_jit_dirs.
  • Container binding: docker-compose.dspark-nfs.override.yml mounts the NFS volume at /cache/huggingface inside all worker containers, setting HF_HOME for automatic framework integration.

Frequently Asked Questions

What NFS version does DeepSeek v4 Flash require?

The implementation specifically requires NFSv4.2, enforced through the nfsvers=4.2 mount option in both the server export configuration and the worker volume creation. This version provides better performance characteristics and locking mechanisms compared to NFSv3 when serving large model checkpoints to multiple concurrent workers.

Why is the HuggingFace cache mounted read-only?

The cache is mounted read-only (ro) to ensure consistency across the cluster and prevent accidental corruption of model weights. Since the head node manages the cache directory (downloading and updating models), workers function as pure consumers. This design eliminates cache coherency issues and reduces the risk of split-brain scenarios in distributed deployments.

Which network interface is used for the NFS export?

The export uses the network interface defined by the IFACE environment variable, which defaults to the ConnectX-7 NIC in DGX environments. The nfs_detect_server_ip function in files/nfs-share.sh extracts the IPv4 address from this interface to bind the NFS server, ensuring high-bandwidth, low-latency access between the head node and worker nodes.

Can workers write to any cache directories?

While the main HuggingFace model cache is strictly read-only, workers write JIT compilation artifacts (Triton kernels, TileLang caches, FlashInfer temporary files) to local host directories created by nfs_ensure_worker_jit_dirs. These directories reside on the worker's local filesystem at paths like /var/cache/dspark/triton-cache, not through the NFS mount, allowing compile-time write operations without violating the read-only constraint on model weights.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →