# How Worker Nodes Expose the HuggingFace Cache via NFSv4 in DeepSeek v4 Flash

> Learn how worker nodes expose the HuggingFace cache via NFSv4 in DeepSeek v4 Flash. Access model checkpoints with read-only Docker volume mounts without local storage.

- Repository: [Mia's AI Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark)
- Tags: how-to-guide
- Published: 2026-09-09

---

**The worker node exposes the HuggingFace cache via NFSv4 by mounting a Docker volume backed by a head-node NFS export, allowing read-only access to model checkpoints without local storage.**

The MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark repository orchestrates distributed inference across DGX systems. To avoid duplicating multi-gigabyte model weights on every worker, the launch scripts expose the head node's HuggingFace cache via NFSv4, creating a centralized, read-only storage layer accessible to all containerized workers.

## NFS Export Architecture Overview

The implementation relies on [`files/nfs-share.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/files/nfs-share.sh) to coordinate server-side exports and worker-side mounts, while [`docker-compose.dspark-nfs.override.yml`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/docker-compose.dspark-nfs.override.yml) handles the container-level volume binding. The head node acts as the NFSv4 server, exporting its local HuggingFace cache directory (`$HF_CACHE_DIR`), while workers consume this export through a Docker-managed NFS volume named `dspark-hf`.

## Step 1: Detect the NFS Server IP

The function `nfs_detect_server_ip` in [`files/nfs-share.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/files/nfs-share.sh) automatically identifies the network interface for the NFS export. It reads the IPv4 address of the interface defined by the `IFACE` environment variable—typically the **ConnectX-7 NIC**—and stores the result in `NFS_SERVER_IP`.

```bash

# From files/nfs-share.sh

nfs_detect_server_ip() {
    NFS_SERVER_IP=$(ip -4 addr show "${IFACE}" | grep -oP '(?<=inet\s)\d+(\.\d+){3}')
    export NFS_SERVER_IP
}

```

This IP is subsequently used for all NFS client mount operations on the worker nodes.

## Step 2: Launch the NFSv4 Server on the Head Node

The `nfs_ensure_server` function checks for an existing NFSv4 service before starting a new exporter. It queries the detected IP using `rpcinfo -t "$NFS_SERVER_IP" nfs 4` to verify if a server is already listening. If not, it builds the `dspark-nfs:local` image from `files/nfs-server/` and runs a privileged container that exports the host's HuggingFace cache as `/export`.

The export enforces **NFSv4.2** with read-only permissions:

```bash

# nfs_ensure_server excerpt from files/nfs-share.sh

if ! rpcinfo -t "$NFS_SERVER_IP" nfs 4 >/dev/null 2>&1; then
    docker build -t dspark-nfs:local files/nfs-server/
    docker run -d --privileged \
        -v "${HF_CACHE_DIR}:/export:ro" \
        -p "${NFS_SERVER_IP}:2049:2049" \
        dspark-nfs:local \
        /export ${NFS_SERVER_IP}(ro,sync,no_subtree_check,fsid=0,no_root_squash,nohide)
fi

```

The container binds to port 2049 on the ConnectX-7 interface, ensuring high-bandwidth transport between the head node and workers.

## Step 3: Provision the Worker NFS Volume

On the worker side, `nfs_ensure_worker_volume` executes via SSH to create a Docker volume using the **local NFS driver**. This volume points to the server IP discovered in Step 1 and uses explicit NFSv4.2 mount options for performance optimization.

```bash

# Executed on worker via SSH

ssh $WORKER_HOST "
  docker volume create --driver local \
    --opt type=nfs \
    --opt o=addr=${NFS_SERVER_IP},nfsvers=4.2,ro,nconnect=8,rsize=1048576,wsize=1048576,hard,timeo=600 \
    --opt device=:/ \
    dspark-hf
"

```

The `nconnect=8` parameter enables multiple TCP connections per NFS session, while `rsize` and `wsize` are tuned to 1MB for efficient large-model checkpoint transfers.

## Step 4: Prepare Runtime Cache Subdirectories

While the HuggingFace model cache is read-only, workers require writable directories for JIT compilation artifacts. The `nfs_jit_subdirs` function defines these subdirectories: `triton-cache`, `tilelang-cache`, `vllm-cache`, `flashinfer`, `b12x-cute-cache`, and `nccl-fr`.

The `nfs_ensure_worker_jit_dirs` function creates these directories on the worker's local host filesystem (not via NFS), ensuring that runtime compilers can write temporary files without traversing the network:

```bash

# From files/nfs-share.sh

nfs_jit_subdirs() {
    echo "triton-cache tilelang-cache vllm-cache flashinfer b12x-cute-cache nccl-fr"
}

nfs_ensure_worker_jit_dirs() {
    local worker=$1
    for dir in $(nfs_jit_subdirs); do
        ssh "$worker" "mkdir -p /var/cache/dspark/${dir}"
    done
}

```

## Step 5: Mount Inside Worker Containers

The [`docker-compose.dspark-nfs.override.yml`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/docker-compose.dspark-nfs.override.yml) file overrides the standard compose configuration to bind the NFS volume into every DSpark container at `/cache/huggingface` with read-only semantics:

```yaml
services:
  dspark:
    volumes:
      - dspark-hf:/cache/huggingface:ro
    environment:
      - HF_HOME=/cache/huggingface

volumes:
  dspark-hf:
    external: true

```

When the launcher sets `HF_HOME=/cache/huggingface`, frameworks like VLLM, FlashInfer, and TileLang automatically resolve model paths against the NFS-mounted cache. Workers stream model weights directly from the head node without maintaining local copies.

## Launching with NFS Support

To enable the NFS workflow, invoke the launcher with the `--nfs` flag. This automatically sources [`files/nfs-share.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/files/nfs-share.sh) and executes the server detection, volume creation, and directory provisioning sequences:

```bash
./start-deepseek-v4-flash-dspark.sh --nfs

```

The script internally calls `nfs_ensure_server`, `nfs_ensure_worker_volume`, and `nfs_ensure_worker_jit_dirs` before bringing up the container stack.

## Verifying the Mount

Confirm that workers can access the shared cache by inspecting the volume contents inside a temporary container:

```bash
ssh $WORKER_HOST "docker run --rm -v 'dspark-hf:/hf:ro' alpine:latest ls /hf/hub/models--deepseek-ai--DeepSeek-V4-Flash-0731"

```

Successful execution indicates the NFSv4 export is functioning correctly and the worker node can resolve model artifacts from the head node's HuggingFace cache.

## Summary

- **Head-node detection**: The `nfs_detect_server_ip` function identifies the ConnectX-7 interface IP to use as the NFS server address.
- **Server initialization**: `nfs_ensure_server` launches a privileged Docker container exporting `$HF_CACHE_DIR` via **NFSv4.2** with read-only permissions.
- **Worker volume creation**: `nfs_ensure_worker_volume` provisions a Docker NFS volume `dspark-hf` pointing to the server IP with optimized mount parameters.
- **JIT directory preparation**: Writable subdirectories for compilation caches are created locally on workers via `nfs_ensure_worker_jit_dirs`.
- **Container binding**: [`docker-compose.dspark-nfs.override.yml`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/docker-compose.dspark-nfs.override.yml) mounts the NFS volume at `/cache/huggingface` inside all worker containers, setting `HF_HOME` for automatic framework integration.

## Frequently Asked Questions

### What NFS version does DeepSeek v4 Flash require?

The implementation specifically requires **NFSv4.2**, enforced through the `nfsvers=4.2` mount option in both the server export configuration and the worker volume creation. This version provides better performance characteristics and locking mechanisms compared to NFSv3 when serving large model checkpoints to multiple concurrent workers.

### Why is the HuggingFace cache mounted read-only?

The cache is mounted read-only (`ro`) to ensure **consistency across the cluster** and prevent accidental corruption of model weights. Since the head node manages the cache directory (downloading and updating models), workers function as pure consumers. This design eliminates cache coherency issues and reduces the risk of split-brain scenarios in distributed deployments.

### Which network interface is used for the NFS export?

The export uses the network interface defined by the `IFACE` environment variable, which defaults to the **ConnectX-7 NIC** in DGX environments. The `nfs_detect_server_ip` function in [`files/nfs-share.sh`](https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark/blob/main/files/nfs-share.sh) extracts the IPv4 address from this interface to bind the NFS server, ensuring high-bandwidth, low-latency access between the head node and worker nodes.

### Can workers write to any cache directories?

While the main HuggingFace model cache is strictly read-only, workers write **JIT compilation artifacts** (Triton kernels, TileLang caches, FlashInfer temporary files) to local host directories created by `nfs_ensure_worker_jit_dirs`. These directories reside on the worker's local filesystem at paths like `/var/cache/dspark/triton-cache`, not through the NFS mount, allowing compile-time write operations without violating the read-only constraint on model weights.