How prepare-dspark-model-cache.sh Handles Checkpoint Distribution in DeepSeek-V4-Flash

The prepare-dspark-model-cache.sh script downloads the DeepSeek-V4-Flash checkpoint once on the head node, verifies all *.safetensors shards, then distributes the model to DSPARK workers via either NFS shared storage or explicit SCP copy based on the DSPARK_WORKER_HF_NFS environment flag.

The prepare-dspark-model-cache.sh script serves as the entry point for model preparation in the MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark repository. This Bash orchestrator manages the complete checkpoint distribution pipeline, ensuring that the 157 GiB model weights are available consistently across your cluster while supporting both gated "abliterated" variants and official releases.

Understanding Checkpoint Resolution Logic

Before downloading begins, the script must determine which model variant to fetch. The resolve_checkpoint() function (lines 66-95 in prepare-dspark-model-cache.sh) evaluates CLI flags (--official, --abliterated) and the ABLITERATED value stored in .env.dspark to select the appropriate Hugging Face repository.

When ABLITERATED=1 is set, the script handles two distinct download phases:

  • Primary checkpoint: The 157 GiB base model remains on the public Hub
  • Gated artifacts: The run_gated_ablit_artifacts() function (lines 31-48) fetches only the 18 KiB direction tensor and terms file from the private gated repository, avoiding redundant downloads of the full weights

Head Node Download and Verification Pipeline

The script strictly enforces a download-once, verify-then-distribute pattern on the head node.

First, run_download() (lines 34-41) executes a Docker container running huggingface_hub.snapshot_download to pull the selected revision into the local HF_CACHE directory. The script supports custom revisions via the DSPARK_REVISION environment variable for pinning specific commit SHAs.

After download completion, verify_cache() (lines 84-89) performs integrity validation by checking for the presence of all required *.safetensors shards. If any shard is missing, the script aborts immediately with a non-zero exit code, preventing corrupted checkpoints from propagating to workers.

Checkpoint Distribution Strategies

The script implements two distinct modes for making the checkpoint available to DSPARK workers, controlled by the DSPARK_WORKER_HF_NFS environment variable checked at lines 58-62.

NFS-Mounted Shared Storage

When DSPARK_WORKER_HF_NFS=1, the script skips remote copying and keeps the checkpoint exclusively on the head node. Workers mount the same HF_CACHE directory via NFS, eliminating storage redundancy across the cluster. This mode requires that all worker nodes have the NFS mount configured at the same path as the head node's cache directory.

Explicit SCP Copy to Workers

By default (DSPARK_WORKER_HF_NFS=0), the script pushes the checkpoint to each worker explicitly:

  1. Directory preparation: Creates the target HF_CACHE path on the worker via SSH
  2. Image verification: Runs verify_worker_image to ensure the worker has the required vllm-dspark-runtime Docker image
  3. Script propagation: Uses scp to transfer prepare-dspark-model-cache.sh and .env.dspark to the worker
  4. Worker execution: Invokes the script remotely with PREPARE_WORKER=0 and --yes flags

Setting PREPARE_WORKER=0 (lines 56-73) signals the script to skip the interactive prompt and download phase, forcing it to reuse the checkpoint already present in the worker's local HF_CACHE after the SCP transfer completes.

Multi-Worker Support and Environment Propagation

For dual-DGX configurations, the script automatically detects WORKER2_HOST and repeats the distribution logic (lines 73-84), ensuring both GPU nodes receive identical checkpoint states.

Authentication tokens propagate securely via WORKER_HF_TOKEN_ENV (lines 19-25). The script forwards the HF_TOKEN to workers through Docker -e arguments, allowing workers to resolve any missing shards without manual re-authentication. This ensures uninterrupted model loading even if the initial SCP transfer encounters network interruptions requiring partial redownloads.

Practical Usage Examples


# Download official checkpoint with interactive confirmation

./prepare-dspark-model-cache.sh

# Prepare gated "abliterated" variant non-interactively

./prepare-dspark-model-cache.sh --abliterated --yes

# Use NFS shared storage (workers mount head node cache)

export DSPARK_WORKER_HF_NFS=1
./prepare-dspark-model-cache.sh

# Distribute to specific revision across two workers

export DSPARK_REVISION=86f746b3
export WORKER2_HOST=worker2.dgx.local
./prepare-dspark-model-cache.sh --official

Summary

  • The prepare-dspark-model-cache.sh script in MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark serves as the single entry point for model preparation across the cluster.
  • Checkpoint resolution via resolve_checkpoint() supports both official and gated abliterated variants without downloading redundant data.
  • Integrity verification through verify_cache() validates all *.safetensors shards before distribution begins.
  • Dual distribution modes: Set DSPARK_WORKER_HF_NFS=1 for NFS-mounted shared storage, or use the default SCP-based push to copy weights explicitly to each worker.
  • Worker idempotency: The PREPARE_WORKER=0 flag ensures workers skip downloads and reuse pre-staged checkpoints, preventing redundant 157 GiB transfers.

Frequently Asked Questions

What happens if a worker already has a partial checkpoint?

When PREPARE_WORKER=0 is set during the remote execution phase, the worker runs verify_cache() to check for existing *.safetensors shards. If shards are missing or corrupted, the Hugging Face token forwarded via WORKER_HF_TOKEN_ENV allows the worker to download only the specific missing files rather than the full 157 GiB dataset.

How does the script handle the gated "abliterated" model differently from the official release?

The run_gated_ablit_artifacts() function fetches only the small direction tensor (18 KiB) and terms file from the private repository when ABLITERATED=1, while pulling the base checkpoint from the public Hub. This design prevents unauthorized redistribution of gated content while keeping the large weights in the open domain.

Can I use this script with a custom Hugging Face cache directory?

Yes. The script respects the HF_CACHE environment variable defined in .env.dspark for both the head node download and the worker target path. Ensure this directory is either NFS-mounted across nodes (when using DSPARK_WORKER_HF_NFS=1) or has sufficient disk space for the explicit copy mode (default).

What is the purpose of the DSPARK_REVISION variable?

This environment variable pins the checkpoint to a specific Git commit SHA or tag, ensuring reproducible deployments. When set, run_download() passes this value to huggingface_hub.snapshot_download as the revision parameter, preventing automatic updates when new model versions are published to the Hub.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →