DeepSeek-v4-Flash-DSpark-2x-DGX-Spark

DeepSeek-v4-Flash 0731 recipe for 2x DGX Sparks

23 articles 1.3k View on GitHub ↗
23 articles
Truncating Tool Calls with `finish_reason: "length"`: Implications for Client-Side Retries

Understand how finish_reason: "length" signals token exhaustion and enables safe client-side retries by preventing malformed JSON tool calls, ensuring robust API interactions.

best-practices
Sep 9, 2026
What Is the Significance of `VLLM_USE_BREAKABLE_CUDAGRAPH=0` in DeepSeek-v4-Flash-DSpark?

Learn the significance of VLLM_USE_BREAKABLE_CUDAGRAPH=0. Discover how disabling breakable CUDA graphs impacts vLLM inference, offering trade-offs in throughput, debuggability, and compatibility for DeepSeek-v4-Flash-DSpark.

deep-dive
Sep 9, 2026
MAX_NUM_BATCHED_TOKENS Default Value in DeepSeek-v4-Flash-DSpark: Configuration Guide

Discover the default value for MAX_NUM_BATCHED_TOKENS in DeepSeek-v4-Flash-DSpark. Learn how this setting impacts token processing and configuration.

configuration-guide
Sep 9, 2026
How to Configure DSPARKMAXINFLIGHTPREFILLS in DeepSeek v4 Flash DSpark

Configure DSPARKMAXINFLIGHTPREFILLS in DeepSeek v4 Flash DSpark by setting the environment variable. Limit concurrent partial prefill requests for better scheduler performance.

how-to-guide
Sep 9, 2026
Partial-Prefill Admission Hotfix Explained: Restoring Decode Lane Balance in DeepSeek-v4

Understand the partial-prefill admission hotfix restoring decode lane balance in DeepSeek-v4. Learn how it prevents token starvation and ensures fair allocation for active decode requests.

internals
Sep 9, 2026
How DSpark Handles Ragged Context Paths in vLLM’s Continuous Batching

Learn how DSpark optimizes vLLM continuous batching by detecting and handling ragged context paths with specialized execution. Discover efficient GPU kernel reuse and eliminated padding.

deep-dive
Sep 9, 2026
Understanding `_req_id_to_slot` and `_free_slots` in DSpark: Slot-Based Request Management

Discover how _req_id_to_slot and _free_slots in DSpark manage execution slots, KV-cache state, and GPU memory for efficient inference. Optimize your DSpark performance today.

internals
Sep 9, 2026
How DSpark Prevents KV Cache Contamination with Request-Stable Slot Mapping

DSpark prevents KV cache contamination using request-stable slot mapping. Learn how its _row_to_slot algorithm ensures persistent slot indexes for stable decoding.

internals
Sep 9, 2026
DSpark Environment Variable Reference: Complete Configuration Guide

Find the complete DSpark environment variable reference in the ENVS.md file within the MiaAI-Lab DeepSeek-v4-Flash-DSpark-2x-DGX-Spark repository. Explore all recognized variables and their configurations.

api-reference
Sep 9, 2026
How prepare-dspark-model-cache.sh Handles Checkpoint Distribution in DeepSeek-V4-Flash

Learn how prepare-dspark-model-cache.sh distributes DeepSeek-V4-Flash checkpoints efficiently. Discover its NFS and SCP copy methods for optimal model deployment.

how-to-guide
Sep 9, 2026
What is docker-compose.dspark.yml? The DeepSeek V4 Flash Deployment Manifest

Understand docker-compose.dspark.yml, the manifest for orchestrating the DeepSeek V4 Flash DSpark inference service. Learn how it manages GPU resources, storage, and multi-node deployments.

how-to-guide
Sep 9, 2026
How to Orchestrate the DeepSeek‑v4‑Flash DSpark Docker Compose Stack: A Complete Guide to `start‑deepseek‑v4‑flash‑dspark.sh`

Learn how the start-deepseek-v4-flash-dspark.sh script orchestrates multi-node Docker Compose deployments. This guide covers configuration, environment validation, and SSH coordination for your DSpark stack.

how-to-guide
Sep 9, 2026

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →