DeepSeek-v4-Flash-DSpark-2x-DGX-Spark
DeepSeek-v4-Flash 0731 recipe for 2x DGX Sparks
Understand how finish_reason: "length" signals token exhaustion and enables safe client-side retries by preventing malformed JSON tool calls, ensuring robust API interactions.
What Is the Significance of `VLLM_USE_BREAKABLE_CUDAGRAPH=0` in DeepSeek-v4-Flash-DSpark?Learn the significance of VLLM_USE_BREAKABLE_CUDAGRAPH=0. Discover how disabling breakable CUDA graphs impacts vLLM inference, offering trade-offs in throughput, debuggability, and compatibility for DeepSeek-v4-Flash-DSpark.
MAX_NUM_BATCHED_TOKENS Default Value in DeepSeek-v4-Flash-DSpark: Configuration GuideDiscover the default value for MAX_NUM_BATCHED_TOKENS in DeepSeek-v4-Flash-DSpark. Learn how this setting impacts token processing and configuration.
How to Configure DSPARKMAXINFLIGHTPREFILLS in DeepSeek v4 Flash DSparkConfigure DSPARKMAXINFLIGHTPREFILLS in DeepSeek v4 Flash DSpark by setting the environment variable. Limit concurrent partial prefill requests for better scheduler performance.
Partial-Prefill Admission Hotfix Explained: Restoring Decode Lane Balance in DeepSeek-v4Understand the partial-prefill admission hotfix restoring decode lane balance in DeepSeek-v4. Learn how it prevents token starvation and ensures fair allocation for active decode requests.
How DSpark Handles Ragged Context Paths in vLLM’s Continuous BatchingLearn how DSpark optimizes vLLM continuous batching by detecting and handling ragged context paths with specialized execution. Discover efficient GPU kernel reuse and eliminated padding.
Understanding `_req_id_to_slot` and `_free_slots` in DSpark: Slot-Based Request ManagementDiscover how _req_id_to_slot and _free_slots in DSpark manage execution slots, KV-cache state, and GPU memory for efficient inference. Optimize your DSpark performance today.
How DSpark Prevents KV Cache Contamination with Request-Stable Slot MappingDSpark prevents KV cache contamination using request-stable slot mapping. Learn how its _row_to_slot algorithm ensures persistent slot indexes for stable decoding.
DSpark Environment Variable Reference: Complete Configuration GuideFind the complete DSpark environment variable reference in the ENVS.md file within the MiaAI-Lab DeepSeek-v4-Flash-DSpark-2x-DGX-Spark repository. Explore all recognized variables and their configurations.
How prepare-dspark-model-cache.sh Handles Checkpoint Distribution in DeepSeek-V4-FlashLearn how prepare-dspark-model-cache.sh distributes DeepSeek-V4-Flash checkpoints efficiently. Discover its NFS and SCP copy methods for optimal model deployment.
What is docker-compose.dspark.yml? The DeepSeek V4 Flash Deployment ManifestUnderstand docker-compose.dspark.yml, the manifest for orchestrating the DeepSeek V4 Flash DSpark inference service. Learn how it manages GPU resources, storage, and multi-node deployments.
How to Orchestrate the DeepSeek‑v4‑Flash DSpark Docker Compose Stack: A Complete Guide to `start‑deepseek‑v4‑flash‑dspark.sh`Learn how the start-deepseek-v4-flash-dspark.sh script orchestrates multi-node Docker Compose deployments. This guide covers configuration, environment validation, and SSH coordination for your DSpark stack.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →