# ds4 | Salvatore Sanfilippo | Knowledge Base | Instagit

DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm

GitHub Stars: 20.4k

Repository: https://github.com/antirez/ds4

---

## Articles

### [The Difference Between Automatic and Explicit Expert Cache Budgets for Metal SSD Streaming](/antirez/ds4/metal-ssd-streaming-automatic-vs-explicit-expert-cache-budgets)

Understand Metal SSD streaming expert cache budgets. Learn the difference between automatic dynamic sizing and explicit user defined budgets in the ds4 engine for optimal performance.

- Tags: deep-dive
- Published: 2026-08-09

### [How to Configure --gpu-devices and --gpu-vram for Multi-GPU Placement in DS4](/antirez/ds4/multi-gpu-placement-understanding-gpu-devices-and-gpu-vram)

Master DS4 multi-GPU placement by configuring --gpu-devices and --gpu-vram. Learn to select CUDA GPUs and set memory budgets for efficient distributed training.

- Tags: how-to-guide
- Published: 2026-08-09

### [DeepSeek V4 Flash Q2 vs Q4 Quantization: Quality and Speed Comparison](/antirez/ds4/deepseek-v4-flash-q2-vs-q4-quantization-quality-vs-speed)

Compare DeepSeek V4 Flash Q2 vs Q4 quantization for speed and quality. Discover which model offers faster inference or superior accuracy for your needs.

- Tags: comparison
- Published: 2026-08-09

### [How Prefill Chunking Works with the Metal Graph Backend in ds4](/antirez/ds4/metal-graph-backend-how-prefill-chunking-works)

Explore prefill chunking in ds4 Metal graph backend. Learn how it processes long prompts in token chunks, minimizing kernel launches and preserving KV-cache consistency.

- Tags: deep-dive
- Published: 2026-08-09

### [GGUF Tensor Layout Requirements for Troubleshooting Model Loading Issues in ds4](/antirez/ds4/troubleshooting-model-loading-gguf-tensor-layout-requirements)

Troubleshoot ds4 model loading issues by understanding GGUF tensor layout requirements. Learn about alignment, dimension ordering, quantization, and data offsets.

- Tags: troubleshooting-guide
- Published: 2026-08-09

### [How DS4 Handles Worker Registration and Rolling Hash Validation in Its Distributed Protocol](/antirez/ds4/distributed-protocol-worker-registration-and-rolling-hash-validation)

Discover how the DS4 distributed protocol manages worker registration via HELLO frames and ensures inference consistency with FNV-1a rolling hashes for token prefix validation before KV-cache updates.

- Tags: internals
- Published: 2026-08-09

### [How to Configure --kv-disk-dir for Session Persistence with KV Disk Caching in DS4](/antirez/ds4/kv-disk-caching-configure-kv-disk-dir-for-session-persistence)

Configure DS4 --kv-disk-dir for session persistence and KV disk caching. Automatically cache KV checkpoints to disk for instant session restoration after restarts.

- Tags: how-to-guide
- Published: 2026-08-09

### [Performance Trade-offs Between MTP Speculative Decoding and DSpark in ds4](/antirez/ds4/mtp-speculative-decoding-vs-dspark-performance-tradeoffs)

Explore MTP speculative decoding vs DSpark in ds4. Discover how MTP saves VRAM with single-token drafting for speed, while DSpark boosts throughput with batched multi-token drafting, requiring more memory.

- Tags: performance
- Published: 2026-08-09

### [How ds4-bench Measures Context Frontier Throughput in LLM Inference](/antirez/ds4/understanding-ds4-bench-context-frontier-throughput-measurement)

Discover how ds4-bench measures context frontier throughput for LLM inference. Learn about its evaluation of token processing speed during prefill and decode phases at varying context lengths.

- Tags: performance
- Published: 2026-08-09

### [How to Run ds4-eval and Interpret GPQA, SuperGPQA, and AIME Results](/antirez/ds4/running-ds4-eval-interpreting-gpqa-supergpqa-and-aime-results)

Learn to run ds4-eval and interpret GPQA, SuperGPQA, and AIME results. Understand PASSED/FAILED reports and fractional scores for your GGUF models.

- Tags: how-to-guide
- Published: 2026-08-09

### [How to Collect and Use imatrix for Better Quantization Quality in ds4](/antirez/ds4/collecting-and-using-imatrix-for-better-quantization-quality)

Enhance ds4 quantization quality by collecting and using imatrix. Learn how to generate an importance matrix and apply it for improved per-column quantization accuracy.

- Tags: how-to-guide
- Published: 2026-08-09

### [How to Build Custom GGUF Quantizations Using deepseek4-quantize](/antirez/ds4/build-custom-gguf-quantizations-with-deepseek4-quantize)

Learn to build custom GGUF quantizations with deepseek4-quantize from the antirez/ds4 repository. Convert safetensors to GGUF models efficiently using Flash quantization.

- Tags: how-to-guide
- Published: 2026-08-09

### [Differences Between GLM 5.2 Support and DeepSeek V4 Models in ds4: A Technical Comparison](/antirez/ds4/glm-5-2-support-differences-from-deepseek-v4-models)

Compare GLM 5.2 support and DeepSeek V4 models in ds4. Discover DeepSeek V4's full speculative decoding, MTP blocks, and 2-bit MoE quantization versus GLM 5.2's restricted inference.

- Tags: deep-dive
- Published: 2026-08-09

### [How Directional Steering with Power Values Below 100 Is Implemented in ds4](/antirez/ds4/directional-steering-power-values-below-100-implementation)

Discover how ds4 implements directional steering below 100. Learn about the dot-product calculation and power factor for subtle, bounded steering effects in this technical deep-dive.

- Tags: internals
- Published: 2026-08-09

### [How to Configure Batch Sessions for Multi-User Serving with ds4-server](/antirez/ds4/ds4-server-batch-session-configuration-for-multi-user-serving)

Configure batched sessions for ds4-server to boost multi-user serving throughput. Group client decode requests into GPU batches for dramatic performance gains.

- Tags: how-to-guide
- Published: 2026-08-09

### [How the Native ds4-agent Manages Sessions and Persists the KV Cache](/antirez/ds4/native-ds4-agent-session-management-and-kv-cache-persistence)

Discover how the native ds4-agent manages sessions and persists KV cache using a worker thread session snapshots and a content-addressable on-disk cache for resilient conversational state.

- Tags: internals
- Published: 2026-08-09

### [How to Use the `--power` Flag for GPU Power Throttling in DwarfStar (ds4) to Reduce Heat and Fan Noise](/antirez/ds4/power-throttling-with-power-flag-for-heat-and-fan-noise-reduction)

Learn how to use the --power flag in DwarfStar ds4 to cap GPU utilization. Reduce heat and fan noise by inserting sleeps between work units without impacting model outputs.

- Tags: how-to-guide
- Published: 2026-08-09

### [Quantization Options and Memory Requirements for DeepSeek V4 PRO vs Flash Models in ds4](/antirez/ds4/deepseek-v4-pro-vs-flash-quantization-options-and-memory-requirements)

Explore quantization options and memory needs for DeepSeek V4 PRO vs Flash models in ds4. Understand RAM usage for low-bit and mixed precision to optimize your deployments.

- Tags: performance
- Published: 2026-08-09

### [How to Set Up and Tune the DSpark Speculative Decoding Confidence Threshold in DS4](/antirez/ds4/dspark-speculative-decoding-setup-and-confidence-threshold-tuning)

Learn to set up and tune the DSpark speculative decoding confidence threshold using the --dspark-confidence flag. Control token acceptance for faster inference.

- Tags: how-to-guide
- Published: 2026-08-09

### [How to Configure CUDA Multi-GPU Tensor Parallelism in ds4](/antirez/ds4/cuda-multi-gpu-tensor-parallelism-configuration)

Configure CUDA multi-GPU tensor parallelism in ds4. Follow steps: compile with CUDA, ensure even GPUs, use DeepSeek-4 with even experts, disable SSD streaming, and launch with the tensor parallel flag.

- Tags: how-to-guide
- Published: 2026-08-09

### [How to Set Up Tensor Parallelism Over RDMA Between Two Macs Using ds4](/antirez/ds4/setup-tensor-parallelism-over-rdma-between-two-macs)

Set up tensor parallelism over RDMA between two Macs using ds4. Split weights for a single inference engine with low latency and TCP fallback.

- Tags: how-to-guide
- Published: 2026-08-09

### [How to Configure Pipeline Parallelism Across Multiple Machines for ds4](/antirez/ds4/configure-pipeline-parallelism-across-multiple-machines)

Configure ds4 pipeline parallelism across machines by deploying worker nodes and a coordinator. Tune network latency with --dist-prefill-chunk and --dist-prefill-window flags for optimal performance.

- Tags: how-to-guide
- Published: 2026-08-09

### [How SSD Streaming Works in DwarfStar: Architecture and Memory Management](/antirez/ds4/how-ssd-streaming-works-in-dwarfstar-and-memory-management)

Learn how DwarfStar implements SSD streaming for efficient LLM expert weight management. Discover its architecture and memory management strategies for optimal performance.

- Tags: architecture
- Published: 2026-08-09

### [Optimizing Inference Speed for Long Context Windows in ds4: Architecture and Tuning Guide](/antirez/ds4/optimize-inference-speed-long-context-windows-ds4)

Boost ds4 inference speed for long contexts with KV cache compression, FlashAttention, and prefill chunking. Optimize prompts and achieve hardware-optimal performance.

- Tags: performance
- Published: 2026-08-08

### [Configuring ds4 Chat Mode: Server, Agent, and CLI Architecture for Tool Execution](/antirez/ds4/configure-tool-calling-function-definitions-ds4-chat-mode)

Configure ds4 chat mode with its CLI server and agent architecture. Learn how to define and execute tools for enhanced assistant capabilities with this unified interface.

- Tags: how-to-guide
- Published: 2026-08-08

### [How to Handle KV Cache Mismatches and Session Recovery in ds4](/antirez/ds4/handle-kv-cache-mismatches-session-recovery-ds4)

Learn how ds4 handles KV cache mismatches and session recovery. Discover how ds4 invalidates sessions, removes corrupt files, and ensures inference integrity with automatic pre-fill.

- Tags: how-to-guide
- Published: 2026-08-08

### [Managing Multi-GPU Memory Placement and Session Allocation in ds4: A Deep Dive into the Layer Packing Algorithm](/antirez/ds4/manage-multi-gpu-memory-placement-session-allocation-ds4)

Optimize multi-GPU memory placement and session allocation with ds4. Learn about the layer packing algorithm for efficient transformer layer assignment and prevent OOM errors.

- Tags: deep-dive
- Published: 2026-08-08

### [Understanding Flash vs PRO Model Differences in ds4 and When to Use Each](/antirez/ds4/ds4-flash-vs-pro-model-differences-usage)

Explore ds4 Flash vs PRO model differences. Learn about context size and memory footprint to choose the right model for your GPU or server needs.

- Tags: deep-dive
- Published: 2026-08-08

### [How to Evaluate Model Capabilities Using ds4-eval: Complete Benchmark Guide](/antirez/ds4/evaluate-model-capabilities-ds4-eval)

Learn how to evaluate model capabilities using ds4-eval. Run the binary against GGUF models for a stable inference pipeline across science math and security benchmarks.

- Tags: how-to-guide
- Published: 2026-08-08

### [How to Run Performance Benchmarks with ds4-bench: A Complete Guide](/antirez/ds4/run-performance-benchmarks-ds4-bench)

Master performance benchmarking with ds4-bench. This guide details measuring DeepSeek V4 and GLM 5.2 inference throughput across context sizes. Optimize your models now.

- Tags: performance
- Published: 2026-08-08

### [Estimating Memory Requirements for Different Context Sizes in ds4: A Complete Guide](/antirez/ds4/estimate-memory-requirements-different-context-sizes-ds4)

Estimate ds4 memory needs for any token context length using ds4_context_memory_estimate API. Prevent errors by calculating GPU/CPU memory consumption before allocation.

- Tags: how-to-guide
- Published: 2026-08-08

### [How to Apply Directional Steering to Modify Model Behavior in ds4](/antirez/ds4/apply-directional-steering-modify-model-behavior-ds4)

Learn to apply directional steering in ds4 to modify transformer model behavior. Bias hidden states using custom vectors by loading a binary file and setting scaling factors.

- Tags: how-to-guide
- Published: 2026-08-08

### [Understanding Layer Slicing with --layers for Partial Model Loading in ds4](/antirez/ds4/understand-layer-slicing-ds4-partial-model-loading)

Master ds4 layer slicing with the --layers flag for efficient partial model loading. Optimize GPU/CPU memory by selectively loading transformer layers and skipping unselected weights.

- Tags: deep-dive
- Published: 2026-08-08

### [Debugging Distributed Inference in DS4: Layer Routing and Worker Communication](/antirez/ds4/debug-distributed-inference-ds4-layer-routing-worker-communication)

Debug distributed inference in DS4 by enabling trace flags to analyze layer routing stalls, HELLO frames, and telemetry. Inspect worker communication and optimize performance.

- Tags: debugging
- Published: 2026-08-08

### [Running Batched Inference with Multiple Concurrent Sessions in DS4: A Deep Dive](/antirez/ds4/run-batched-inference-multiple-concurrent-sessions-ds4)

Discover how to run batched inference with multiple concurrent sessions in DS4 using ds4_sessions_eval_batch. Maximize GPU throughput and reduce overhead.

- Tags: deep-dive
- Published: 2026-08-08

### [How to Persist and Restore KV Cache Sessions to Disk in ds4](/antirez/ds4/persist-restore-kv-cache-sessions-disk-ds4)

Learn to persist and restore KV cache sessions to disk in ds4. Save and load raw binary data to resume generation without recomputing attention history.

- Tags: how-to-guide
- Published: 2026-08-08

### [Comparing GGUF Quantization Options (Q2_K, Q4_K, MXFP4, IQ2_XXS) in ds4 for Quality vs Memory](/antirez/ds4/compare-gguf-quantization-options-ds4-quality-memory)

Explore GGUF quantization in ds4 comparing Q2 K, Q4 K, MXFP4, and IQ2 XXS formats. Discover the best balance of quality and memory for your needs.

- Tags: comparison
- Published: 2026-08-08

### [How to Reduce GPU Power Consumption and Fan Noise Using the --power Flag in ds4](/antirez/ds4/reduce-gpu-power-consumption-fan-noise-ds4-power-flag)

Reduce GPU power consumption and fan noise with the ds4 --power flag. Cap GPU duty cycle for lower energy use and quieter operation, balancing performance and acoustics.

- Tags: how-to-guide
- Published: 2026-08-08

### [Using the Native ds4-agent for Local Coding Tasks: A Complete Guide](/antirez/ds4/using-native-ds4-agent-local-coding)

Leverage the native ds4-agent for local coding with ultra-low latency inference and persistent KV-cache sessions. Restore context seamlessly with this complete guide.

- Tags: how-to-guide
- Published: 2026-08-08

### [How to Set Up the DS4 Server with OpenAI-Compatible API Endpoints](/antirez/ds4/setup-ds4-server-openai-compatible-api)

Set up the DS4 server with OpenAI compatible API endpoints for local inference. Easily replace OpenAI API calls with DS4's built-in HTTP server.

- Tags: how-to-guide
- Published: 2026-08-08

### [MTP (Multi-Token Prediction) Support in ds4: Configuration and Usage Guide](/antirez/ds4/ds4-mtp-multi-token-prediction-support-configuration)

Learn about MTP Multi-Token Prediction in ds4. Discover how this speculative decoding technique speeds up greedy generation by drafting and verifying multiple tokens in parallel. Configure and use MTP effectively.

- Tags: how-to-guide
- Published: 2026-08-08

### [How ds4 Speculative Decoding Works: Implementation Guide and Setup](/antirez/ds4/how-ds4-speculative-decoding-works-enable)

Learn how ds4 speculative decoding accelerates GPU generation with its MTP draft model. This guide explains implementation and setup, ensuring identical results to autoregressive decoding.

- Tags: how-to-guide
- Published: 2026-08-08

### [Configuring CUDA Tensor Parallelism for Multi-GPU NVIDIA Systems in ds4](/antirez/ds4/configuring-cuda-tensor-parallelism-multi-gpu-nvidia)

Learn to configure CUDA tensor parallelism for multi-GPU NVIDIA systems in ds4. Explore synchronizing transformer gates and modifying backend validation for enhanced performance.

- Tags: how-to-guide
- Published: 2026-08-08

### [Setting up Pipeline Parallelism for Multi-Machine Distributed Inference in ds4](/antirez/ds4/setting-up-pipeline-parallelism-for-multi-machine-distributed-inference)

Learn how to set up pipeline parallelism for multi-machine distributed inference in ds4. ds4 splits LLMs across GPUs, keeping them busy during data flow for efficient inference.

- Tags: how-to-guide
- Published: 2026-08-08

### [How to Troubleshoot CUDA Out-of-Memory Errors in Multi-GPU Setups with DwarfStar (ds4)](/antirez/ds4/troubleshoot-cuda-out-of-memory-errors-multi-gpu)

Troubleshoot CUDA out-of-memory errors in multi-GPU setups with DwarfStar. Learn to verify GPU layouts, adjust VRAM margins, set caps, and ensure correct device ordering.

- Tags: how-to-guide
- Published: 2026-08-08

### [How to Integrate DS4 with OpenAI-Compatible API Endpoints](/antirez/ds4/integrate-ds4-openai-compatible-api)

Integrate DS4 with OpenAI-compatible API endpoints easily. DS4 offers a built-in HTTP server at localhost:9333/v1 for seamless replacement of OpenAI's API.

- Tags: how-to-guide
- Published: 2026-08-08

### [Tradeoffs Between IQ2_XXS vs Q2_K Quantization for Quality in DS4](/antirez/ds4/tradeoffs-iq2_xxs-vs-q2_k-quantization-quality)

Explore IQ2_XXS vs Q2_K quantization tradeoffs for DS4 quality. IQ2_XXS offers smaller sizes with reduced fidelity, while Q2_K provides better quality with a slight size increase.

- Tags: performance
- Published: 2026-08-08

### [How MTP Speculative Decoding Works in the DeepSeek V4 Engine (antirez/ds4)](/antirez/ds4/mtp-multi-token-prediction-speculative-decoding-work)

Understand MTP speculative decoding in DeepSeek V4. Accelerate generation by drafting and verifying token sequences efficiently with a lightweight model and target model integration.

- Tags: deep-dive
- Published: 2026-08-08

### [Best Practices for Multi-Session Batching in ds4-server](/antirez/ds4/best-practices-multi-session-batching-ds4-server)

Discover ds4 server best practices for multi-session batching. Maximize throughput and control latency with GPU kernel launch coalescing and pre-allocated server slots.

- Tags: best-practices
- Published: 2026-08-08

### [How to Handle Model Loading Failures with Partial Layer Distribution in DS4](/antirez/ds4/handle-model-loading-failures-partial-layer-distribution)

Learn how DS4 handles model loading failures by falling back layers to CPU and continuing the load. Discover resilient mixed GPU/CPU placement for uninterrupted model initialization.

- Tags: how-to-guide
- Published: 2026-08-08

### [Metal Graph Execution Modes in ds4: Decode vs Batch vs No-Copy](/antirez/ds4/metal-graph-execution-modes-difference)

Understand ds4 Metal graph execution modes: Decode, Batch, and No-Copy. Learn how GPU tiers and hidden-state tensors impact performance in this inference engine.

- Tags: deep-dive
- Published: 2026-08-08

### [How to Benchmark Inference Performance with the ds4-bench Tool](/antirez/ds4/benchmark-inference-performance-ds4-bench)

Benchmark inference performance with ds4-bench. Measure pre-fill latency and greedy decode speed for your models. Optimize your AI applications today.

- Tags: how-to-guide
- Published: 2026-08-08

### [Backend Options for ROCm Strix Halo Systems in DS4](/antirez/ds4/backend-options-rocm-strix-halo-systems)

Discover the unified ROCm backend for Strix Halo systems in DS4. Explore full-graph and SSD-streaming execution modes with runtime flags. Optimize your AMD architecture performance.

- Tags: backend-options
- Published: 2026-08-08

### [How Directional Steering Works in ds4 for Behavior Modification](/antirez/ds4/ds4-directional-steering-behavior-modification)

Learn how directional steering in ds4 modifies transformer behavior by biasing layer activations. Achieve desired outputs without retraining the model.

- Tags: deep-dive
- Published: 2026-08-08

### [Memory Requirements for Different Model Quantizations in DS4](/antirez/ds4/memory-requirements-model-quantizations)

Discover DS4 model quantization memory requirements. Explore 2-bit to 16-bit formats, memory scaling, and estimate total consumption with ds4_context_memory_estimate().

- Tags: performance
- Published: 2026-08-08

### [How to Optimize Power Consumption and Thermal Management with the DS4 `--power` Flag](/antirez/ds4/optimize-power-thermal-management-power-flag)

Optimize power consumption and thermal management with the DS4 --power flag. Cap GPU utilization between 1-99% to reduce power draw and heat. Learn how to implement.

- Tags: how-to-guide
- Published: 2026-08-08

### [DS4 Distributed Worker-Coordinator Communication Protocol: Binary TCP Implementation](/antirez/ds4/distributed-worker-coordinator-communication-protocol)

Discover the DS4 distributed worker-coordinator communication protocol. Learn about its lightweight binary TCP implementation, 12-byte headers, and core message types for efficient distributed computing.

- Tags: architecture
- Published: 2026-08-08

### [How to Configure ds4-agent for Local Development: A Complete Setup Guide](/antirez/ds4/configure-ds4-agent-local-development)

Set up ds4-agent for local development. Compile the binary, configure your .env file, and run the agent to expose a local API for your coding needs.

- Tags: how-to-guide
- Published: 2026-08-08

### [Supported GGUF Quantization Layouts for Routed MoE Experts in DS4](/antirez/ds4/supported-gguf-quantization-layouts-moe-experts)

Discover the seven GGUF quantization layouts for routed MoE experts supported by DS4: Q8_0, IQ2_XXS, Q2_K, Q4_K, Q5_K, Q6_K, and MXFP4. Learn how DS4 ensures compatibility.

- Tags: deep-dive
- Published: 2026-08-07

### [How KV Cache Management Works in ds4 for Long Context: Disk-Based Checkpointing Explained](/antirez/ds4/ds4-kv-cache-management-long-context)

Discover how ds4 manages KV cache for long context using disk-based checkpointing. Resume long conversations efficiently without recomputing. Learn about sparse checkpoints and prefix matching.

- Tags: deep-dive
- Published: 2026-08-07

### [DSpark Speculative Decoding: How to Enable DeepSeek V4 Flash Draft-Verify Acceleration](/antirez/ds4/dspark-speculative-decoding-enable)

Explore DSpark speculative decoding for DeepSeek V4 Flash. Accelerate token generation by using a lightweight draft model for faster verification and increased throughput. Learn how to enable this experimental feature.

- Tags: how-to-guide
- Published: 2026-08-07

### [How to Implement Multi-GPU Inference with CUDA Tensor Parallelism in ds4](/antirez/ds4/multi-gpu-inference-cuda-tensor-parallelism)

Implement multi-GPU inference with CUDA tensor parallelism in ds4. Learn how ds4 splits model tensors across devices for efficient parallel processing without duplicating kernels.

- Tags: how-to-guide
- Published: 2026-08-07

### [Understanding ds4 Distributed Pipeline Parallelism: Architecture and Implementation](/antirez/ds4/ds4-distributed-pipeline-parallelism-architecture)

Explore ds4's distributed pipeline parallelism architecture. Discover how it splits LLM inference across layers using a coordinator-worker model with efficient communication on Apple Silicon.

- Tags: architecture
- Published: 2026-08-07

### [How to Configure SSD Streaming to Run Models Larger Than Available RAM in ds4](/antirez/ds4/configure-ssd-streaming-large-models-ram)

Learn how to configure SSD streaming in ds4 to run models larger than available RAM. ds4 streams model weights from SSD, optimizing performance for Metal CUDA and ROCm.

- Tags: how-to-guide
- Published: 2026-08-07

### [Quantization Formats Supported by ds4 for DeepSeek and GLM Models](/antirez/ds4/ds4-quantization-formats-deepseek-glm)

Discover the quantization formats ds4 supports for DeepSeek and GLM models including Q8_0, Q4_K, Q2_K, IQ2_XXS, and Q8_K. Optimize your model performance.

- Tags: deep-dive
- Published: 2026-08-07

### [How DeepSeek V4 Flash Model Inference Works on Apple Silicon with Metal Backend](/antirez/ds4/how-deepseek-v4-flash-inference-apple-silicon-metal)

Discover how DeepSeek V4 Flash model inference accelerates on Apple Silicon using Metal. Learn about fused FlashAttention kernels, GPU execution, and optimized KV-cache management.

- Tags: internals
- Published: 2026-08-07

### [Optimal Prefill Chunk Sizes for Different Context Lengths in DS4](/antirez/ds4/optimal-prefill-chunk-sizes-context-lengths)

Discover optimal prefill chunk sizes for DS4 context lengths. Learn default settings and customization options for efficient processing with CUDA Tensor-Parallel.

- Tags: performance
- Published: 2026-08-05

### [How ds4 Handles Tool Calling and DSML Format Conversion: A Complete Technical Breakdown](/antirez/ds4/ds4-tool-calling-dsml-format-conversion)

Discover how ds4 manages tool calling and DSML format conversion using a three-stage pipeline for seamless integration with OpenAI-compatible APIs. Learn the technical details.

- Tags: deep-dive
- Published: 2026-08-05

### [DeepSeek V4 Flash vs PRO Model Support in DS4: Key Differences Explained](/antirez/ds4/deepseek-v4-flash-vs-pro-model-support)

Understand DeepSeek V4 Flash vs PRO model support in DS4. Learn when to use Flash (default) and PRO (high-memory/distributed) for optimal performance.

- Tags: deep-dive
- Published: 2026-08-05

### [How to Run GLM 5.2 Models with Supported Routed Quantization Layouts](/antirez/ds4/run-glm-5.2-models-routed-quantization-layouts)

Learn how to run GLM 5.2 models with ds4 using supported routed quantization layouts like Q2 K, Q4 K, Q5 K, and Q6 K. Ensure correct tensor quantization for successful model loading.

- Tags: how-to-guide
- Published: 2026-08-05

### [Best Practices for Multi-GPU Memory Allocation with Placement Hints in ds4](/antirez/ds4/best-practices-multi-gpu-memory-allocation-placement-hints)

Master multi-GPU memory allocation with ds4 best practices. Utilize placement hints for optimal tensor control and ensure fallback strategies.

- Tags: best-practices
- Published: 2026-08-05

### [How to Save and Restore Agent Sessions Using KV Cache Files in ds4 Session persistence](/antirez/ds4/save-restore-agent-sessions-kv-cache-files)

Learn to save and restore agent sessions with ds4 KV cache files. Effortlessly checkpoint and resume execution, avoiding prompt reprocessing for efficient AI development.

- Tags: how-to-guide
- Published: 2026-08-05

### [What Protocol Is Used for Distributed Coordinator-Worker Communication in DS4?](/antirez/ds4/ds4-distributed-coordinator-worker-communication-protocol)

Discover the custom lightweight binary protocol DS4 uses for distributed coordinator-worker communication over TCP sockets. Learn more about this efficient protocol.

- Tags: internals
- Published: 2026-08-05

### [How KV Cache Persistence Works in ds4-server for Conversation State](/antirez/ds4/ds4-server-kv-cache-persistence-conversation-state)

Learn how KV cache persistence in ds4-server saves conversation token-text to durable checkpoint files. Resume chats instantly without recomputing prompts using SHA-1 hashed filenames and smart eviction.

- Tags: internals
- Published: 2026-08-05

### [How to Use ds4-eval for Capability Regression Testing in the ds4 Inference Engine](/antirez/ds4/use-ds4-eval-capability-regression-testing)

Learn how to use ds4-eval for capability regression testing. This benchmark tool helps catch regressions in the ds4 inference engine by running fixed prompt-answer pairs and grading output.

- Tags: how-to-guide
- Published: 2026-08-05

### [Metal vs CUDA vs ROCm Backends in ds4: A Complete Technical Comparison](/antirez/ds4/ds4-metal-vs-cuda-vs-roc-backends)

Compare Metal, CUDA, and ROCm backends in ds4. Discover their API layers, platform support, and exclusive tensor-parallelism features to choose the best GPU acceleration for your needs.

- Tags: deep-dive
- Published: 2026-08-05

### [How to Enable and Tune Power Management to Reduce GPU Heat and Fan Noise in DwarfStar (ds4)](/antirez/ds4/enable-tune-power-management-gpu-heat-fan-noise)

Reduce GPU heat and fan noise in DwarfStar ds4. Use the --power N flag to cap duty-cycle and insert sleep intervals, lowering temperature without changing model outputs.

- Tags: how-to-guide
- Published: 2026-08-05

### [How ds4-server Handles Batched Sessions for Multi-User Inference](/antirez/ds4/ds4-server-batched-sessions-multi-user-inference)

Discover how ds4-server efficiently handles batched sessions for multi-user inference. Learn how it boosts throughput by coalescing requests for GPU/Metal backends.

- Tags: internals
- Published: 2026-08-05

### [Directional Steering in ds4: How to Influence Model Behavior with Hidden-State Biasing](/antirez/ds4/directional-steering-influence-model-behavior)

Learn directional steering in ds4 to bias model hidden states and influence behavior at inference time without retraining. Guide your DeepSeek-V4 model outputs effectively.

- Tags: deep-dive
- Published: 2026-08-05

### [How to Use the Native ds4‑Agent Mode and Its Advantages Over the Server](/antirez/ds4/ds4-agent-mode-vs-server)

Discover the native ds4-agent mode for DeepSeek inference. Bypass HTTP for lower latency, enhanced privacy, and integrated tools. Learn its advantages over the ds4-server.

- Tags: how-to-guide
- Published: 2026-08-05

### [How to Configure CUDA Tensor Parallelism on Multi-GPU Systems Like DGX Spark](/antirez/ds4/configure-cuda-tensor-parallelism-multi-gpu)

Configure CUDA tensor parallelism on multi-GPU systems like DGX Spark. Build DS4 with cuda-spark and launch processes using tensor parallel flags to distribute model layers.

- Tags: how-to-guide
- Published: 2026-08-05

### [DS4 GGUF Quantization Formats: Complete Hardware Compatibility Guide](/antirez/ds4/supported-gguf-quantization-formats-hardware-choice)

Explore DS4 GGUF quantization formats. Find the best option for your hardware from 20+ formats. Get recommendations for NVIDIA, Apple Metal, and CPU for optimal performance.

- Tags: guide
- Published: 2026-08-05

### [How DSpark Speculative Decoding Improves Generation Speed in DS4](/antirez/ds4/dspark-speculative-decoding-speed-improvement)

DSpark speculative decoding speeds up text generation using a draft model to propose tokens, then verifies them with the target model. Learn how it beats standard decoding.

- Tags: deep-dive
- Published: 2026-08-05

### [How to Configure Pipeline Parallelism Across Multiple Machines for Distributed Inference in DwarfStar (ds4)](/antirez/ds4/configure-pipeline-parallelism-distributed-inference)

Configure pipeline parallelism across multiple machines for distributed inference in DwarfStar ds4. Split transformer layers across hosts for models too large for single-machine RAM.

- Tags: how-to-guide
- Published: 2026-08-05

### [Difference Between Pipeline Parallelism and Tensor Parallelism in ds4](/antirez/ds4/ds4-pipeline-vs-tensor-parallelism)

Understand pipeline parallelism vs tensor parallelism in ds4. Discover how each technique optimizes multi-device inference and doubles per-layer throughput for enhanced performance.

- Tags: deep-dive
- Published: 2026-08-05

### [How ds4 SSD Streaming Runs Models Larger Than Available RAM: Architecture and Implementation](/antirez/ds4/how-ds4-ssd-streaming-works-for-large-models)

Discover how ds4 SSD streaming runs models larger than RAM. Learn about its architecture for dynamic expert caching and efficient on-demand weight streaming from SSDs.

- Tags: architecture
- Published: 2026-08-05

### [How to Use the `--chdir` Option to Launch `ds4-agent` from a Different Directory](/antirez/ds4/how-to-use-chdir-option-to-launch-ds4-agent-from-different-directory)

Learn how to use the ds4-agent --chdir option to launch from a different directory. Easily change the working directory before loading model files or initializing the inference engine.

- Tags: how-to-guide
- Published: 2026-08-04

### [How the 80% Backend Working Set Calculation Works for Automatic SSD Streaming Cache Sizing in ds4](/antirez/ds4/how-80-percent-backend-working-set-calculation-works-for-ssd-streaming-cache-sizing)

Understand the 80% backend working set calculation for automatic SSD streaming cache sizing in ds4. Learn how ds4 optimizes cache allocation for performance.

- Tags: internals
- Published: 2026-08-04

### [Rolling Token-Prefix Hash Validation in Distributed Sessions: How ds4 Secures Cluster-Wide Session Integrity](/antirez/ds4/what-is-rolling-token-prefix-hash-validation-in-distributed-sessions)

Learn how rolling token-prefix hash validation in ds4 secures cluster-wide session integrity. Verify token authenticity in O(1) time with this lightweight mechanism.

- Tags: deep-dive
- Published: 2026-08-04

### [Performance Trade-offs Between Prefill and Generation in Distributed Mode](/antirez/ds4/performance-trade-offs-between-prefill-and-generation-in-distributed-mode)

Understand prefill vs generation performance trade-offs in DS4 distributed mode. Optimize KV cache, GPU memory, and latency based on prompt length for better efficiency.

- Tags: performance
- Published: 2026-08-04

### [ds4_dist_session_sync vs ds4_dist_session_eval: Understanding Distributed Session Functions](/antirez/ds4/difference-between-ds4_dist_session_sync-and-ds4_dist_session_eval)

Understand the key differences between ds4_dist_session_sync and ds4_dist_session_eval. Learn how ds4_dist_session_sync modifies state while ds4_dist_session_eval offers read-only evaluation.

- Tags: deep-dive
- Published: 2026-08-04

### [How to Integrate Tool Calling with DS4's OpenAI-Compatible API Endpoints](/antirez/ds4/how-to-integrate-tool-calling-with-openai-compatible-api-endpoints)

Integrate tool calling with DS4's OpenAI-compatible API. Send tool definitions, get function calls, and return results seamlessly. Self-host your AI with DS4.

- Tags: how-to-guide
- Published: 2026-08-04

### [DeepSeek V4 PRO Memory Requirements with SSD Streaming: Complete Guide](/antirez/ds4/memory-requirements-for-running-deepseek-v4-pro-with-ssd-streaming)

Discover DeepSeek V4 PRO memory needs with SSD streaming. Learn about RAM requirements, expert cache allocation, and non-routed weights for optimal performance on your 128GB system.

- Tags: how-to-guide
- Published: 2026-08-04

### [How the MTP Speculative Decoding Path Differs from DSpark in ds4](/antirez/ds4/how-mtp-speculative-decoding-path-differs-from-dspark)

Discover the key differences between MTP speculative decoding and DSpark. Understand their verification methods, tensor parallelism support, and head implementations for ds4.

- Tags: deep-dive
- Published: 2026-08-04

### [Directional Steering in ds4: How to Use It With the --power Flag Below 100](/antirez/ds4/what-is-directional-steering-and-how-to-use-with-power-flag-below-100)

Learn directional steering in ds4 and use the --power flag below 100. Bias language model hidden states safely to cut GPU energy use without impacting steering behavior.

- Tags: how-to-guide
- Published: 2026-08-04

### [Supported GGUF Tensor Layouts in Dwarf Star: Why Arbitrary GGUF Files Are Rejected](/antirez/ds4/supported-gguf-tensor-layouts-and-dwarfstar-limitations)

Explore supported GGUF tensor layouts in Dwarf Star DS4. Understand why arbitrary GGUF files are rejected due to strict validation in ds4.c for optimal performance.

- Tags: internals
- Published: 2026-08-04

### [How ds4-server Handles Batched Sessions for Multi-User Inference](/antirez/ds4/how-batched-session-handling-works-in-ds4-server-for-multi-user-inference)

Discover how ds4-server efficiently manages batched sessions for multi-user inference. Learn about its three-stage pipeline, GPU batching, and mutexes for seamless request processing.

- Tags: internals
- Published: 2026-08-04

### [How to Use the `ds4-eval` Tool for Capability Regression Testing in the ds4 Inference Engine](/antirez/ds4/how-to-use-ds4-eval-tool-for-capability-regression-testing)

Learn to use ds4-eval for capability regression testing with the ds4 inference engine. Detect regressions in execution output quality and performance using this built-in benchmark harness.

- Tags: how-to-guide
- Published: 2026-08-04

### [Coordinator-to-Worker Protocol in DS4 Distributed Inference: Custom Binary Framing Over TCP/RDMA](/antirez/ds4/protocol-for-coordinator-to-worker-communication-in-distributed-inference)

Discover the custom binary protocol DS4 uses for coordinator-to-worker communication. Learn about its 12-byte header, TCP/RDMA support, and efficient distributed inference.

- Tags: internals
- Published: 2026-08-04

### [How DS4 Native Agent Session Management Handles KV Cache Persistence and Resumption](/antirez/ds4/how-native-agent-session-management-handles-kv-cache-persistence-and-resumption)

Discover how DS4 native agent session management persists and resumes KV cache. Learn how DS4 saves and rebuilds the KV cache for seamless generation continuation.

- Tags: internals
- Published: 2026-08-04

### [IQ2_XXS vs Q2_K vs Q4_K Quantization: When to Use Each in DS4](/antirez/ds4/differences-between-iq2_xxs-q2_k-and-q4_k-quantizations-and-usage)

Compare IQ2_XXS, Q2_K, and Q4_K quantizations in DS4. Understand trade-offs in memory, computation, and quality to choose the best format for your needs.

- Tags: deep-dive
- Published: 2026-08-04

### [How to Configure DSpark Speculative Decoding for Faster Token Generation](/antirez/ds4/how-to-configure-dspark-speculative-decoding-for-faster-token-generation)

Boost token generation speed using DSpark speculative decoding in DS4. Learn how to configure DSpark with --mtp and --dspark for faster, efficient processing.

- Tags: how-to-guide
- Published: 2026-08-04

### [Pipeline Parallelism vs Tensor Parallelism in Distributed Inference: Key Differences Explained](/antirez/ds4/difference-between-pipeline-parallelism-and-tensor-parallelism-in-distributed-inference)

Understand pipeline parallelism vs tensor parallelism for distributed inference. Learn how they split models and reduce latency for faster AI.

- Tags: deep-dive
- Published: 2026-08-04

### [How SSD Streaming Works in Dwarf Star (ds4) and Recommended Cache Sizes for MacBook Models](/antirez/ds4/how-does-ssd-streaming-work-in-dwarfstar-and-what-cache-sizes-are-recommended-for-different-macbook-models)

Discover how SSD streaming in Dwarf Star ds4 loads expert MoE models on-demand. Learn recommended cache sizes for MacBooks to optimize performance.

- Tags: performance
- Published: 2026-08-04

