ds4

DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm

104 articles 20.4k View on GitHub ↗
104 articles
The Difference Between Automatic and Explicit Expert Cache Budgets for Metal SSD Streaming

Understand Metal SSD streaming expert cache budgets. Learn the difference between automatic dynamic sizing and explicit user defined budgets in the ds4 engine for optimal performance.

deep-dive
Aug 9, 2026
How to Configure --gpu-devices and --gpu-vram for Multi-GPU Placement in DS4

Master DS4 multi-GPU placement by configuring --gpu-devices and --gpu-vram. Learn to select CUDA GPUs and set memory budgets for efficient distributed training.

how-to-guide
Aug 9, 2026
DeepSeek V4 Flash Q2 vs Q4 Quantization: Quality and Speed Comparison

Compare DeepSeek V4 Flash Q2 vs Q4 quantization for speed and quality. Discover which model offers faster inference or superior accuracy for your needs.

comparison
Aug 9, 2026
How Prefill Chunking Works with the Metal Graph Backend in ds4

Explore prefill chunking in ds4 Metal graph backend. Learn how it processes long prompts in token chunks, minimizing kernel launches and preserving KV-cache consistency.

deep-dive
Aug 9, 2026
GGUF Tensor Layout Requirements for Troubleshooting Model Loading Issues in ds4

Troubleshoot ds4 model loading issues by understanding GGUF tensor layout requirements. Learn about alignment, dimension ordering, quantization, and data offsets.

troubleshooting-guide
Aug 9, 2026
How DS4 Handles Worker Registration and Rolling Hash Validation in Its Distributed Protocol

Discover how the DS4 distributed protocol manages worker registration via HELLO frames and ensures inference consistency with FNV-1a rolling hashes for token prefix validation before KV-cache updates.

internals
Aug 9, 2026
How to Configure --kv-disk-dir for Session Persistence with KV Disk Caching in DS4

Configure DS4 --kv-disk-dir for session persistence and KV disk caching. Automatically cache KV checkpoints to disk for instant session restoration after restarts.

how-to-guide
Aug 9, 2026
Performance Trade-offs Between MTP Speculative Decoding and DSpark in ds4

Explore MTP speculative decoding vs DSpark in ds4. Discover how MTP saves VRAM with single-token drafting for speed, while DSpark boosts throughput with batched multi-token drafting, requiring more memory.

performance
Aug 9, 2026
How ds4-bench Measures Context Frontier Throughput in LLM Inference

Discover how ds4-bench measures context frontier throughput for LLM inference. Learn about its evaluation of token processing speed during prefill and decode phases at varying context lengths.

performance
Aug 9, 2026
How to Run ds4-eval and Interpret GPQA, SuperGPQA, and AIME Results

Learn to run ds4-eval and interpret GPQA, SuperGPQA, and AIME results. Understand PASSED/FAILED reports and fractional scores for your GGUF models.

how-to-guide
Aug 9, 2026
How to Collect and Use imatrix for Better Quantization Quality in ds4

Enhance ds4 quantization quality by collecting and using imatrix. Learn how to generate an importance matrix and apply it for improved per-column quantization accuracy.

how-to-guide
Aug 9, 2026
How to Build Custom GGUF Quantizations Using deepseek4-quantize

Learn to build custom GGUF quantizations with deepseek4-quantize from the antirez/ds4 repository. Convert safetensors to GGUF models efficiently using Flash quantization.

how-to-guide
Aug 9, 2026
…

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →