ds4
DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm
Understand Metal SSD streaming expert cache budgets. Learn the difference between automatic dynamic sizing and explicit user defined budgets in the ds4 engine for optimal performance.
How to Configure --gpu-devices and --gpu-vram for Multi-GPU Placement in DS4Master DS4 multi-GPU placement by configuring --gpu-devices and --gpu-vram. Learn to select CUDA GPUs and set memory budgets for efficient distributed training.
DeepSeek V4 Flash Q2 vs Q4 Quantization: Quality and Speed ComparisonCompare DeepSeek V4 Flash Q2 vs Q4 quantization for speed and quality. Discover which model offers faster inference or superior accuracy for your needs.
How Prefill Chunking Works with the Metal Graph Backend in ds4Explore prefill chunking in ds4 Metal graph backend. Learn how it processes long prompts in token chunks, minimizing kernel launches and preserving KV-cache consistency.
GGUF Tensor Layout Requirements for Troubleshooting Model Loading Issues in ds4Troubleshoot ds4 model loading issues by understanding GGUF tensor layout requirements. Learn about alignment, dimension ordering, quantization, and data offsets.
How DS4 Handles Worker Registration and Rolling Hash Validation in Its Distributed ProtocolDiscover how the DS4 distributed protocol manages worker registration via HELLO frames and ensures inference consistency with FNV-1a rolling hashes for token prefix validation before KV-cache updates.
How to Configure --kv-disk-dir for Session Persistence with KV Disk Caching in DS4Configure DS4 --kv-disk-dir for session persistence and KV disk caching. Automatically cache KV checkpoints to disk for instant session restoration after restarts.
Performance Trade-offs Between MTP Speculative Decoding and DSpark in ds4Explore MTP speculative decoding vs DSpark in ds4. Discover how MTP saves VRAM with single-token drafting for speed, while DSpark boosts throughput with batched multi-token drafting, requiring more memory.
How ds4-bench Measures Context Frontier Throughput in LLM InferenceDiscover how ds4-bench measures context frontier throughput for LLM inference. Learn about its evaluation of token processing speed during prefill and decode phases at varying context lengths.
How to Run ds4-eval and Interpret GPQA, SuperGPQA, and AIME ResultsLearn to run ds4-eval and interpret GPQA, SuperGPQA, and AIME results. Understand PASSED/FAILED reports and fractional scores for your GGUF models.
How to Collect and Use imatrix for Better Quantization Quality in ds4Enhance ds4 quantization quality by collecting and using imatrix. Learn how to generate an importance matrix and apply it for improved per-column quantization accuracy.
How to Build Custom GGUF Quantizations Using deepseek4-quantizeLearn to build custom GGUF quantizations with deepseek4-quantize from the antirez/ds4 repository. Convert safetensors to GGUF models efficiently using Flash quantization.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →