How DeepSeek V4 Flash Model Inference Works on Apple Silicon with Metal Backend
DeepSeek V4 Flash model inference on Apple Silicon leverages the Metal graphics-compute API to execute fused FlashAttention kernels directly on the GPU, eliminating CPU bottlenecks through vectorized thread-group parallelism and device-local KV-cache management.
The DwarfStar (ds4) inference engine powers DeepSeek V4 Flash by running the model entirely on Apple Silicon GPUs via the Metal backend. This implementation parses custom 2-bit routed-MoE quantized weights from GGUF files and dispatches optimized compute kernels that maintain the KV-cache in GPU memory throughout the token generation process, avoiding costly CPU-GPU transfers.
Model Loading and Tensor Layout
In ds4.c, the load_gguf function parses the DeepSeek V4 Flash model file and prepares tensors for GPU execution. The loader recognizes the custom 2-bit routed-MoE quantization scheme used by this model, ensuring weight tensors can stream directly into Metal device buffers without intermediate conversions.
The tensor layout preserves the compressed format, allowing the engine to upload quantized parameters straight to GPU memory. This zero-copy approach minimizes initialization overhead and maximizes available VRAM for the KV-cache on Apple Silicon devices.
KV-Cache Management on Metal Device Buffers
For each generated token, the engine maintains a key/value (K/V) cache entirely within Metal device buffers. The cache is padded to a multiple of 32 rows to align with the vector FlashAttention kernel requirements, enabling full vector register utilization when reading cache entries.
This resident-GPU design prevents the memory bandwidth bottlenecks typical of CPU-managed caches. The padding strategy, handled in metal/flash_attn.metal via the kernel_flash_attn_ext_pad kernel, ensures that subsequent attention operations process data in optimally sized chunks regardless of the actual sequence length.
FlashAttention Kernel Architecture
The core inference logic resides in metal/flash_attn.metal, where three cooperating Metal kernels implement a fused FlashAttention algorithm that combines the Q-K dot-product, softmax, and weighted value summation into a single GPU dispatch:
kernel_flash_attn_ext_pad– Handles KV-cache padding to align memory accesses to 32-row boundarieskernel_flash_attn_ext_blk– Scans attention masks and marks blocks for skippingkernel_flash_attn_ext_impl– Executes the core attention computation with fused operations
These kernels are templated on data types (FP16, quantized FP4/INT4) and head dimensions, allowing the same source code to optimize performance across both 96 GB-class M-series GPUs and smaller Apple Silicon chips.
Thread-Group Configuration and Massive Parallelism
Each kernel receives Metal's standard thread-group coordinates: threadgroup_position_in_grid (tgpig), thread_index_in_threadgroup (tiitg), and SIMD group indices (tiisg). These map directly to the head dimension (Q) and KV-cache rows (C), enabling massive parallelism across thousands of concurrent threads.
This thread-group model ensures that compute units remain saturated during attention operations, distributing work efficiently across the GPU's execution units without CPU coordination.
Compile-Time Specialization with Function Constants
The kernels utilize Metal function constants such as FC_FLASH_ATTN_EXT_PAD (declared via [[function_constant(... )]]) to specialize code paths at compile time. This eliminates conditional branches inside hot loops by baking configuration parameters—such as mask presence, bias terms, and padding requirements—directly into the kernel binary.
Specialization reduces register pressure and instruction cache misses, critical for maintaining high throughput during iterative token generation.
Vector vs Non-Vector Execution Paths
The engine automatically selects between vectorized and generic implementations based on head size characteristics. When the head size (n_head_log2) and cache stride permit 512-wide FP16 vector operations (C = 32), the system launches the optimized vector kernel (kernel_flash_attn_ext_impl with C = 32).
If the head configuration does not meet vector alignment requirements, the engine falls back to a non-vector path that handles arbitrary head sizes correctly, albeit with reduced memory bandwidth efficiency.
Memory Layout and Zero-Copy Buffer Access
Metal kernels operate on raw device const char * buffers rather than typed arrays. Offset calculations (nb11, nb12, etc.) are pre-computed in ds4.c and passed via the ds4_metal_args_* structs, allowing the GPU code to treat buffers as opaque byte arrays.
This approach avoids costly texture conversions and format translations, enabling direct memory access patterns that maximize the effective bandwidth of Apple Silicon's unified memory architecture.
Mask Handling and Block Skipping
The kernel_flash_attn_ext_blk kernel scans the attention mask and writes a single-byte status (res) for each block. The subsequent attention kernel consults these status bytes to skip completely masked blocks entirely, saving significant computation during long-context sequences where large portions of the attention matrix contain negative infinity values.
End-to-End Inference Workflow
Build the Metal binary using the included makefile, which compiles the ds4 engine with Metal backend support:
make
Execute inference on a Mac with DeepSeek V4 Flash, specifying the Metal backend:
./ds4 -m gguf/DeepSeek-V4-Flash-MXFP4Experts-F16HC-F16Compressor-F16Indexer-Q8Attn-Q8Shared-Q8Out-chat-v2-mxfp4-0731.gguf \
--backend metal \
--temp 0.7 \
--max-tokens 256 \
"Explain the physics of superconductivity in simple terms."
The --backend metal flag (parsed in ds4.c via parse_cli) initializes a Metal command queue and uploads quantized weights into GPU buffers. During generation, the engine dispatches the FlashAttention kernels for each layer, updates the KV-cache in-place on the device, and streams resulting tokens back to the CPU for output.
Additional Metal kernels in metal/softmax.metal and metal/dense.metal handle post-attention feed-forward networks, layer normalization, and optional speculative "DSpark" decoding stages entirely on the GPU.
Summary
- DeepSeek V4 Flash runs on Apple Silicon via the ds4 inference engine using the Metal graphics API, keeping all computation on the GPU
- The GGUF loader in
ds4.cparses 2-bit routed-MoE quantized weights and streams them directly to Metal device buffers - KV-cache management pads cache entries to 32-row multiples and maintains them in GPU memory throughout inference
- Three fused kernels (
kernel_flash_attn_ext_pad,kernel_flash_attn_ext_blk,kernel_flash_attn_ext_impl) implement FlashAttention with mask skipping and vectorized execution - Function constants and template specialization eliminate runtime branches for specific head sizes and data types (FP16, FP4, INT4)
- Thread-group parallelism maps attention heads and sequence positions to Metal's SIMD execution model for maximum hardware utilization
Frequently Asked Questions
What Metal kernels are used for FlashAttention in ds4?
The implementation uses three primary kernels defined in metal/flash_attn.metal: kernel_flash_attn_ext_pad for memory alignment, kernel_flash_attn_ext_blk for mask scanning and block skipping, and kernel_flash_attn_ext_impl for the core attention computation. These kernels fuse the Q-K matrix multiplication, softmax normalization, and value weighting into a single GPU dispatch to minimize memory bandwidth usage.
How does the ds4 engine handle the KV-cache on Apple Silicon?
The KV-cache resides entirely in Metal device buffers allocated on the GPU, padded to multiples of 32 rows to accommodate vectorized loads. This design keeps cache updates and attention operations on-device, avoiding the latency penalties of CPU-GPU memory transfers during autoregressive token generation.
What quantization format does DeepSeek V4 Flash use in ds4?
DeepSeek V4 Flash utilizes a custom 2-bit routed-MoE (Mixture of Experts) quantization scheme stored in GGUF format. The load_gguf function in ds4.c recognizes this layout and streams compressed weights directly into GPU memory, where Metal kernels decompress and process them during inference without CPU intervention.
Can ds4 run DeepSeek V4 Flash on smaller Apple Silicon devices?
Yes, the Metal kernels are templated to support various data types and head sizes, allowing the engine to fall back to non-vectorized execution paths when hardware constraints prevent the use of 512-wide FP16 vector operations. This flexibility enables DeepSeek V4 Flash inference across the full range of M-series chips, from base models to Max and Ultra variants.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →