Backend Options for ROCm Strix Halo Systems in DS4
The DS4 repository provides a single, unified ROCm backend specifically optimized for AMD Strix Halo architecture, offering both full-graph and SSD-streaming execution modes through runtime flags.
The antirez/ds4 project ships a dedicated inference engine for large language models that includes specialized support for AMD's Strix Halo platform. When targeting these APUs—such as those found in the Framework Desktop—the codebase exposes specific ROCm Strix Halo backend options that leverage the ROCm graph API and ROCWMMA libraries for accelerated compute.
The Primary ROCm Backend Architecture
The Strix Halo backend is not a separate fork but a conditional compilation target within the main DS4 codebase. It integrates ROCm-specific kernels and memory management strategies tailored to the unified memory architecture of Strix Halo APUs.
Build Configuration and Targets
To compile DS4 with ROCm support for Strix Halo, use the dedicated Make target defined in the Makefile. The strix-halo target (aliased as rocm) triggers the compilation of HIP kernels and links against the ROCm runtime libraries.
# Build the Strix Halo (ROCm) binary
make strix-halo -j$(nproc) # alias: make rocm
This build process compiles ds4_rocm.cu and the kernel headers located in the rocm/ directory, producing a single executable that automatically initializes the ROCm backend when running on compatible hardware.
Core Implementation Files
The backend implementation spans several key source files that handle graph construction and kernel execution:
ds4_rocm.cu– Contains the main ROCm implementation, including graph construction logic and the launch mechanisms for GPU kernels.ds4_rocm.h– Public header exposing ROCm-specific APIs and data structures to the rest of the inference engine.rocm/*.cuh– Collection of device kernel headers (e.g.,ds4_rocm_attention.cuh,ds4_rocm_moe.cuh) implementing compute primitives for attention and mixture-of-experts layers.STRIXHALO.md– Documentation detailing ROCm setup requirements, necessary kernel parameters, and package dependencies for Strix Halo machines.
Runtime Execution Modes
The ROCm backend supports two distinct execution strategies selected at runtime through command-line flags. Both modes utilize the same compiled binary (ds4) but alter the memory management and computation graph behavior.
Full-Graph ROCm Mode
By default, the backend operates in full-graph mode, where the entire model computation graph is compiled and executed through the ROCm graph API. This mode maximizes throughput for models that fit entirely within the available GPU memory.
You can enable debugging features for this mode through environment variables:
# Dump the compiled ROCm graph for inspection
DS4_ROCM_GRAPH_DUMP_PREFIX=./graph_dump ./ds4 -m model.gguf
SSD-Streaming Mode for Routed Experts
For large routed-expert models that exceed GPU memory capacity, the backend supports SSD-streaming. When invoked with the --ssd-streaming flag, DS4 dynamically loads expert weights from storage during inference, enabling execution of models like GLM-5.2 on Strix Halo's unified memory architecture.
# Enable SSD-streaming for large routed-expert models
./ds4 -m gguf/GLM-5.2-UD-Q2_K_RoutedQ2K.gguf \
--rocm \
--ssd-streaming \
--ctx 4096
Environment Tuning Variables
The ROCm backend exposes several environment variables for fine-tuning performance characteristics on Strix Halo hardware:
DS4_ROCM_GRAPH_DUMP_PREFIX– Specifies a directory path where the compiled ROCm graph will be dumped for debugging and optimization analysis.DS4_ROCM_DISABLE_STREAMING_READAHEAD– Disables read-ahead optimizations during SSD streaming, useful for systems with slower storage or when debugging I/O bottlenecks.DS4_ROCM_ENABLE_STREAMING_PREFILL– Forces the streaming prefill path even when the full model fits in GPU memory, allowing performance comparison between modes.
These variables allow developers to adapt the ROCm Strix Halo backend options to specific memory layouts and storage configurations without recompiling the binary.
Building and Running the ROCm Backend
Follow these steps to build and execute DS4 with ROCm support on a Strix Halo system:
- Compile the binary using the Strix Halo target:
make strix-halo -j$(nproc)
- Run inference using the default full-graph mode:
./ds4 -m gguf/DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix.gguf
- Enable streaming for routed-expert architectures:
./ds4 -m gguf/GLM-5.2-UD-Q2_K_RoutedQ2K.gguf --rocm --ssd-streaming --ctx 4096
- Debug the execution graph by setting the dump prefix:
DS4_ROCM_GRAPH_DUMP_PREFIX=./graph_dump ./ds4 -m gguf/model.gguf
Summary
- The DS4 repository provides a unified ROCm backend specifically targeting AMD Strix Halo APUs through the
make strix-halobuild target. - Runtime execution offers two modes: full-graph (default) for in-memory models and SSD-streaming (via
--ssd-streaming) for large routed-expert architectures. - Configuration occurs through environment variables like
DS4_ROCM_GRAPH_DUMP_PREFIXwithout requiring separate binaries. - Core implementation resides in
ds4_rocm.cu,ds4_rocm.h, and therocm/kernel directory, with setup instructions inSTRIXHALO.md.
Frequently Asked Questions
What hardware targets the ROCm Strix Halo backend?
The backend specifically targets AMD Strix Halo APUs, including systems like the Framework Desktop that feature integrated RDNA graphics and unified memory architecture. The ds4 binary detects compatible hardware and automatically initializes the ROCm graph API when built with the strix-halo target.
How do I enable SSD streaming for large models?
Pass the --ssd-streaming flag when launching DS4. This mode is essential for routed-expert models (such as GLM-5.2 variants) that exceed available GPU memory, allowing the system to stream expert weights from SSD storage during inference.
Can I debug the ROCm execution graph?
Yes. Set the DS4_ROCM_GRAPH_DUMP_PREFIX environment variable to a directory path before running DS4. This dumps the compiled ROCm graph structures to disk, enabling inspection of the execution flow and kernel scheduling as implemented in ds4_rocm.cu.
Is the ROCm backend separate from the main DS4 binary?
No. The ROCm backend compiles into the same ds4 executable as other backends. The build system uses the strix-halo (or rocm) Make target to conditionally compile HIP code from ds4_rocm.cu and the rocm/ kernel headers, producing a single binary that selects the appropriate backend at runtime based on hardware detection and command-line flags.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →