Metal vs CUDA vs ROCm Backends in ds4: A Complete Technical Comparison

Metal, CUDA, and ROCm are GPU acceleration backends in ds4 that share the same high-level inference engine but differ in API layers, platform support, and feature availability—with Metal offering exclusive tensor-parallelism capabilities.

The ds4 project (by antirez) implements a modular GPU inference system with three distinct backends. Each backend is selected at runtime via the ds4_backend enum and compiled with backend-specific flags. This article examines the architectural differences, code paths, and feature gaps between these implementations.

Backend Overview

ds4 defines three backend constants in its core header:

/* From ds4.h - backend type enumeration */
typedef enum {
    DS4_BACKEND_CPU = 0,
    DS4_BACKEND_METAL,
    DS4_BACKEND_CUDA
} ds4_backend;

Notice that ROCm does not have a dedicated enum value—it piggybacks on DS4_BACKEND_CUDA with compile-time differentiation.

Metal Backend: Native Apple GPU Implementation

The Metal backend is the most mature implementation in ds4, featuring full Objective-C integration and exclusive access to tensor-parallel operations.

Platform and API

Metal targets macOS and iOS devices with Apple GPUs. The glue layer resides in ds4_metal.m, which creates Metal objects and manages the GPU pipeline:

/* From ds4_metal.m - device and queue initialization */
id<MTLDevice> device = MTLCreateSystemDefaultDevice();
id<MTLCommandQueue> commandQueue = [device newCommandQueue];

The kernel code is written in Metal Shading Language (.metal files) and compiled at build time.

Exclusive Feature: Tensor-Parallelism

Tensor-parallelism is Metal-only in ds4. The implementation handles gate encoding and expert sharding through specialized Metal kernels. When other backends attempt these operations, they hit a hard limit:

/* From ds4_rocm.cu - CUDA/ROCm stub for unsupported tensor-parallelism */
if (tensor_parallel) {
    fprintf(stderr, "tensor parallelism is Metal-only\n");
    return NULL;
}

This limitation stems from architectural choices in the ds4 codebase—the tensor-parallel implementation was built first for Metal and never ported to CUDA/ROCm paths.

CUDA Backend: NVIDIA GPU Support

The CUDA backend provides NVIDIA GPU acceleration through the CUDA runtime and cuBLAS library.

Compilation and Source Structure

CUDA code is compiled with nvcc and activated by the DS4_CUDA_BUILD flag. The implementation lives in ds4_rocm.cu but follows the CUDA path when __HIP_PLATFORM_AMD__ is not defined:

/* From ds4_rocm.cu - CUDA-specific header includes */
#ifndef __HIP_PLATFORM_AMD__
#include <cuda_runtime.h>
#include <cublas_v2.h>
#define DS4_BACKEND_NAME "CUDA"

Feature Limitations

The CUDA backend lacks tensor-parallelism support despite having comparable raw compute performance. The same stub function that warns about Metal-only features applies here.

Backend Selection

Users select CUDA via the command-line string "cuda":

/* From ds4_cli.c - backend string parsing */
if (!strcmp(s, "cuda")) return DS4_BACKEND_CUDA;

ROCm Backend: AMD GPU via HIP

The ROCm backend delivers AMD GPU support through HIP—AMD's CUDA-compatible runtime layer.

Shared Source Architecture

ROCm reuses the identical source file as CUDA. Compilation with hipcc and the __HIP_PLATFORM_AMD__ macro toggles the AMD code path:

/* From ds4_rocm.cu - ROCm/HIP header includes */
#ifdef __HIP_PLATFORM_AMD__
#include <hip/hip_runtime.h>
#include <hipblaslt/hipblaslt.h>
#define DS4_BACKEND_NAME "ROCm"

The DS4_ROCM_BUILD compile flag distinguishes the final binary.

Backend Enum Alias

Despite the different name, "rocm" maps to DS4_BACKEND_CUDA at runtime:

/* From ds4_cli.c - ROCm is an alias for CUDA enum value */
if (!strcmp(s, "rocm")) return DS4_BACKEND_CUDA;  // Same enum, different build

Runtime behavior diverges based on which compiler built the binary—the DS4_ROCM_BUILD flag switches to HIP function calls internally.

Tensor-Parallelism Status

Like CUDA, ROCm inherits the same tensor-parallelism stubs and cannot execute gate encoding or expert sharding operations. The shared source code means both backends advance (or stall) together.

Feature Comparison Summary

Capability Metal CUDA ROCm
Target hardware Apple GPUs NVIDIA GPUs AMD GPUs
Operating systems macOS, iOS Linux, macOS, Windows Linux, macOS
Primary language Objective-C + Metal Shading Language CUDA C++ HIP C++ (CUDA-compatible)
Tensor-parallel gate encoding ✅ Full implementation ❌ Stub only ❌ Stub only
Tensor-parallel expert sharding ✅ Full implementation ❌ Stub only ❌ Stub only
Backend enum value DS4_BACKEND_METAL DS4_BACKEND_CUDA DS4_BACKEND_CUDA (alias)
Compile flag DS4_METAL_BUILD DS4_CUDA_BUILD DS4_ROCM_BUILD
Build compiler Xcode toolchain nvcc hipcc

Runtime GPU Detection

The ds4 engine uses a unified helper to identify GPU backends in ds4.c:

/* From ds4.c - GPU backend detection */
int backend_is_gpu(ds4_backend backend) {
    return backend == DS4_BACKEND_METAL ||
           backend == DS4_BACKEND_CUDA;
}

This treats Metal and CUDA/ROCm uniformly for memory allocation decisions, display formatting, and other GPU-aware operations.

Key Source Files

  • ds4.h — Backend enum definitions and public API
  • ds4_metal.m — Objective-C glue layer for Metal device management
  • ds4_rocm.cu — Shared source for CUDA and ROCm with conditional compilation
  • ds4.c — Core engine with backend-agnostic logic and backend_is_gpu() helper
  • ds4_cli.c / ds4_server.c — Backend string parsing ("metal", "cuda", "rocm")

Backend Selection Example

/* Complete backend initialization flow */
ds4_backend select_and_init_backend(const char *name) {
    ds4_backend backend;
    
    if (!strcmp(name, "metal")) {
        backend = DS4_BACKEND_METAL;
        ds4_metal_init();        // Loads .metal kernels, creates MTLDevice
    } else if (!strcmp(name, "cuda")) {
        backend = DS4_BACKEND_CUDA;
        ds4_cuda_init();         // cudaSetDevice, cuBLAS handle creation
    } else if (!strcmp(name, "rocm")) {
        backend = DS4_BACKEND_CUDA;  // Same enum!
        ds4_rocm_init();         // HIP device setup, hipBLAS initialization
    } else {
        backend = DS4_BACKEND_CPU;
    }
    
    return backend;
}

Summary

  • Metal provides the most complete ds4 implementation with native Apple GPU optimization and exclusive tensor-parallelism support for distributed inference patterns.

  • CUDA offers mature NVIDIA acceleration through cuBLAS but lacks tensor-parallel features due to unimplemented stubs in the shared source architecture.

  • ROCm delivers AMD GPU compatibility via HIP, reusing CUDA code paths through conditional compilation, yet inherits the same feature limitations as CUDA.

  • All three backends share high-level inference logic; differences emerge at the API binding layer and in Metal-specific kernel optimizations.

Frequently Asked Questions

Why does ROCm use the same enum value as CUDA?

ds4 treats ROCm as a compile-time variant rather than a distinct runtime type. The DS4_BACKEND_CUDA enum covers both NVIDIA and AMD paths, with DS4_ROCM_BUILD and __HIP_PLATFORM_AMD__ macros selecting the correct API calls at build time. This reduces branching in the core engine.

Can tensor-parallelism be ported to CUDA/ROCm?

The stub functions in ds4_rocm.cu (lines 35-41) indicate unimplemented territory, not fundamental barriers. The Metal implementation could theoretically be translated to CUDA warp primitives and ROCm wavefront operations, but this requires significant engineering effort to match the gate encoding and expert sharding logic.

How does backend selection affect model loading?

All backends load identical weight formats. The difference lies in where computation happens: Metal kernels execute on Apple Neural Engine or GPU cores, CUDA kernels on NVIDIA tensor cores, and ROCm kernels on AMD compute units. Memory layout and quantization schemes remain consistent across targets.

Is Metal performance better than CUDA for the same model?

Cross-platform performance comparisons depend on specific GPUs (Apple M-series vs. NVIDIA RTX/A100 vs. AMD Instinct). Metal's tensor-parallelism advantage matters most for expert-parallel MoE models where gate distribution across devices reduces latency. For standard dense models, raw TFLOPS and memory bandwidth dominate.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →