# Metal vs CUDA vs ROCm Backends in ds4: A Complete Technical Comparison

> Compare Metal, CUDA, and ROCm backends in ds4. Discover their API layers, platform support, and exclusive tensor-parallelism features to choose the best GPU acceleration for your needs.

- Repository: [Salvatore Sanfilippo/ds4](https://github.com/antirez/ds4)
- Tags: deep-dive
- Published: 2026-08-05

---

**Metal, CUDA, and ROCm are GPU acceleration backends in ds4 that share the same high-level inference engine but differ in API layers, platform support, and feature availability—with Metal offering exclusive tensor-parallelism capabilities.**

The **ds4** project (by antirez) implements a modular GPU inference system with three distinct backends. Each backend is selected at runtime via the `ds4_backend` enum and compiled with backend-specific flags. This article examines the architectural differences, code paths, and feature gaps between these implementations.

## Backend Overview

ds4 defines three backend constants in its core header:

```c
/* From ds4.h - backend type enumeration */
typedef enum {
    DS4_BACKEND_CPU = 0,
    DS4_BACKEND_METAL,
    DS4_BACKEND_CUDA
} ds4_backend;

```

Notice that **ROCm does not have a dedicated enum value**—it piggybacks on `DS4_BACKEND_CUDA` with compile-time differentiation.

## Metal Backend: Native Apple GPU Implementation

The **Metal backend** is the most mature implementation in ds4, featuring full Objective-C integration and exclusive access to tensor-parallel operations.

### Platform and API

Metal targets **macOS and iOS devices with Apple GPUs**. The glue layer resides in `ds4_metal.m`, which creates Metal objects and manages the GPU pipeline:

```c
/* From ds4_metal.m - device and queue initialization */
id<MTLDevice> device = MTLCreateSystemDefaultDevice();
id<MTLCommandQueue> commandQueue = [device newCommandQueue];

```

The kernel code is written in **Metal Shading Language** (`.metal` files) and compiled at build time.

### Exclusive Feature: Tensor-Parallelism

**Tensor-parallelism is Metal-only** in ds4. The implementation handles gate encoding and expert sharding through specialized Metal kernels. When other backends attempt these operations, they hit a hard limit:

```c
/* From ds4_rocm.cu - CUDA/ROCm stub for unsupported tensor-parallelism */
if (tensor_parallel) {
    fprintf(stderr, "tensor parallelism is Metal-only\n");
    return NULL;
}

```

This limitation stems from architectural choices in the ds4 codebase—the tensor-parallel implementation was built first for Metal and never ported to CUDA/ROCm paths.

## CUDA Backend: NVIDIA GPU Support

The **CUDA backend** provides NVIDIA GPU acceleration through the CUDA runtime and cuBLAS library.

### Compilation and Source Structure

CUDA code is compiled with `nvcc` and activated by the `DS4_CUDA_BUILD` flag. The implementation lives in `ds4_rocm.cu` but follows the CUDA path when `__HIP_PLATFORM_AMD__` is **not** defined:

```c
/* From ds4_rocm.cu - CUDA-specific header includes */
#ifndef __HIP_PLATFORM_AMD__
#include <cuda_runtime.h>
#include <cublas_v2.h>
#define DS4_BACKEND_NAME "CUDA"

```

### Feature Limitations

The CUDA backend **lacks tensor-parallelism support** despite having comparable raw compute performance. The same stub function that warns about Metal-only features applies here.

### Backend Selection

Users select CUDA via the command-line string "cuda":

```c
/* From ds4_cli.c - backend string parsing */
if (!strcmp(s, "cuda")) return DS4_BACKEND_CUDA;

```

## ROCm Backend: AMD GPU via HIP

The **ROCm backend** delivers AMD GPU support through **HIP**—AMD's CUDA-compatible runtime layer.

### Shared Source Architecture

ROCm reuses the identical source file as CUDA. Compilation with `hipcc` and the `__HIP_PLATFORM_AMD__` macro toggles the AMD code path:

```c
/* From ds4_rocm.cu - ROCm/HIP header includes */
#ifdef __HIP_PLATFORM_AMD__
#include <hip/hip_runtime.h>
#include <hipblaslt/hipblaslt.h>
#define DS4_BACKEND_NAME "ROCm"

```

The `DS4_ROCM_BUILD` compile flag distinguishes the final binary.

### Backend Enum Alias

Despite the different name, **"rocm" maps to `DS4_BACKEND_CUDA`** at runtime:

```c
/* From ds4_cli.c - ROCm is an alias for CUDA enum value */
if (!strcmp(s, "rocm")) return DS4_BACKEND_CUDA;  // Same enum, different build

```

Runtime behavior diverges based on which compiler built the binary—the `DS4_ROCM_BUILD` flag switches to HIP function calls internally.

### Tensor-Parallelism Status

Like CUDA, ROCm **inherits the same tensor-parallelism stubs** and cannot execute gate encoding or expert sharding operations. The shared source code means both backends advance (or stall) together.

## Feature Comparison Summary

| Capability | Metal | CUDA | ROCm |
| ---------- | ----- | ---- | ---- |
| Target hardware | Apple GPUs | NVIDIA GPUs | AMD GPUs |
| Operating systems | macOS, iOS | Linux, macOS, Windows | Linux, macOS |
| Primary language | Objective-C + Metal Shading Language | CUDA C++ | HIP C++ (CUDA-compatible) |
| Tensor-parallel gate encoding | ✅ Full implementation | ❌ Stub only | ❌ Stub only |
| Tensor-parallel expert sharding | ✅ Full implementation | ❌ Stub only | ❌ Stub only |
| Backend enum value | `DS4_BACKEND_METAL` | `DS4_BACKEND_CUDA` | `DS4_BACKEND_CUDA` (alias) |
| Compile flag | `DS4_METAL_BUILD` | `DS4_CUDA_BUILD` | `DS4_ROCM_BUILD` |
| Build compiler | Xcode toolchain | `nvcc` | `hipcc` |

## Runtime GPU Detection

The ds4 engine uses a unified helper to identify GPU backends in [`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c):

```c
/* From ds4.c - GPU backend detection */
int backend_is_gpu(ds4_backend backend) {
    return backend == DS4_BACKEND_METAL ||
           backend == DS4_BACKEND_CUDA;
}

```

This treats Metal and CUDA/ROCm uniformly for memory allocation decisions, display formatting, and other GPU-aware operations.

## Key Source Files

- **[`ds4.h`](https://github.com/antirez/ds4/blob/main/ds4.h)** — Backend enum definitions and public API
- **`ds4_metal.m`** — Objective-C glue layer for Metal device management
- **`ds4_rocm.cu`** — Shared source for CUDA and ROCm with conditional compilation
- **[`ds4.c`](https://github.com/antirez/ds4/blob/main/ds4.c)** — Core engine with backend-agnostic logic and `backend_is_gpu()` helper
- **[`ds4_cli.c`](https://github.com/antirez/ds4/blob/main/ds4_cli.c)** / **[`ds4_server.c`](https://github.com/antirez/ds4/blob/main/ds4_server.c)** — Backend string parsing ("metal", "cuda", "rocm")

## Backend Selection Example

```c
/* Complete backend initialization flow */
ds4_backend select_and_init_backend(const char *name) {
    ds4_backend backend;
    
    if (!strcmp(name, "metal")) {
        backend = DS4_BACKEND_METAL;
        ds4_metal_init();        // Loads .metal kernels, creates MTLDevice
    } else if (!strcmp(name, "cuda")) {
        backend = DS4_BACKEND_CUDA;
        ds4_cuda_init();         // cudaSetDevice, cuBLAS handle creation
    } else if (!strcmp(name, "rocm")) {
        backend = DS4_BACKEND_CUDA;  // Same enum!
        ds4_rocm_init();         // HIP device setup, hipBLAS initialization
    } else {
        backend = DS4_BACKEND_CPU;
    }
    
    return backend;
}

```

## Summary

- **Metal** provides the most complete ds4 implementation with native Apple GPU optimization and exclusive tensor-parallelism support for distributed inference patterns.

- **CUDA** offers mature NVIDIA acceleration through cuBLAS but lacks tensor-parallel features due to unimplemented stubs in the shared source architecture.

- **ROCm** delivers AMD GPU compatibility via HIP, reusing CUDA code paths through conditional compilation, yet inherits the same feature limitations as CUDA.

- All three backends share high-level inference logic; differences emerge at the API binding layer and in Metal-specific kernel optimizations.

## Frequently Asked Questions

### Why does ROCm use the same enum value as CUDA?

ds4 treats ROCm as a compile-time variant rather than a distinct runtime type. The `DS4_BACKEND_CUDA` enum covers both NVIDIA and AMD paths, with `DS4_ROCM_BUILD` and `__HIP_PLATFORM_AMD__` macros selecting the correct API calls at build time. This reduces branching in the core engine.

### Can tensor-parallelism be ported to CUDA/ROCm?

The stub functions in `ds4_rocm.cu` (lines 35-41) indicate unimplemented territory, not fundamental barriers. The Metal implementation could theoretically be translated to CUDA warp primitives and ROCm wavefront operations, but this requires significant engineering effort to match the gate encoding and expert sharding logic.

### How does backend selection affect model loading?

All backends load identical weight formats. The difference lies in **where computation happens**: Metal kernels execute on Apple Neural Engine or GPU cores, CUDA kernels on NVIDIA tensor cores, and ROCm kernels on AMD compute units. Memory layout and quantization schemes remain consistent across targets.

### Is Metal performance better than CUDA for the same model?

Cross-platform performance comparisons depend on specific GPUs (Apple M-series vs. NVIDIA RTX/A100 vs. AMD Instinct). Metal's tensor-parallelism advantage matters most for **expert-parallel MoE models** where gate distribution across devices reduces latency. For standard dense models, raw TFLOPS and memory bandwidth dominate.