# How ODS Manages Docker Compose Configurations for Different GPU Types

> Discover how ODS manages Docker Compose for diverse GPU types. Learn about its layered architecture and automatic configuration merging for AMD, NVIDIA, and Apple Silicon.

- Repository: [Osmantic/ODS](https://github.com/Osmantic/ODS)
- Tags: how-to-guide
- Published: 2026-08-30

---

**ODS uses a layered Docker Compose architecture that keeps core service definitions in a base file and applies GPU-specific overlays for AMD, NVIDIA, or Apple Silicon hardware, automatically merged at runtime based on the `GPU_BACKEND` environment variable.**

The Osmantic/ODS repository implements a sophisticated, modular approach to Docker Compose management that allows the same core application stack to run across heterogeneous GPU hardware without duplicating configuration. By separating the baseline service definitions from hardware-specific overrides, ODS ensures that the `llama-server`, Open-WebUI, and dashboard services remain maintainable while adapting dynamically to the host's acceleration capabilities.

## Layered Compose Architecture

ODS organizes its Docker Compose files into a base configuration and minimal overlay files that contain only the delta required for each GPU backend. This pattern prevents configuration drift and keeps the core stack reusable for CPU-only deployments.

### The Base Configuration

At [`ods/docker-compose.base.yml`](https://github.com/Osmantic/ODS/blob/main/ods/docker-compose.base.yml), ODS defines the complete application stack—including the `llama-server`, Open-WebUI, and dashboard-api services—without any GPU-specific settings. This file is intentionally generic, omitting device reservations, driver-specific images, or vendor environment variables so that it can serve as the foundation for every deployment target.

### GPU-Specific Overlay Files

For each supported acceleration backend, ODS provides a thin overlay that modifies only the services requiring hardware adaptation:

- **AMD** – [`ods/docker-compose.amd.yml`](https://github.com/Osmantic/ODS/blob/main/ods/docker-compose.amd.yml) replaces the default `llama-server` image with a Lemonade-based ROCm build, mounts the `/dev/dri` and `/dev/kfd` devices, and injects AMD-specific environment variables like `HSA_OVERRIDE_GFX_VERSION` and `ROCBLAS_USE_HIPBLASLT`. It also declares named volumes including `lemonade-cache`, `lemonade-llama`, and `lemonade-recipe` to persist compiled model binaries.

- **NVIDIA** – [`ods/docker-compose.nvidia.yml`](https://github.com/Osmantic/ODS/blob/main/ods/docker-compose.nvidia.yml) swaps the server image for a CUDA build (e.g., `ghcr.io/ggml-org/llama.cpp:server-cuda-b9014`) and requests the NVIDIA driver runtime using the `deploy.resources.reservations.devices` syntax with `driver: nvidia` and `capabilities: [gpu]`. It additionally sets `AUDIO_STT_MODEL` to enable GPU-accelerated Whisper transcription.

- **Apple Silicon** – [`ods/docker-compose.apple.yml`](https://github.com/Osmantic/ODS/blob/main/ods/docker-compose.apple.yml) selects an ARM-64 optimized image and sets the `LLAMA_NO_METAL` environment variable to ensure stable CPU inference within Docker containers on macOS, effectively disabling Metal GPU passthrough that is unreliable in containerized contexts.

## Environment-Driven Stack Resolution

Rather than requiring users to manually specify the correct overlay, ODS automates selection through the `GPU_BACKEND` environment variable, populated during the installation phase.

### Hardware Detection

During **phase-06** of the installation process, the script at [`ods/installers/lib/detection.sh`](https://github.com/Osmantic/ODS/blob/main/ods/installers/lib/detection.sh) probes the host hardware and writes the detected backend into the `.env` file:

```bash
GPU_BACKEND=amd   # or nvidia, apple, cpu

```

### Merge Logic

The helper script [`ods/scripts/resolve-compose-stack.sh`](https://github.com/Osmantic/ODS/blob/main/ods/scripts/resolve-compose-stack.sh) reads this variable and constructs the final compose command by merging the base file with the appropriate overlay:

```bash
docker compose -f docker-compose.base.yml -f docker-compose.${GPU_BACKEND}.yml up -d

```

Since each overlay only touches the services it needs to modify—typically just the `llama-server`—the base stack remains untouched and can be reused for CPU-only installations or testing environments.

## Service-Level Overrides

Each overlay file applies surgical modifications to the base services:

**Image Replacement**
Overlay files override the container image to pull architecture-specific builds (ROCm for AMD, CUDA for NVIDIA, ARM64 for Apple).

**Device Allocation**
- **NVIDIA**: Uses the standard Docker device reservation syntax requesting `nvidia` driver access.
- **AMD**: Mounts character devices `/dev/dri` and `/dev/kfd` directly into the container and adds the host's render and video GID groups to the container process.

**Environment Variables**
- **AMD**: `HSA_OVERRIDE_GFX_VERSION`, `ROCBLAS_USE_HIPBLASLT`, and Lemonade-specific flags.
- **NVIDIA**: `AUDIO_STT_MODEL` for Whisper optimization.
- **Apple**: `LLAMA_NO_METAL` to disable Metal within Docker.

**Persistent Volumes**
The AMD overlay mounts specific named volumes to cache compiled GPU kernels and model artifacts, preventing redundant compilation across container restarts.

## Running the Stack

To start ODS with a specific GPU backend manually, specify the base file and the corresponding overlay:

**AMD GPU:**

```bash

# Ensure .env contains GPU_BACKEND=amd

docker compose -f docker-compose.base.yml -f docker-compose.amd.yml up -d

```

**NVIDIA GPU:**

```bash
docker compose -f docker-compose.base.yml -f docker-compose.nvidia.yml up -d

```

**Apple Silicon (CPU fallback):**

```bash
docker compose -f docker-compose.base.yml -f docker-compose.apple.yml up -d

```

For programmatic selection in custom scripts, mirror the installer logic:

```bash
#!/usr/bin/env bash
source ./ods/installers/lib/detection.sh   # Detects GPU and exports GPU_BACKEND

OVERLAY="docker-compose.${GPU_BACKEND}.yml"
docker compose -f docker-compose.base.yml -f "$OVERLAY" up -d

```

## Summary

- **ODS employs a layered compose strategy** where [`ods/docker-compose.base.yml`](https://github.com/Osmantic/ODS/blob/main/ods/docker-compose.base.yml) defines the core stack and overlay files ([`amd.yml`](https://github.com/Osmantic/ODS/blob/main/amd.yml), [`nvidia.yml`](https://github.com/Osmantic/ODS/blob/main/nvidia.yml), [`apple.yml`](https://github.com/Osmantic/ODS/blob/main/apple.yml)) supply hardware-specific modifications.
- **Automatic detection** via [`ods/installers/lib/detection.sh`](https://github.com/Osmantic/ODS/blob/main/ods/installers/lib/detection.sh) sets the `GPU_BACKEND` variable, which [`ods/scripts/resolve-compose-stack.sh`](https://github.com/Osmantic/ODS/blob/main/ods/scripts/resolve-compose-stack.sh) uses to merge the correct files at runtime.
- **Overlays are minimal and focused**, typically changing only the `llama-server` image, device mounts, and environment variables while leaving other services untouched.
- **AMD configurations** leverage ROCm through Lemonade with dedicated volumes for caching, while **NVIDIA** uses standard CUDA runtime reservations and **Apple Silicon** forces CPU inference via `LLAMA_NO_METAL`.

## Frequently Asked Questions

### How does ODS detect which GPU backend to use?

During installation, the [`ods/installers/lib/detection.sh`](https://github.com/Osmantic/ODS/blob/main/ods/installers/lib/detection.sh) script probes the host hardware to identify AMD, NVIDIA, or Apple Silicon GPUs, then writes the result as `GPU_BACKEND=amd|nvidia|apple|cpu` into the project's `.env` file. The [`resolve-compose-stack.sh`](https://github.com/Osmantic/ODS/blob/main/resolve-compose-stack.sh) script reads this variable to determine which overlay file to include when starting the stack.

### Can I run ODS without a GPU?

Yes. The base configuration at [`ods/docker-compose.base.yml`](https://github.com/Osmantic/ODS/blob/main/ods/docker-compose.base.yml) contains no GPU-specific settings and runs entirely on CPU. If the detection script finds no supported GPU, it sets `GPU_BACKEND=cpu` or leaves the variable empty, causing the resolution script to invoke only the base file without any hardware overlays.

### What is the purpose of the Lemonade cache volumes in the AMD configuration?

The [`docker-compose.amd.yml`](https://github.com/Osmantic/ODS/blob/main/docker-compose.amd.yml) file declares volumes named `lemonade-cache`, `lemonade-llama`, and `lemonade-recipe` to persist compiled model binaries and ROCm kernel artifacts across container restarts. This caching significantly reduces startup time for subsequent runs of the `llama-server` on AMD hardware by avoiding redundant compilation of GPU kernels.

### Why does the Apple Silicon overlay disable Metal with `LLAMA_NO_METAL`?

Docker Desktop on macOS does not support passing the Metal GPU into Linux containers reliably. The [`ods/docker-compose.apple.yml`](https://github.com/Osmantic/ODS/blob/main/ods/docker-compose.apple.yml) overlay sets `LLAMA_NO_METAL=1` to force the `llama-server` to use optimized ARM-64 CPU inference instead, ensuring stable performance on Apple Silicon Macs without attempting unsupported GPU passthrough.