How ODS Manages Docker Compose Configurations for Different GPU Types

ODS uses a layered Docker Compose architecture that keeps core service definitions in a base file and applies GPU-specific overlays for AMD, NVIDIA, or Apple Silicon hardware, automatically merged at runtime based on the GPU_BACKEND environment variable.

The Osmantic/ODS repository implements a sophisticated, modular approach to Docker Compose management that allows the same core application stack to run across heterogeneous GPU hardware without duplicating configuration. By separating the baseline service definitions from hardware-specific overrides, ODS ensures that the llama-server, Open-WebUI, and dashboard services remain maintainable while adapting dynamically to the host's acceleration capabilities.

Layered Compose Architecture

ODS organizes its Docker Compose files into a base configuration and minimal overlay files that contain only the delta required for each GPU backend. This pattern prevents configuration drift and keeps the core stack reusable for CPU-only deployments.

The Base Configuration

At ods/docker-compose.base.yml, ODS defines the complete application stack—including the llama-server, Open-WebUI, and dashboard-api services—without any GPU-specific settings. This file is intentionally generic, omitting device reservations, driver-specific images, or vendor environment variables so that it can serve as the foundation for every deployment target.

GPU-Specific Overlay Files

For each supported acceleration backend, ODS provides a thin overlay that modifies only the services requiring hardware adaptation:

  • AMD – ods/docker-compose.amd.yml replaces the default llama-server image with a Lemonade-based ROCm build, mounts the /dev/dri and /dev/kfd devices, and injects AMD-specific environment variables like HSA_OVERRIDE_GFX_VERSION and ROCBLAS_USE_HIPBLASLT. It also declares named volumes including lemonade-cache, lemonade-llama, and lemonade-recipe to persist compiled model binaries.

  • NVIDIA – ods/docker-compose.nvidia.yml swaps the server image for a CUDA build (e.g., ghcr.io/ggml-org/llama.cpp:server-cuda-b9014) and requests the NVIDIA driver runtime using the deploy.resources.reservations.devices syntax with driver: nvidia and capabilities: [gpu]. It additionally sets AUDIO_STT_MODEL to enable GPU-accelerated Whisper transcription.

  • Apple Silicon – ods/docker-compose.apple.yml selects an ARM-64 optimized image and sets the LLAMA_NO_METAL environment variable to ensure stable CPU inference within Docker containers on macOS, effectively disabling Metal GPU passthrough that is unreliable in containerized contexts.

Environment-Driven Stack Resolution

Rather than requiring users to manually specify the correct overlay, ODS automates selection through the GPU_BACKEND environment variable, populated during the installation phase.

Hardware Detection

During phase-06 of the installation process, the script at ods/installers/lib/detection.sh probes the host hardware and writes the detected backend into the .env file:

GPU_BACKEND=amd   # or nvidia, apple, cpu

Merge Logic

The helper script ods/scripts/resolve-compose-stack.sh reads this variable and constructs the final compose command by merging the base file with the appropriate overlay:

docker compose -f docker-compose.base.yml -f docker-compose.${GPU_BACKEND}.yml up -d

Since each overlay only touches the services it needs to modify—typically just the llama-server—the base stack remains untouched and can be reused for CPU-only installations or testing environments.

Service-Level Overrides

Each overlay file applies surgical modifications to the base services:

Image Replacement Overlay files override the container image to pull architecture-specific builds (ROCm for AMD, CUDA for NVIDIA, ARM64 for Apple).

Device Allocation

  • NVIDIA: Uses the standard Docker device reservation syntax requesting nvidia driver access.
  • AMD: Mounts character devices /dev/dri and /dev/kfd directly into the container and adds the host's render and video GID groups to the container process.

Environment Variables

  • AMD: HSA_OVERRIDE_GFX_VERSION, ROCBLAS_USE_HIPBLASLT, and Lemonade-specific flags.
  • NVIDIA: AUDIO_STT_MODEL for Whisper optimization.
  • Apple: LLAMA_NO_METAL to disable Metal within Docker.

Persistent Volumes The AMD overlay mounts specific named volumes to cache compiled GPU kernels and model artifacts, preventing redundant compilation across container restarts.

Running the Stack

To start ODS with a specific GPU backend manually, specify the base file and the corresponding overlay:

AMD GPU:


# Ensure .env contains GPU_BACKEND=amd

docker compose -f docker-compose.base.yml -f docker-compose.amd.yml up -d

NVIDIA GPU:

docker compose -f docker-compose.base.yml -f docker-compose.nvidia.yml up -d

Apple Silicon (CPU fallback):

docker compose -f docker-compose.base.yml -f docker-compose.apple.yml up -d

For programmatic selection in custom scripts, mirror the installer logic:

#!/usr/bin/env bash
source ./ods/installers/lib/detection.sh   # Detects GPU and exports GPU_BACKEND

OVERLAY="docker-compose.${GPU_BACKEND}.yml"
docker compose -f docker-compose.base.yml -f "$OVERLAY" up -d

Summary

  • ODS employs a layered compose strategy where ods/docker-compose.base.yml defines the core stack and overlay files (amd.yml, nvidia.yml, apple.yml) supply hardware-specific modifications.
  • Automatic detection via ods/installers/lib/detection.sh sets the GPU_BACKEND variable, which ods/scripts/resolve-compose-stack.sh uses to merge the correct files at runtime.
  • Overlays are minimal and focused, typically changing only the llama-server image, device mounts, and environment variables while leaving other services untouched.
  • AMD configurations leverage ROCm through Lemonade with dedicated volumes for caching, while NVIDIA uses standard CUDA runtime reservations and Apple Silicon forces CPU inference via LLAMA_NO_METAL.

Frequently Asked Questions

How does ODS detect which GPU backend to use?

During installation, the ods/installers/lib/detection.sh script probes the host hardware to identify AMD, NVIDIA, or Apple Silicon GPUs, then writes the result as GPU_BACKEND=amd|nvidia|apple|cpu into the project's .env file. The resolve-compose-stack.sh script reads this variable to determine which overlay file to include when starting the stack.

Can I run ODS without a GPU?

Yes. The base configuration at ods/docker-compose.base.yml contains no GPU-specific settings and runs entirely on CPU. If the detection script finds no supported GPU, it sets GPU_BACKEND=cpu or leaves the variable empty, causing the resolution script to invoke only the base file without any hardware overlays.

What is the purpose of the Lemonade cache volumes in the AMD configuration?

The docker-compose.amd.yml file declares volumes named lemonade-cache, lemonade-llama, and lemonade-recipe to persist compiled model binaries and ROCm kernel artifacts across container restarts. This caching significantly reduces startup time for subsequent runs of the llama-server on AMD hardware by avoiding redundant compilation of GPU kernels.

Why does the Apple Silicon overlay disable Metal with LLAMA_NO_METAL?

Docker Desktop on macOS does not support passing the Metal GPU into Linux containers reliably. The ods/docker-compose.apple.yml overlay sets LLAMA_NO_METAL=1 to force the llama-server to use optimized ARM-64 CPU inference instead, ensuring stable performance on Apple Silicon Macs without attempting unsupported GPU passthrough.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →