How ODS Manages Docker Compose Configurations for Different GPU Types
ODS uses a layered Docker Compose architecture that keeps core service definitions in a base file and applies GPU-specific overlays for AMD, NVIDIA, or Apple Silicon hardware, automatically merged at runtime based on the GPU_BACKEND environment variable.
The Osmantic/ODS repository implements a sophisticated, modular approach to Docker Compose management that allows the same core application stack to run across heterogeneous GPU hardware without duplicating configuration. By separating the baseline service definitions from hardware-specific overrides, ODS ensures that the llama-server, Open-WebUI, and dashboard services remain maintainable while adapting dynamically to the host's acceleration capabilities.
Layered Compose Architecture
ODS organizes its Docker Compose files into a base configuration and minimal overlay files that contain only the delta required for each GPU backend. This pattern prevents configuration drift and keeps the core stack reusable for CPU-only deployments.
The Base Configuration
At ods/docker-compose.base.yml, ODS defines the complete application stack—including the llama-server, Open-WebUI, and dashboard-api services—without any GPU-specific settings. This file is intentionally generic, omitting device reservations, driver-specific images, or vendor environment variables so that it can serve as the foundation for every deployment target.
GPU-Specific Overlay Files
For each supported acceleration backend, ODS provides a thin overlay that modifies only the services requiring hardware adaptation:
-
AMD –
ods/docker-compose.amd.ymlreplaces the defaultllama-serverimage with a Lemonade-based ROCm build, mounts the/dev/driand/dev/kfddevices, and injects AMD-specific environment variables likeHSA_OVERRIDE_GFX_VERSIONandROCBLAS_USE_HIPBLASLT. It also declares named volumes includinglemonade-cache,lemonade-llama, andlemonade-recipeto persist compiled model binaries. -
NVIDIA –
ods/docker-compose.nvidia.ymlswaps the server image for a CUDA build (e.g.,ghcr.io/ggml-org/llama.cpp:server-cuda-b9014) and requests the NVIDIA driver runtime using thedeploy.resources.reservations.devicessyntax withdriver: nvidiaandcapabilities: [gpu]. It additionally setsAUDIO_STT_MODELto enable GPU-accelerated Whisper transcription. -
Apple Silicon –
ods/docker-compose.apple.ymlselects an ARM-64 optimized image and sets theLLAMA_NO_METALenvironment variable to ensure stable CPU inference within Docker containers on macOS, effectively disabling Metal GPU passthrough that is unreliable in containerized contexts.
Environment-Driven Stack Resolution
Rather than requiring users to manually specify the correct overlay, ODS automates selection through the GPU_BACKEND environment variable, populated during the installation phase.
Hardware Detection
During phase-06 of the installation process, the script at ods/installers/lib/detection.sh probes the host hardware and writes the detected backend into the .env file:
GPU_BACKEND=amd # or nvidia, apple, cpu
Merge Logic
The helper script ods/scripts/resolve-compose-stack.sh reads this variable and constructs the final compose command by merging the base file with the appropriate overlay:
docker compose -f docker-compose.base.yml -f docker-compose.${GPU_BACKEND}.yml up -d
Since each overlay only touches the services it needs to modify—typically just the llama-server—the base stack remains untouched and can be reused for CPU-only installations or testing environments.
Service-Level Overrides
Each overlay file applies surgical modifications to the base services:
Image Replacement Overlay files override the container image to pull architecture-specific builds (ROCm for AMD, CUDA for NVIDIA, ARM64 for Apple).
Device Allocation
- NVIDIA: Uses the standard Docker device reservation syntax requesting
nvidiadriver access. - AMD: Mounts character devices
/dev/driand/dev/kfddirectly into the container and adds the host's render and video GID groups to the container process.
Environment Variables
- AMD:
HSA_OVERRIDE_GFX_VERSION,ROCBLAS_USE_HIPBLASLT, and Lemonade-specific flags. - NVIDIA:
AUDIO_STT_MODELfor Whisper optimization. - Apple:
LLAMA_NO_METALto disable Metal within Docker.
Persistent Volumes The AMD overlay mounts specific named volumes to cache compiled GPU kernels and model artifacts, preventing redundant compilation across container restarts.
Running the Stack
To start ODS with a specific GPU backend manually, specify the base file and the corresponding overlay:
AMD GPU:
# Ensure .env contains GPU_BACKEND=amd
docker compose -f docker-compose.base.yml -f docker-compose.amd.yml up -d
NVIDIA GPU:
docker compose -f docker-compose.base.yml -f docker-compose.nvidia.yml up -d
Apple Silicon (CPU fallback):
docker compose -f docker-compose.base.yml -f docker-compose.apple.yml up -d
For programmatic selection in custom scripts, mirror the installer logic:
#!/usr/bin/env bash
source ./ods/installers/lib/detection.sh # Detects GPU and exports GPU_BACKEND
OVERLAY="docker-compose.${GPU_BACKEND}.yml"
docker compose -f docker-compose.base.yml -f "$OVERLAY" up -d
Summary
- ODS employs a layered compose strategy where
ods/docker-compose.base.ymldefines the core stack and overlay files (amd.yml,nvidia.yml,apple.yml) supply hardware-specific modifications. - Automatic detection via
ods/installers/lib/detection.shsets theGPU_BACKENDvariable, whichods/scripts/resolve-compose-stack.shuses to merge the correct files at runtime. - Overlays are minimal and focused, typically changing only the
llama-serverimage, device mounts, and environment variables while leaving other services untouched. - AMD configurations leverage ROCm through Lemonade with dedicated volumes for caching, while NVIDIA uses standard CUDA runtime reservations and Apple Silicon forces CPU inference via
LLAMA_NO_METAL.
Frequently Asked Questions
How does ODS detect which GPU backend to use?
During installation, the ods/installers/lib/detection.sh script probes the host hardware to identify AMD, NVIDIA, or Apple Silicon GPUs, then writes the result as GPU_BACKEND=amd|nvidia|apple|cpu into the project's .env file. The resolve-compose-stack.sh script reads this variable to determine which overlay file to include when starting the stack.
Can I run ODS without a GPU?
Yes. The base configuration at ods/docker-compose.base.yml contains no GPU-specific settings and runs entirely on CPU. If the detection script finds no supported GPU, it sets GPU_BACKEND=cpu or leaves the variable empty, causing the resolution script to invoke only the base file without any hardware overlays.
What is the purpose of the Lemonade cache volumes in the AMD configuration?
The docker-compose.amd.yml file declares volumes named lemonade-cache, lemonade-llama, and lemonade-recipe to persist compiled model binaries and ROCm kernel artifacts across container restarts. This caching significantly reduces startup time for subsequent runs of the llama-server on AMD hardware by avoiding redundant compilation of GPU kernels.
Why does the Apple Silicon overlay disable Metal with LLAMA_NO_METAL?
Docker Desktop on macOS does not support passing the Metal GPU into Linux containers reliably. The ods/docker-compose.apple.yml overlay sets LLAMA_NO_METAL=1 to force the llama-server to use optimized ARM-64 CPU inference instead, ensuring stable performance on Apple Silicon Macs without attempting unsupported GPU passthrough.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →