How Dream Server Layers Docker Compose Files for GPU Overlays
Dream Server implements a layered Docker Compose architecture where docker-compose.base.yml provides CPU-agnostic core services, and GPU-specific overlays (docker-compose.amd.yml or docker-compose.nvidia.yml) override the llama-server service to inject hardware-specific device mappings, environment variables, and backend binaries.
Light-Heart-Labs/DreamServer uses Docker Compose file merging to support heterogeneous GPU hardware without duplicating service definitions. By maintaining a single base configuration and applying Dream Server Docker Compose GPU overlays, the platform dynamically routes inference workloads to ROCm-optimized or CUDA-optimized backends while preserving identical service names and network aliases for downstream dependencies.
The Base Layer (docker-compose.base.yml)
The foundation of the stack resides in dream-server/docker-compose.base.yml, which defines the always-present services including LLM inference, Open WebUI, Dashboard API, and Dashboard UI. This file contains a comment block explaining the overlay mechanism:
# GPU overlays layered on top:
# docker compose -f docker-compose.base.yml -f docker-compose.amd.yml up -d
# docker compose -f docker-compose.base.yml -f docker-compose.nvidia.yml up -d
Within this base definition, the llama-server service uses a generic, CPU-agnostic image (ghcr.io/ggml-org/llama.cpp:server-b9014) with standard command-line arguments that function on any hardware. The base file intentionally avoids GPU-specific devices or drivers, ensuring portability across CPU-only and GPU-enabled hosts.
GPU-Specific Overlay Files
Dream Server provides two overlay files that redefine the llama-server service to target specific GPU architectures. When appended to the compose command, these files merge into the base definition using Docker's override rules, where later files take precedence.
AMD GPU Overlay (docker-compose.amd.yml)
The dream-server/docker-compose.amd.yml file configures a ROCm-optimized inference backend using a custom Lemonade-based build process. Key modifications include:
- Build context: Replaces the base image with a build instruction that compiles
llama.cppwith ROCm support - Device mappings: Exposes
/dev/driand/dev/kfdfor GPU compute access - Environment variables: Sets
HSA_OVERRIDE_GFX_VERSIONfor hardware compatibility andLEMONADE_LLAMACPP_BACKEND=rocmto select the appropriate backend - Volume mounts: Adds GPU-specific cache directories for model optimization artifacts
NVIDIA GPU Overlay (docker-compose.nvidia.yml)
The dream-server/docker-compose.nvidia.yml file (analogous to the AMD variant) configures CUDA acceleration by:
- Binary optimization: Using an NVIDIA-optimized
llama.cppbinary compiled with CUDA support - Device access: Mapping NVIDIA driver devices (
/dev/nvidia*) into the container - Environment configuration: Setting
CUDA_VISIBLE_DEVICESfor device selection andLLAMA_CUDA=1to enable GPU acceleration - LiteLLM routing: Configuring the LiteLLM service to direct requests to the NVIDIA-specific backend
Both overlays preserve the service name llama-server, ensuring that dependent services like Open WebUI and the Dashboard API maintain their network references regardless of the underlying GPU hardware.
How Docker Compose Merges the Layers
Docker Compose applies files sequentially, with later definitions overriding earlier ones at the service level. When executing:
docker compose \
-f dream-server/docker-compose.base.yml \
-f dream-server/docker-compose.amd.yml \
up -d
The merge process:
- Ingests the base service definitions, network configurations, and volume mappings
- Overlays the AMD-specific
llama-serverdefinition, replacing theimagefield with thebuilddirective and injecting ROCm-specific devices - Preserves all base-defined services (Open WebUI, Dashboard API) that are not mentioned in the overlay
- Maintains consistent Docker network aliases so
llama-serverresolves correctly for internal service communication
This declarative approach eliminates conditional logic in the compose files; the hardware target is determined entirely by which overlay file appears last in the command sequence.
Automating Extension Discovery with resolve-compose-stack.sh
For deployments requiring additional services (ComfyUI, Whisper, etc.), Dream Server includes scripts/resolve-compose-stack.sh. This helper script discovers extension-specific compose files located at extensions/services/*/compose.yaml and merges them with the base and GPU overlay configurations.
The resolution flow executes as follows:
- Base inclusion: Loads
docker-compose.base.ymlas the foundation - GPU overlay selection: Appends either
docker-compose.amd.ymlordocker-compose.nvidia.yml - Extension discovery: Scans
extensions/services/*/directories forcompose.yamlfiles - Stack synthesis: Outputs a unified compose configuration to stdout for piping into
docker compose
Because the script appends extension files after the GPU overlay, extension services can reference the already-configured llama-server with its hardware-specific backend active.
Practical Usage Examples
Start a Dream Server instance with AMD GPU support:
docker compose \
-f dream-server/docker-compose.base.yml \
-f dream-server/docker-compose.amd.yml \
up -d
Start with NVIDIA GPU acceleration:
docker compose \
-f dream-server/docker-compose.base.yml \
-f dream-server/docker-compose.nvidia.yml \
up -d
Deploy the full stack including auto-discovered extensions:
scripts/resolve-compose-stack.sh \
-f dream-server/docker-compose.base.yml \
-f dream-server/docker-compose.amd.yml \
| docker compose -f - up -d
Summary
- Dream Server maintains a single
docker-compose.base.ymlcontaining CPU-agnostic core services including the genericllama-serverdefinition. - GPU overlays (
docker-compose.amd.ymlanddocker-compose.nvidia.yml) override thellama-serverservice to inject ROCm or CUDA-specific build instructions, devices, and environment variables. - Docker Compose merging applies later files on top of earlier ones, allowing the overlay to replace specific fields without duplicating the entire service definition.
- Service name consistency ensures that Open WebUI, Dashboard API, and other dependencies always reference
llama-serverregardless of the hardware backend. resolve-compose-stack.shautomates the inclusion of extension services fromextensions/services/*/compose.yaml, appending them after the GPU overlay in the final stack configuration.
Frequently Asked Questions
Can I run Dream Server without a GPU using these overlay files?
Yes. Simply omit the GPU overlay files and run only the base compose file: docker compose -f docker-compose.base.yml up -d. The base configuration runs llama-server in CPU-only mode using the standard ghcr.io/ggml-org/llama.cpp image without requiring special device access or GPU drivers.
Why does the llama-server service name remain identical across different overlays?
Maintaining the same service name ensures that downstream services defined in docker-compose.base.yml—such as Open WebUI and the Dashboard API—can resolve the inference backend at the fixed hostname llama-server on the Docker network. This abstraction allows the same service configuration to work with AMD ROCm, NVIDIA CUDA, or CPU backends without modifying dependent service definitions.
How do I add a custom extension to the Docker Compose stack?
Create a compose.yaml file inside a new subdirectory under extensions/services/ (e.g., extensions/services/my-custom-service/compose.yaml). The scripts/resolve-compose-stack.sh utility automatically discovers this file during stack resolution and appends it after the GPU overlay, integrating your extension into the final configuration with access to the configured llama-server backend.
What happens if I specify both AMD and NVIDIA overlays simultaneously?
Avoid specifying both overlays in the same command. Because Docker Compose applies files sequentially with later overrides taking precedence, the last-specified overlay would overwrite the previous one's llama-server definition, potentially resulting in conflicting device mappings (/dev/dri vs /dev/nvidia*) and environment variables. Deploy only one GPU overlay that matches your hardware configuration.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →