How Dream Server Layers Docker Compose Files for GPU Overlays

Dream Server implements a layered Docker Compose architecture where docker-compose.base.yml provides CPU-agnostic core services, and GPU-specific overlays (docker-compose.amd.yml or docker-compose.nvidia.yml) override the llama-server service to inject hardware-specific device mappings, environment variables, and backend binaries.

Light-Heart-Labs/DreamServer uses Docker Compose file merging to support heterogeneous GPU hardware without duplicating service definitions. By maintaining a single base configuration and applying Dream Server Docker Compose GPU overlays, the platform dynamically routes inference workloads to ROCm-optimized or CUDA-optimized backends while preserving identical service names and network aliases for downstream dependencies.

The Base Layer (docker-compose.base.yml)

The foundation of the stack resides in dream-server/docker-compose.base.yml, which defines the always-present services including LLM inference, Open WebUI, Dashboard API, and Dashboard UI. This file contains a comment block explaining the overlay mechanism:


# GPU overlays layered on top:

# docker compose -f docker-compose.base.yml -f docker-compose.amd.yml up -d

# docker compose -f docker-compose.base.yml -f docker-compose.nvidia.yml up -d

Within this base definition, the llama-server service uses a generic, CPU-agnostic image (ghcr.io/ggml-org/llama.cpp:server-b9014) with standard command-line arguments that function on any hardware. The base file intentionally avoids GPU-specific devices or drivers, ensuring portability across CPU-only and GPU-enabled hosts.

GPU-Specific Overlay Files

Dream Server provides two overlay files that redefine the llama-server service to target specific GPU architectures. When appended to the compose command, these files merge into the base definition using Docker's override rules, where later files take precedence.

AMD GPU Overlay (docker-compose.amd.yml)

The dream-server/docker-compose.amd.yml file configures a ROCm-optimized inference backend using a custom Lemonade-based build process. Key modifications include:

  • Build context: Replaces the base image with a build instruction that compiles llama.cpp with ROCm support
  • Device mappings: Exposes /dev/dri and /dev/kfd for GPU compute access
  • Environment variables: Sets HSA_OVERRIDE_GFX_VERSION for hardware compatibility and LEMONADE_LLAMACPP_BACKEND=rocm to select the appropriate backend
  • Volume mounts: Adds GPU-specific cache directories for model optimization artifacts

NVIDIA GPU Overlay (docker-compose.nvidia.yml)

The dream-server/docker-compose.nvidia.yml file (analogous to the AMD variant) configures CUDA acceleration by:

  • Binary optimization: Using an NVIDIA-optimized llama.cpp binary compiled with CUDA support
  • Device access: Mapping NVIDIA driver devices (/dev/nvidia*) into the container
  • Environment configuration: Setting CUDA_VISIBLE_DEVICES for device selection and LLAMA_CUDA=1 to enable GPU acceleration
  • LiteLLM routing: Configuring the LiteLLM service to direct requests to the NVIDIA-specific backend

Both overlays preserve the service name llama-server, ensuring that dependent services like Open WebUI and the Dashboard API maintain their network references regardless of the underlying GPU hardware.

How Docker Compose Merges the Layers

Docker Compose applies files sequentially, with later definitions overriding earlier ones at the service level. When executing:

docker compose \
  -f dream-server/docker-compose.base.yml \
  -f dream-server/docker-compose.amd.yml \
  up -d

The merge process:

  1. Ingests the base service definitions, network configurations, and volume mappings
  2. Overlays the AMD-specific llama-server definition, replacing the image field with the build directive and injecting ROCm-specific devices
  3. Preserves all base-defined services (Open WebUI, Dashboard API) that are not mentioned in the overlay
  4. Maintains consistent Docker network aliases so llama-server resolves correctly for internal service communication

This declarative approach eliminates conditional logic in the compose files; the hardware target is determined entirely by which overlay file appears last in the command sequence.

Automating Extension Discovery with resolve-compose-stack.sh

For deployments requiring additional services (ComfyUI, Whisper, etc.), Dream Server includes scripts/resolve-compose-stack.sh. This helper script discovers extension-specific compose files located at extensions/services/*/compose.yaml and merges them with the base and GPU overlay configurations.

The resolution flow executes as follows:

  1. Base inclusion: Loads docker-compose.base.yml as the foundation
  2. GPU overlay selection: Appends either docker-compose.amd.yml or docker-compose.nvidia.yml
  3. Extension discovery: Scans extensions/services/*/ directories for compose.yaml files
  4. Stack synthesis: Outputs a unified compose configuration to stdout for piping into docker compose

Because the script appends extension files after the GPU overlay, extension services can reference the already-configured llama-server with its hardware-specific backend active.

Practical Usage Examples

Start a Dream Server instance with AMD GPU support:

docker compose \
  -f dream-server/docker-compose.base.yml \
  -f dream-server/docker-compose.amd.yml \
  up -d

Start with NVIDIA GPU acceleration:

docker compose \
  -f dream-server/docker-compose.base.yml \
  -f dream-server/docker-compose.nvidia.yml \
  up -d

Deploy the full stack including auto-discovered extensions:

scripts/resolve-compose-stack.sh \
  -f dream-server/docker-compose.base.yml \
  -f dream-server/docker-compose.amd.yml \
  | docker compose -f - up -d

Summary

  • Dream Server maintains a single docker-compose.base.yml containing CPU-agnostic core services including the generic llama-server definition.
  • GPU overlays (docker-compose.amd.yml and docker-compose.nvidia.yml) override the llama-server service to inject ROCm or CUDA-specific build instructions, devices, and environment variables.
  • Docker Compose merging applies later files on top of earlier ones, allowing the overlay to replace specific fields without duplicating the entire service definition.
  • Service name consistency ensures that Open WebUI, Dashboard API, and other dependencies always reference llama-server regardless of the hardware backend.
  • resolve-compose-stack.sh automates the inclusion of extension services from extensions/services/*/compose.yaml, appending them after the GPU overlay in the final stack configuration.

Frequently Asked Questions

Can I run Dream Server without a GPU using these overlay files?

Yes. Simply omit the GPU overlay files and run only the base compose file: docker compose -f docker-compose.base.yml up -d. The base configuration runs llama-server in CPU-only mode using the standard ghcr.io/ggml-org/llama.cpp image without requiring special device access or GPU drivers.

Why does the llama-server service name remain identical across different overlays?

Maintaining the same service name ensures that downstream services defined in docker-compose.base.yml—such as Open WebUI and the Dashboard API—can resolve the inference backend at the fixed hostname llama-server on the Docker network. This abstraction allows the same service configuration to work with AMD ROCm, NVIDIA CUDA, or CPU backends without modifying dependent service definitions.

How do I add a custom extension to the Docker Compose stack?

Create a compose.yaml file inside a new subdirectory under extensions/services/ (e.g., extensions/services/my-custom-service/compose.yaml). The scripts/resolve-compose-stack.sh utility automatically discovers this file during stack resolution and appends it after the GPU overlay, integrating your extension into the final configuration with access to the configured llama-server backend.

What happens if I specify both AMD and NVIDIA overlays simultaneously?

Avoid specifying both overlays in the same command. Because Docker Compose applies files sequentially with later overrides taking precedence, the last-specified overlay would overwrite the previous one's llama-server definition, potentially resulting in conflicting device mappings (/dev/dri vs /dev/nvidia*) and environment variables. Deploy only one GPU overlay that matches your hardware configuration.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →