# How Dream Server Layers Docker Compose Files for GPU Overlays

> Discover how Dream Server layers Docker Compose files for efficient GPU overlays. Understand CPU-agnostic services and GPU-specific overrides for your AI workloads.

- Repository: [Light Heart Labs/DreamServer](https://github.com/Light-Heart-Labs/DreamServer)
- Tags: how-to-guide
- Published: 2026-05-18

---

**Dream Server implements a layered Docker Compose architecture where [`docker-compose.base.yml`](https://github.com/Light-Heart-Labs/DreamServer/blob/main/docker-compose.base.yml) provides CPU-agnostic core services, and GPU-specific overlays ([`docker-compose.amd.yml`](https://github.com/Light-Heart-Labs/DreamServer/blob/main/docker-compose.amd.yml) or [`docker-compose.nvidia.yml`](https://github.com/Light-Heart-Labs/DreamServer/blob/main/docker-compose.nvidia.yml)) override the `llama-server` service to inject hardware-specific device mappings, environment variables, and backend binaries.**

Light-Heart-Labs/DreamServer uses Docker Compose file merging to support heterogeneous GPU hardware without duplicating service definitions. By maintaining a single base configuration and applying Dream Server Docker Compose GPU overlays, the platform dynamically routes inference workloads to ROCm-optimized or CUDA-optimized backends while preserving identical service names and network aliases for downstream dependencies.

## The Base Layer (docker-compose.base.yml)

The foundation of the stack resides in [`dream-server/docker-compose.base.yml`](https://github.com/Light-Heart-Labs/DreamServer/blob/main/dream-server/docker-compose.base.yml), which defines the always-present services including LLM inference, Open WebUI, Dashboard API, and Dashboard UI. This file contains a comment block explaining the overlay mechanism:

```yaml

# GPU overlays layered on top:

# docker compose -f docker-compose.base.yml -f docker-compose.amd.yml up -d

# docker compose -f docker-compose.base.yml -f docker-compose.nvidia.yml up -d

```

Within this base definition, the `llama-server` service uses a generic, CPU-agnostic image (`ghcr.io/ggml-org/llama.cpp:server-b9014`) with standard command-line arguments that function on any hardware. The base file intentionally avoids GPU-specific devices or drivers, ensuring portability across CPU-only and GPU-enabled hosts.

## GPU-Specific Overlay Files

Dream Server provides two overlay files that redefine the `llama-server` service to target specific GPU architectures. When appended to the compose command, these files merge into the base definition using Docker's override rules, where later files take precedence.

### AMD GPU Overlay (docker-compose.amd.yml)

The [`dream-server/docker-compose.amd.yml`](https://github.com/Light-Heart-Labs/DreamServer/blob/main/dream-server/docker-compose.amd.yml) file configures a ROCm-optimized inference backend using a custom Lemonade-based build process. Key modifications include:

- **Build context**: Replaces the base image with a build instruction that compiles [`llama.cpp`](https://github.com/Light-Heart-Labs/DreamServer/blob/main/llama.cpp) with ROCm support
- **Device mappings**: Exposes `/dev/dri` and `/dev/kfd` for GPU compute access
- **Environment variables**: Sets `HSA_OVERRIDE_GFX_VERSION` for hardware compatibility and `LEMONADE_LLAMACPP_BACKEND=rocm` to select the appropriate backend
- **Volume mounts**: Adds GPU-specific cache directories for model optimization artifacts

### NVIDIA GPU Overlay (docker-compose.nvidia.yml)

The [`dream-server/docker-compose.nvidia.yml`](https://github.com/Light-Heart-Labs/DreamServer/blob/main/dream-server/docker-compose.nvidia.yml) file (analogous to the AMD variant) configures CUDA acceleration by:

- **Binary optimization**: Using an NVIDIA-optimized [`llama.cpp`](https://github.com/Light-Heart-Labs/DreamServer/blob/main/llama.cpp) binary compiled with CUDA support
- **Device access**: Mapping NVIDIA driver devices (`/dev/nvidia*`) into the container
- **Environment configuration**: Setting `CUDA_VISIBLE_DEVICES` for device selection and `LLAMA_CUDA=1` to enable GPU acceleration
- **LiteLLM routing**: Configuring the LiteLLM service to direct requests to the NVIDIA-specific backend

Both overlays preserve the service name `llama-server`, ensuring that dependent services like Open WebUI and the Dashboard API maintain their network references regardless of the underlying GPU hardware.

## How Docker Compose Merges the Layers

Docker Compose applies files sequentially, with later definitions overriding earlier ones at the service level. When executing:

```bash
docker compose \
  -f dream-server/docker-compose.base.yml \
  -f dream-server/docker-compose.amd.yml \
  up -d

```

The merge process:

1. Ingests the base service definitions, network configurations, and volume mappings
2. Overlays the AMD-specific `llama-server` definition, replacing the `image` field with the `build` directive and injecting ROCm-specific devices
3. Preserves all base-defined services (Open WebUI, Dashboard API) that are not mentioned in the overlay
4. Maintains consistent Docker network aliases so `llama-server` resolves correctly for internal service communication

This declarative approach eliminates conditional logic in the compose files; the hardware target is determined entirely by which overlay file appears last in the command sequence.

## Automating Extension Discovery with resolve-compose-stack.sh

For deployments requiring additional services (ComfyUI, Whisper, etc.), Dream Server includes [`scripts/resolve-compose-stack.sh`](https://github.com/Light-Heart-Labs/DreamServer/blob/main/scripts/resolve-compose-stack.sh). This helper script discovers extension-specific compose files located at `extensions/services/*/compose.yaml` and merges them with the base and GPU overlay configurations.

The resolution flow executes as follows:

1. **Base inclusion**: Loads [`docker-compose.base.yml`](https://github.com/Light-Heart-Labs/DreamServer/blob/main/docker-compose.base.yml) as the foundation
2. **GPU overlay selection**: Appends either [`docker-compose.amd.yml`](https://github.com/Light-Heart-Labs/DreamServer/blob/main/docker-compose.amd.yml) or [`docker-compose.nvidia.yml`](https://github.com/Light-Heart-Labs/DreamServer/blob/main/docker-compose.nvidia.yml)
3. **Extension discovery**: Scans `extensions/services/*/` directories for [`compose.yaml`](https://github.com/Light-Heart-Labs/DreamServer/blob/main/compose.yaml) files
4. **Stack synthesis**: Outputs a unified compose configuration to stdout for piping into `docker compose`

Because the script appends extension files after the GPU overlay, extension services can reference the already-configured `llama-server` with its hardware-specific backend active.

## Practical Usage Examples

Start a Dream Server instance with AMD GPU support:

```bash
docker compose \
  -f dream-server/docker-compose.base.yml \
  -f dream-server/docker-compose.amd.yml \
  up -d

```

Start with NVIDIA GPU acceleration:

```bash
docker compose \
  -f dream-server/docker-compose.base.yml \
  -f dream-server/docker-compose.nvidia.yml \
  up -d

```

Deploy the full stack including auto-discovered extensions:

```bash
scripts/resolve-compose-stack.sh \
  -f dream-server/docker-compose.base.yml \
  -f dream-server/docker-compose.amd.yml \
  | docker compose -f - up -d

```

## Summary

- **Dream Server** maintains a single [`docker-compose.base.yml`](https://github.com/Light-Heart-Labs/DreamServer/blob/main/docker-compose.base.yml) containing CPU-agnostic core services including the generic `llama-server` definition.
- **GPU overlays** ([`docker-compose.amd.yml`](https://github.com/Light-Heart-Labs/DreamServer/blob/main/docker-compose.amd.yml) and [`docker-compose.nvidia.yml`](https://github.com/Light-Heart-Labs/DreamServer/blob/main/docker-compose.nvidia.yml)) override the `llama-server` service to inject ROCm or CUDA-specific build instructions, devices, and environment variables.
- **Docker Compose merging** applies later files on top of earlier ones, allowing the overlay to replace specific fields without duplicating the entire service definition.
- **Service name consistency** ensures that Open WebUI, Dashboard API, and other dependencies always reference `llama-server` regardless of the hardware backend.
- **[`resolve-compose-stack.sh`](https://github.com/Light-Heart-Labs/DreamServer/blob/main/resolve-compose-stack.sh)** automates the inclusion of extension services from `extensions/services/*/compose.yaml`, appending them after the GPU overlay in the final stack configuration.

## Frequently Asked Questions

### Can I run Dream Server without a GPU using these overlay files?

Yes. Simply omit the GPU overlay files and run only the base compose file: `docker compose -f docker-compose.base.yml up -d`. The base configuration runs `llama-server` in CPU-only mode using the standard [`ghcr.io/ggml-org/llama.cpp`](https://github.com/Light-Heart-Labs/DreamServer/blob/main/ghcr.io/ggml-org/llama.cpp) image without requiring special device access or GPU drivers.

### Why does the `llama-server` service name remain identical across different overlays?

Maintaining the same service name ensures that downstream services defined in [`docker-compose.base.yml`](https://github.com/Light-Heart-Labs/DreamServer/blob/main/docker-compose.base.yml)—such as Open WebUI and the Dashboard API—can resolve the inference backend at the fixed hostname `llama-server` on the Docker network. This abstraction allows the same service configuration to work with AMD ROCm, NVIDIA CUDA, or CPU backends without modifying dependent service definitions.

### How do I add a custom extension to the Docker Compose stack?

Create a [`compose.yaml`](https://github.com/Light-Heart-Labs/DreamServer/blob/main/compose.yaml) file inside a new subdirectory under `extensions/services/` (e.g., [`extensions/services/my-custom-service/compose.yaml`](https://github.com/Light-Heart-Labs/DreamServer/blob/main/extensions/services/my-custom-service/compose.yaml)). The [`scripts/resolve-compose-stack.sh`](https://github.com/Light-Heart-Labs/DreamServer/blob/main/scripts/resolve-compose-stack.sh) utility automatically discovers this file during stack resolution and appends it after the GPU overlay, integrating your extension into the final configuration with access to the configured `llama-server` backend.

### What happens if I specify both AMD and NVIDIA overlays simultaneously?

Avoid specifying both overlays in the same command. Because Docker Compose applies files sequentially with later overrides taking precedence, the last-specified overlay would overwrite the previous one's `llama-server` definition, potentially resulting in conflicting device mappings (`/dev/dri` vs `/dev/nvidia*`) and environment variables. Deploy only one GPU overlay that matches your hardware configuration.