How to Set Up Docker Compose with NVIDIA GPU Acceleration for Local Deep Research

To enable NVIDIA GPU acceleration, merge the base docker-compose.yml with docker-compose.gpu.override.yml using docker compose -f docker-compose.yml -f docker-compose.gpu.override.yml up -d after installing the NVIDIA Container Toolkit.

Local Deep Research (LDR) orchestrates three services—Ollama for LLM inference, SearXNG for search, and a web UI—through Docker Compose. While the base configuration runs Ollama on CPU, you can achieve significantly faster model inference by applying a GPU override file that reserves NVIDIA devices for the Ollama container.

Prerequisites

Before configuring the Docker Compose setup, ensure your host meets these requirements:

  • NVIDIA GPU with proprietary drivers installed
  • NVIDIA Container Toolkit (not the deprecated nvidia-docker2)
  • Docker Compose v2 or later

Install the toolkit on Ubuntu/Debian using the official NVIDIA repositories:

curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey \
  | sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg
curl -s -L https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list \
  | sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' \
  | sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list
sudo apt-get update
sudo apt-get install -y nvidia-container-toolkit
sudo systemctl restart docker
nvidia-smi  # Verify GPU visibility

Step-by-Step Docker Compose Configuration

Download the Compose Files

Retrieve the base configuration and the GPU override from the learningcircuit/local-deep-research repository:

curl -O https://raw.githubusercontent.com/LearningCircuit/local-deep-research/main/docker-compose.yml
curl -O https://raw.githubusercontent.com/LearningCircuit/local-deep-research/main/docker-compose.gpu.override.yml

The base docker-compose.yml defines CPU-only services for cross-platform compatibility, while docker-compose.gpu.override.yml adds the NVIDIA GPU reservation specifically for the Ollama service.

Launch the Stack with GPU Support

Merge the configurations by specifying both files with the -f flag. Docker Compose applies the override values on top of the base configuration:

docker compose -f docker-compose.yml -f docker-compose.gpu.override.yml up -d

This command starts all three containers, with the ollama service now accessing the host GPU. The override file injects a deploy.resources.reservations.devices block that requests the NVIDIA device driver, as seen in the source at [docker-compose.gpu.override.yml](https://github.com/LearningCircuit/local-deep-research/blob/main/docker-compose.gpu.override.yml).

Verify GPU Utilization

Confirm that Ollama is running on the GPU by checking the container logs or executing nvidia-smi inside the container:

docker logs local-deep-research-ollama-1
docker exec -it local-deep-research-ollama-1 nvidia-smi

If configured correctly, the output shows your GPU model and process information, indicating that model inference is hardware-accelerated rather than CPU-bound.

Understanding the GPU Override Structure

The docker-compose.gpu.override.yml file follows Docker Compose's override pattern. It redefines only the ollama service, adding the deploy block necessary for GPU access without duplicating the entire service definition from the base file:

services:
  ollama:
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: 1
              capabilities: [gpu]

This structure allows you to maintain a single CPU-compatible baseline while selectively enabling GPU acceleration on Linux hosts. The count: 1 reserves one GPU; adjust this value or replace it with device_ids if you need specific GPU targeting in multi-GPU systems.

Unraid-Specific Configuration

For Unraid users, combine the GPU override with the Unraid-specific volume mapping file:

docker compose -f docker-compose.yml \
               -f docker-compose.unraid.yml \
               -f docker-compose.gpu.override.yml up -d

The [docker-compose.unraid.yml](https://github.com/LearningCircuit/local-deep-research/blob/main/docker-compose.unraid.yml) file adjusts volume paths for Unraid's filesystem structure, while the GPU override adds the NVIDIA device reservation. Both overrides merge cleanly with the base configuration.

Summary

  • Base configuration: docker-compose.yml provides CPU-only, cross-platform compatibility
  • GPU acceleration: docker-compose.gpu.override.yml adds NVIDIA device reservations to the Ollama service
  • Merge command: Use docker compose -f docker-compose.yml -f docker-compose.gpu.override.yml up -d to combine configurations
  • Requirement: NVIDIA Container Toolkit must be installed on the Linux host
  • Verification: Run nvidia-smi inside the Ollama container to confirm GPU access

Frequently Asked Questions

Why does the setup require two separate Compose files?

Separating the GPU configuration allows the project to maintain a single cross-platform baseline. The base docker-compose.yml works on any system including macOS and CPU-only Linux hosts, while the override file adds Linux-specific NVIDIA device reservations without breaking compatibility for other users.

Can I use GPU acceleration on Windows or macOS?

No. NVIDIA Container Toolkit requires Linux host support for Docker containers. Windows users with WSL2 may achieve GPU passthrough in some configurations, but the official Local Deep Research GPU setup targets Linux hosts exclusively according to the repository documentation.

How do I run the stack on CPU-only mode again?

Simply omit the override file and start with the base configuration only: docker compose -f docker-compose.yml up -d. This launches Ollama in CPU inference mode, which works on any platform but offers significantly slower performance compared to GPU acceleration.

What if I have multiple NVIDIA GPUs?

Modify the count: 1 value in docker-compose.gpu.override.yml to reserve additional GPUs, or specify exact device_ids (e.g., device_ids: ['0', '1']) to target specific cards. Ensure the Ollama container has sufficient VRAM for your selected model—the llama3.1:8b model requires approximately 6GB, while larger models like gemma3:12b need correspondingly more memory.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →