# How to Use ODS with NVIDIA GPUs: Automatic Detection and Configuration Guide

> Learn how to use ODS with NVIDIA GPUs. ODS automatically detects and configures Docker GPU passthrough for seamless integration. Get started today.

- Repository: [Osmantic/ODS](https://github.com/Osmantic/ODS)
- Tags: how-to-guide
- Published: 2026-09-02

---

**ODS (Open Source AI Stack) automatically detects NVIDIA GPUs by querying `nvidia-smi` and configures Docker GPU passthrough by merging the [`ods/docker-compose.nvidia.yml`](https://github.com/Osmantic/ODS/blob/main/ods/docker-compose.nvidia.yml) overlay into your runtime stack.**

ODS from the Osmantic/ODS repository eliminates manual GPU configuration by detecting supported NVIDIA hardware during installation and selecting optimized model catalogs based on your GPU tier. When you run the installer on Linux or Windows with WSL2, the stack identifies your GPU memory and compute capabilities, then applies the appropriate NVIDIA Container Toolkit settings to enable `--gpus all` passthrough for AI workloads.

## Prerequisites for NVIDIA GPU Support

Before installing ODS, verify that your host meets the driver and runtime requirements for NVIDIA GPU passthrough.

- **NVIDIA Driver** version 525 or newer must be installed on the host.
- **`nvidia-smi`** must be available in your system PATH and return GPU details when executed.
- **Windows hosts** require Docker Desktop with the WSL2 backend and the NVIDIA Container Toolkit installed in the WSL2 distribution.
- **Linux hosts** need Docker Engine with the NVIDIA Container Toolkit configured at the daemon level.

## How ODS Detects NVIDIA GPUs Automatically

The ODS installer relies on shell scripts in the `ods/installers/lib/` directory to identify hardware and map capabilities to configuration tiers.

### GPU Detection Logic

In **[`ods/installers/lib/detection.sh`](https://github.com/Osmantic/ODS/blob/main/ods/installers/lib/detection.sh)**, the installer executes `nvidia-smi` to query GPU presence and driver compatibility. When a supported NVIDIA GPU is detected, the script sets the internal environment variable **`ODS_GPU`** to `nvidia`. This variable controls which Docker Compose overlays are applied during stack generation.

### Tier Mapping and Model Selection

After detection, **[`ods/installers/lib/tier-map.sh`](https://github.com/Osmantic/ODS/blob/main/ods/installers/lib/tier-map.sh)** categorizes your hardware into tiers such as *NVIDIA Ultra* or *NVIDIA High* based on available VRAM and CUDA cores. This tier determines which model catalogs ODS downloads and which inference engines (e.g., `llama-server`, `comfyui`) are configured for GPU acceleration. The tier mapping ensures that ODS deploys only models optimized for your specific NVIDIA hardware capabilities.

## Installing ODS with NVIDIA GPU Support

Run the standard installation command to trigger automatic detection, or use environment variables to override default behaviors.

### Standard Installation with Auto-Detection

Execute the installer from the repository root. The detection script will automatically identify your NVIDIA GPU and configure the stack:

```bash
./install.sh

```

### Forcing NVIDIA GPU Mode

If automatic detection fails or you need to bypass the detection logic, explicitly set the `ODS_GPU` variable:

```bash
ODS_GPU=nvidia ./install.sh

```

### Selecting Model Profiles

Use the **`MODEL_PROFILE`** variable to specify which model family ODS deploys. Set `auto` to let ODS select based on your GPU tier, or specify `gemma4` for Google Gemma 4 models optimized for NVIDIA CUDA cores:

```bash
MODEL_PROFILE=auto ./install.sh

```

```bash
MODEL_PROFILE=gemma4 ./install.sh

```

### Enabling Multi-Instance GPU (MIG) Support

For NVIDIA A100 GPUs, enable MIG support by setting the **`NVIDIA_MIG_ENABLED`** variable before installation:

```bash
export NVIDIA_MIG_ENABLED=1
./install.sh

```

## How ODS Configures the Docker Stack

ODS generates the final Docker Compose configuration by merging base service definitions with GPU-specific overlays.

### Compose Stack Resolution

The script **[`scripts/resolve-compose-stack.sh`](https://github.com/Osmantic/ODS/blob/main/scripts/resolve-compose-stack.sh)** handles the merging process. When `ODS_GPU` is set to `nvidia`, the resolver automatically includes **[`ods/docker-compose.nvidia.yml`](https://github.com/Osmantic/ODS/blob/main/ods/docker-compose.nvidia.yml)** in the final stack. This overlay adds the `deploy.resources.reservations.devices` configuration required for Docker to pass through all GPUs to containers running services like `llama-server` and `comfyui`.

### Inspecting the Generated Configuration

After installation, verify that the NVIDIA overlay was applied correctly:

```bash
cat ods/docker-compose.yml | grep -i nvidia

```

You should see references to the NVIDIA runtime and device reservations. To rebuild the compose stack after changing GPU configurations (such as enabling MIG), run the installer with the `--reset` flag:

```bash
./install.sh --reset

```

## Verifying GPU Passthrough at Runtime

Once the stack is running, confirm that containers have GPU access.

Check running services with Docker Compose:

```bash
docker compose ps

```

GPU-accelerated services will show status `running` and have been created with the `--gpus all` flag implicitly via the Compose file. Access the dashboard at `http://localhost:3000` to verify that inference workloads are executing on the GPU rather than CPU fallback.

## Summary

- ODS automatically detects NVIDIA GPUs via **[`ods/installers/lib/detection.sh`](https://github.com/Osmantic/ODS/blob/main/ods/installers/lib/detection.sh)**, which queries `nvidia-smi` and sets `ODS_GPU=nvidia`.
- The **[`ods/installers/lib/tier-map.sh`](https://github.com/Osmantic/ODS/blob/main/ods/installers/lib/tier-map.sh)** script maps your GPU to performance tiers (e.g., *NVIDIA Ultra*) and selects appropriate model catalogs.
- The installer merges **[`ods/docker-compose.nvidia.yml`](https://github.com/Osmantic/ODS/blob/main/ods/docker-compose.nvidia.yml)** through **[`scripts/resolve-compose-stack.sh`](https://github.com/Osmantic/ODS/blob/main/scripts/resolve-compose-stack.sh)** to enable Docker GPU passthrough with `--gpus all`.
- Use **`MODEL_PROFILE`** to control which AI models are deployed (e.g., `gemma4` for NVIDIA-optimized variants).
- Enable MIG on A100 GPUs by setting **`NVIDIA_MIG_ENABLED=1`** before running [`install.sh`](https://github.com/Osmantic/ODS/blob/main/install.sh).
- Driver version 525 or newer is required, and Windows users must use Docker Desktop with the WSL2 backend.

## Frequently Asked Questions

### Does ODS support NVIDIA GPUs on Windows?

Yes, but you must use Docker Desktop with the WSL2 backend enabled. The installer detects the WSL2 environment and ensures the NVIDIA Container Toolkit is properly linked between Windows and the Linux subsystem. Native Windows containers without WSL2 are not supported for GPU passthrough in ODS.

### How do I force ODS to use NVIDIA if the detection script fails?

Set the `ODS_GPU` environment variable to `nvidia` before executing the installer: `ODS_GPU=nvidia ./install.sh`. This bypasses the automatic detection in [`ods/installers/lib/detection.sh`](https://github.com/Osmantic/ODS/blob/main/ods/installers/lib/detection.sh) and forces the inclusion of the NVIDIA Docker Compose overlay regardless of `nvidia-smi` output.

### What is the purpose of the `MODEL_PROFILE` environment variable?

`MODEL_PROFILE` controls which AI model catalogs ODS downloads and configures. Setting it to `auto` allows [`ods/installers/lib/tier-map.sh`](https://github.com/Osmantic/ODS/blob/main/ods/installers/lib/tier-map.sh) to select models based on your detected GPU tier (e.g., Gemma 4 for high-end NVIDIA cards), while specific values like `gemma4` force that particular model family regardless of hardware detection.

### Can I use ODS with multiple NVIDIA GPUs in a single host?

Yes. The **[`ods/docker-compose.nvidia.yml`](https://github.com/Osmantic/ODS/blob/main/ods/docker-compose.nvidia.yml)** overlay uses the `--gpus all` flag, which passes through all available NVIDIA GPUs to the containers. If you need to restrict specific services to specific GPUs, manually edit the generated [`ods/docker-compose.yml`](https://github.com/Osmantic/ODS/blob/main/ods/docker-compose.yml) file to use specific device IDs instead of `all` before starting the stack.