How to Use ODS with NVIDIA GPUs: Automatic Detection and Configuration Guide

ODS (Open Source AI Stack) automatically detects NVIDIA GPUs by querying nvidia-smi and configures Docker GPU passthrough by merging the ods/docker-compose.nvidia.yml overlay into your runtime stack.

ODS from the Osmantic/ODS repository eliminates manual GPU configuration by detecting supported NVIDIA hardware during installation and selecting optimized model catalogs based on your GPU tier. When you run the installer on Linux or Windows with WSL2, the stack identifies your GPU memory and compute capabilities, then applies the appropriate NVIDIA Container Toolkit settings to enable --gpus all passthrough for AI workloads.

Prerequisites for NVIDIA GPU Support

Before installing ODS, verify that your host meets the driver and runtime requirements for NVIDIA GPU passthrough.

  • NVIDIA Driver version 525 or newer must be installed on the host.
  • nvidia-smi must be available in your system PATH and return GPU details when executed.
  • Windows hosts require Docker Desktop with the WSL2 backend and the NVIDIA Container Toolkit installed in the WSL2 distribution.
  • Linux hosts need Docker Engine with the NVIDIA Container Toolkit configured at the daemon level.

How ODS Detects NVIDIA GPUs Automatically

The ODS installer relies on shell scripts in the ods/installers/lib/ directory to identify hardware and map capabilities to configuration tiers.

GPU Detection Logic

In ods/installers/lib/detection.sh, the installer executes nvidia-smi to query GPU presence and driver compatibility. When a supported NVIDIA GPU is detected, the script sets the internal environment variable ODS_GPU to nvidia. This variable controls which Docker Compose overlays are applied during stack generation.

Tier Mapping and Model Selection

After detection, ods/installers/lib/tier-map.sh categorizes your hardware into tiers such as NVIDIA Ultra or NVIDIA High based on available VRAM and CUDA cores. This tier determines which model catalogs ODS downloads and which inference engines (e.g., llama-server, comfyui) are configured for GPU acceleration. The tier mapping ensures that ODS deploys only models optimized for your specific NVIDIA hardware capabilities.

Installing ODS with NVIDIA GPU Support

Run the standard installation command to trigger automatic detection, or use environment variables to override default behaviors.

Standard Installation with Auto-Detection

Execute the installer from the repository root. The detection script will automatically identify your NVIDIA GPU and configure the stack:

./install.sh

Forcing NVIDIA GPU Mode

If automatic detection fails or you need to bypass the detection logic, explicitly set the ODS_GPU variable:

ODS_GPU=nvidia ./install.sh

Selecting Model Profiles

Use the MODEL_PROFILE variable to specify which model family ODS deploys. Set auto to let ODS select based on your GPU tier, or specify gemma4 for Google Gemma 4 models optimized for NVIDIA CUDA cores:

MODEL_PROFILE=auto ./install.sh
MODEL_PROFILE=gemma4 ./install.sh

Enabling Multi-Instance GPU (MIG) Support

For NVIDIA A100 GPUs, enable MIG support by setting the NVIDIA_MIG_ENABLED variable before installation:

export NVIDIA_MIG_ENABLED=1
./install.sh

How ODS Configures the Docker Stack

ODS generates the final Docker Compose configuration by merging base service definitions with GPU-specific overlays.

Compose Stack Resolution

The script scripts/resolve-compose-stack.sh handles the merging process. When ODS_GPU is set to nvidia, the resolver automatically includes ods/docker-compose.nvidia.yml in the final stack. This overlay adds the deploy.resources.reservations.devices configuration required for Docker to pass through all GPUs to containers running services like llama-server and comfyui.

Inspecting the Generated Configuration

After installation, verify that the NVIDIA overlay was applied correctly:

cat ods/docker-compose.yml | grep -i nvidia

You should see references to the NVIDIA runtime and device reservations. To rebuild the compose stack after changing GPU configurations (such as enabling MIG), run the installer with the --reset flag:

./install.sh --reset

Verifying GPU Passthrough at Runtime

Once the stack is running, confirm that containers have GPU access.

Check running services with Docker Compose:

docker compose ps

GPU-accelerated services will show status running and have been created with the --gpus all flag implicitly via the Compose file. Access the dashboard at http://localhost:3000 to verify that inference workloads are executing on the GPU rather than CPU fallback.

Summary

  • ODS automatically detects NVIDIA GPUs via ods/installers/lib/detection.sh, which queries nvidia-smi and sets ODS_GPU=nvidia.
  • The ods/installers/lib/tier-map.sh script maps your GPU to performance tiers (e.g., NVIDIA Ultra) and selects appropriate model catalogs.
  • The installer merges ods/docker-compose.nvidia.yml through scripts/resolve-compose-stack.sh to enable Docker GPU passthrough with --gpus all.
  • Use MODEL_PROFILE to control which AI models are deployed (e.g., gemma4 for NVIDIA-optimized variants).
  • Enable MIG on A100 GPUs by setting NVIDIA_MIG_ENABLED=1 before running install.sh.
  • Driver version 525 or newer is required, and Windows users must use Docker Desktop with the WSL2 backend.

Frequently Asked Questions

Does ODS support NVIDIA GPUs on Windows?

Yes, but you must use Docker Desktop with the WSL2 backend enabled. The installer detects the WSL2 environment and ensures the NVIDIA Container Toolkit is properly linked between Windows and the Linux subsystem. Native Windows containers without WSL2 are not supported for GPU passthrough in ODS.

How do I force ODS to use NVIDIA if the detection script fails?

Set the ODS_GPU environment variable to nvidia before executing the installer: ODS_GPU=nvidia ./install.sh. This bypasses the automatic detection in ods/installers/lib/detection.sh and forces the inclusion of the NVIDIA Docker Compose overlay regardless of nvidia-smi output.

What is the purpose of the MODEL_PROFILE environment variable?

MODEL_PROFILE controls which AI model catalogs ODS downloads and configures. Setting it to auto allows ods/installers/lib/tier-map.sh to select models based on your detected GPU tier (e.g., Gemma 4 for high-end NVIDIA cards), while specific values like gemma4 force that particular model family regardless of hardware detection.

Can I use ODS with multiple NVIDIA GPUs in a single host?

Yes. The ods/docker-compose.nvidia.yml overlay uses the --gpus all flag, which passes through all available NVIDIA GPUs to the containers. If you need to restrict specific services to specific GPUs, manually edit the generated ods/docker-compose.yml file to use specific device IDs instead of all before starting the stack.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →