How to Use ODS with NVIDIA GPUs: Automatic Detection and Configuration Guide
ODS (Open Source AI Stack) automatically detects NVIDIA GPUs by querying nvidia-smi and configures Docker GPU passthrough by merging the ods/docker-compose.nvidia.yml overlay into your runtime stack.
ODS from the Osmantic/ODS repository eliminates manual GPU configuration by detecting supported NVIDIA hardware during installation and selecting optimized model catalogs based on your GPU tier. When you run the installer on Linux or Windows with WSL2, the stack identifies your GPU memory and compute capabilities, then applies the appropriate NVIDIA Container Toolkit settings to enable --gpus all passthrough for AI workloads.
Prerequisites for NVIDIA GPU Support
Before installing ODS, verify that your host meets the driver and runtime requirements for NVIDIA GPU passthrough.
- NVIDIA Driver version 525 or newer must be installed on the host.
nvidia-smimust be available in your system PATH and return GPU details when executed.- Windows hosts require Docker Desktop with the WSL2 backend and the NVIDIA Container Toolkit installed in the WSL2 distribution.
- Linux hosts need Docker Engine with the NVIDIA Container Toolkit configured at the daemon level.
How ODS Detects NVIDIA GPUs Automatically
The ODS installer relies on shell scripts in the ods/installers/lib/ directory to identify hardware and map capabilities to configuration tiers.
GPU Detection Logic
In ods/installers/lib/detection.sh, the installer executes nvidia-smi to query GPU presence and driver compatibility. When a supported NVIDIA GPU is detected, the script sets the internal environment variable ODS_GPU to nvidia. This variable controls which Docker Compose overlays are applied during stack generation.
Tier Mapping and Model Selection
After detection, ods/installers/lib/tier-map.sh categorizes your hardware into tiers such as NVIDIA Ultra or NVIDIA High based on available VRAM and CUDA cores. This tier determines which model catalogs ODS downloads and which inference engines (e.g., llama-server, comfyui) are configured for GPU acceleration. The tier mapping ensures that ODS deploys only models optimized for your specific NVIDIA hardware capabilities.
Installing ODS with NVIDIA GPU Support
Run the standard installation command to trigger automatic detection, or use environment variables to override default behaviors.
Standard Installation with Auto-Detection
Execute the installer from the repository root. The detection script will automatically identify your NVIDIA GPU and configure the stack:
./install.sh
Forcing NVIDIA GPU Mode
If automatic detection fails or you need to bypass the detection logic, explicitly set the ODS_GPU variable:
ODS_GPU=nvidia ./install.sh
Selecting Model Profiles
Use the MODEL_PROFILE variable to specify which model family ODS deploys. Set auto to let ODS select based on your GPU tier, or specify gemma4 for Google Gemma 4 models optimized for NVIDIA CUDA cores:
MODEL_PROFILE=auto ./install.sh
MODEL_PROFILE=gemma4 ./install.sh
Enabling Multi-Instance GPU (MIG) Support
For NVIDIA A100 GPUs, enable MIG support by setting the NVIDIA_MIG_ENABLED variable before installation:
export NVIDIA_MIG_ENABLED=1
./install.sh
How ODS Configures the Docker Stack
ODS generates the final Docker Compose configuration by merging base service definitions with GPU-specific overlays.
Compose Stack Resolution
The script scripts/resolve-compose-stack.sh handles the merging process. When ODS_GPU is set to nvidia, the resolver automatically includes ods/docker-compose.nvidia.yml in the final stack. This overlay adds the deploy.resources.reservations.devices configuration required for Docker to pass through all GPUs to containers running services like llama-server and comfyui.
Inspecting the Generated Configuration
After installation, verify that the NVIDIA overlay was applied correctly:
cat ods/docker-compose.yml | grep -i nvidia
You should see references to the NVIDIA runtime and device reservations. To rebuild the compose stack after changing GPU configurations (such as enabling MIG), run the installer with the --reset flag:
./install.sh --reset
Verifying GPU Passthrough at Runtime
Once the stack is running, confirm that containers have GPU access.
Check running services with Docker Compose:
docker compose ps
GPU-accelerated services will show status running and have been created with the --gpus all flag implicitly via the Compose file. Access the dashboard at http://localhost:3000 to verify that inference workloads are executing on the GPU rather than CPU fallback.
Summary
- ODS automatically detects NVIDIA GPUs via
ods/installers/lib/detection.sh, which queriesnvidia-smiand setsODS_GPU=nvidia. - The
ods/installers/lib/tier-map.shscript maps your GPU to performance tiers (e.g., NVIDIA Ultra) and selects appropriate model catalogs. - The installer merges
ods/docker-compose.nvidia.ymlthroughscripts/resolve-compose-stack.shto enable Docker GPU passthrough with--gpus all. - Use
MODEL_PROFILEto control which AI models are deployed (e.g.,gemma4for NVIDIA-optimized variants). - Enable MIG on A100 GPUs by setting
NVIDIA_MIG_ENABLED=1before runninginstall.sh. - Driver version 525 or newer is required, and Windows users must use Docker Desktop with the WSL2 backend.
Frequently Asked Questions
Does ODS support NVIDIA GPUs on Windows?
Yes, but you must use Docker Desktop with the WSL2 backend enabled. The installer detects the WSL2 environment and ensures the NVIDIA Container Toolkit is properly linked between Windows and the Linux subsystem. Native Windows containers without WSL2 are not supported for GPU passthrough in ODS.
How do I force ODS to use NVIDIA if the detection script fails?
Set the ODS_GPU environment variable to nvidia before executing the installer: ODS_GPU=nvidia ./install.sh. This bypasses the automatic detection in ods/installers/lib/detection.sh and forces the inclusion of the NVIDIA Docker Compose overlay regardless of nvidia-smi output.
What is the purpose of the MODEL_PROFILE environment variable?
MODEL_PROFILE controls which AI model catalogs ODS downloads and configures. Setting it to auto allows ods/installers/lib/tier-map.sh to select models based on your detected GPU tier (e.g., Gemma 4 for high-end NVIDIA cards), while specific values like gemma4 force that particular model family regardless of hardware detection.
Can I use ODS with multiple NVIDIA GPUs in a single host?
Yes. The ods/docker-compose.nvidia.yml overlay uses the --gpus all flag, which passes through all available NVIDIA GPUs to the containers. If you need to restrict specific services to specific GPUs, manually edit the generated ods/docker-compose.yml file to use specific device IDs instead of all before starting the stack.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →