ODS Hardware Support for AI Deployment: Complete GPU and CPU Compatibility Guide
ODS supports AI deployment on NVIDIA, AMD, Intel Arc, and Apple Silicon across Linux, Windows, and macOS, automatically detecting available compute resources to select optimal quantized models without manual configuration.
The Osmantic/ODS repository provides a turnkey AI deployment stack designed for heterogeneous hardware environments. Its installer eliminates configuration guesswork by mapping detected GPUs and CPUs to pre-defined performance tiers, enabling single-command deployment across consumer desktops, workstations, and specialized AI accelerators.
Cross-Platform Hardware Coverage
ODS targets the three major operating systems with vendor-specific optimizations for each compute architecture.
Linux: NVIDIA, AMD, and Intel Arc
Linux hosts receive the broadest hardware support. The installer detects NVIDIA discrete GPUs (via CUDA), AMD GPUs (via ROCm), and Intel Arc graphics cards (via SYCL), falling back to CPU-only inference when no compatible accelerator is present. This covers data center cards like the A100 and H100 down to consumer RTX 4060s and Arc A380s.
Windows: NVIDIA and AMD
Windows support focuses on the two dominant gaming and workstation GPU vendors. The PowerShell-based installer (installers/windows/lib/detection.ps1) queries dxdiag and WMI to identify NVIDIA and AMD graphics adapters. Intel Arc is not currently supported on Windows in this release.
macOS: Apple Silicon M-Series
Apple Silicon Macs utilize the Metal backend for unified memory inference. ODS supports the entire M-series lineup—from base M1 chips with 8GB of shared memory to M4 Max and M2 Ultra configurations with 64GB+—leveraging the unified memory architecture to run larger models than discrete VRAM tiers of equivalent capacity would allow.
GPU-Tiered Model Selection System
The core of ODS hardware support for AI deployment is its automatic tier mapping. The installer categorizes detected hardware into numbered tiers based on available VRAM (discrete cards) or unified memory (Apple Silicon/AMD Strix Halo), then selects a default GGUF model from config/model-library.json that fits within the memory envelope.
NVIDIA Discrete VRAM Tiers
For NVIDIA GPUs, the system distinguishes between consumer workstations and high-end data center deployments:
- Tier 0 (8GB CPU Fallback): Runs Qwen3.5 2B (Q4_K_M) with 8K context on CPU-only systems or low-VRAM configurations.
- Tier 1 (8GB VRAM): Deploys Qwen3.5 9B (Q4_K_M) with 32K context—optimal for RTX 4060 and RTX 3060 12GB cards.
- Tier 2 (12GB VRAM): Selects Phi-4 14B (Q4_K_M) with 16K context, targeting RTX 4070-class hardware.
- Tier 3 (24GB VRAM): Maps to Qwen3.5 27B (Q4_K_M) with 32K context, ideal for RTX 4090 and RTX A6000 users.
- Tier 4 (48GB VRAM): Runs DeepSeek R1 Distill Llama 70B (Q4_K_M) with 32K context on professional cards like the A6000 Ada and L40S.
- NV_ULTRA (90+GB Discrete): For multi-GPU A100/H100 setups, selects Qwen3 Coder Next (Q4_K_M) with 128K context.
- NV_ULTRA (90+GB ARM64): On DGX Spark or GB10-class hosts, uses Qwen3.6 35B-A3B (UD-Q4_K_M) with 128K context.
AMD Strix Halo Unified Memory
AMD's Strix Halo APUs with high-bandwidth unified memory receive specialized tiers:
- SH_COMPACT (64GB): Runs Qwen3.6 35B-A3B (UD-Q4_K_M) with 128K context on Ryzen AI MAX+ 395 (64GB configurations).
- SH_LARGE (96GB): Deploys DeepSeek R1 Distill Llama 70B (Q4_K_M) with 32K context on 96GB Ryzen AI MAX+ 395 systems.
- SH_LARGE (124GB): Utilizes Qwen3.6 35B-A3B (UD-Q4_K_M) with 128K context on 128GB Ryzen AI MAX+ 395 configurations.
Apple Silicon Metal Tiers
Apple's unified memory architecture allows larger parameter models at lower tier numbers compared to discrete VRAM:
- Tier 0 (8GB): Runs Phi-4 Mini (Q4_K_M) with 128K context on base M1/M2 Macs.
- Tier 1 (16GB): Selects Qwen3.5 9B (Q4_K_M) with 32K context on M4 Mac Mini (16GB).
- Tier 2 (32GB): Deploys Phi-4 14B (Q4_K_M) with 16K context on M4 Pro Mac Mini or M3 Max MacBook Pro.
- Tier 3 (48GB): Uses Qwen3.5 27B (Q4_K_M) with 32K context on M4 Pro (48GB) or M2 Max (48GB) machines.
- Tier 4 (64GB+): Maps to Qwen3.6 35B-A3B (UD-Q4_K_M) with 128K context for M2 Ultra Mac Studio or M4 Max (64GB+) systems.
Intel Arc on Linux (SYCL)
Intel Arc GPU support is implemented via the SYCL backend on Linux hosts:
- ARC_LITE (6GB): Runs Phi-4 Mini (Q4_K_M) with 128K context on Arc A380 hardware.
- ARC_LITE (8GB): Selects Qwen3.5 9B (Q4_K_M) with 32K context for Arc A750 GPUs.
- ARC (16GB): Deploys Phi-4 14B (Q4_K_M) with 16K context on Arc A770 16GB and newer Arc GPUs.
Automatic Hardware Detection Implementation
The hardware detection pipeline begins in the installers/lib/detection.sh script on Linux/macOS and installers/windows/lib/detection.ps1 on Windows. These scripts query system APIs—lshw, PCI IDs, dxdiag, and Metal Performance Shaders—to identify the GPU vendor, model, and available memory.
Once detected, installers/lib/tier-map.sh translates the hardware signature into a numeric tier. This tier value indexes into config/model-library.json, which contains the canonical mapping of tier numbers to specific GGUF file URLs, quantization formats (Q4_K_M, UD-Q4_K_M), and context window configurations. The selected values are written to the .env file as LLM_MODEL, GGUF_FILE, and MAX_CONTEXT before the stack initializes.
Runtime CLI and Manual Overrides
While automatic detection handles standard deployments, ODS provides CLI utilities to inspect and override hardware mappings.
View current hardware detection and assigned tier:
ods status
List all available tiers for your detected backend:
ods model list
Force a specific tier during installation (bypassing auto-detection):
MODEL_TIER=3 ./install.sh
Override the model family while retaining the auto-detected tier:
MODEL_PROFILE=gemma4 ./install.sh
These environment variables allow testing of higher-context models on capable hardware or forcing conservative models on shared systems.
Summary
- ODS hardware support for AI deployment spans Linux (NVIDIA/AMD/Intel), Windows (NVIDIA/AMD), and macOS (Apple Silicon), with automatic GPU detection via
installers/lib/detection.shand PowerShell equivalents. - The tier-mapping system (
installers/lib/tier-map.sh) assigns hardware to categories 0-4 (and ULTRA variants) based on VRAM or unified memory capacity. - Default GGUF models are automatically selected from
config/model-library.jsonto fit within detected memory constraints, ranging from 2B parameter models on 8GB systems to 70B parameter models on 48GB+ workstations. - Users can override automatic selection using
MODEL_TIERorMODEL_PROFILEenvironment variables at install time.
Frequently Asked Questions
Does ODS support Intel Arc GPUs on Windows?
No. According to the current source code in installers/windows/lib/detection.ps1, Intel Arc is only supported on Linux via the SYCL backend. Windows installations are limited to NVIDIA and AMD graphics adapters.
What is the minimum hardware requirement to run ODS?
The minimum supported configuration is Tier 0, requiring 8GB of system RAM (CPU-only) or 8GB of VRAM. This tier runs the Qwen3.5 2B model quantized to Q4_K_M with an 8,000-token context window, suitable for basic inference tasks on low-end hardware.
Can I deploy ODS on a machine without any GPU?
Yes. The installer automatically falls back to CPU inference when no compatible GPU is detected, categorizing the system as Tier 0. While performance is significantly slower than GPU acceleration, this allows deployment on servers and VMs without discrete graphics cards.
How do I verify which hardware tier the installer selected for my system?
Run the ods status command after installation. This displays the detected GPU vendor/model and the assigned tier number (e.g., "GPU: NVIDIA RTX 4090 (Tier 3)"). For a preview before installing, examine the output of the detection scripts directly: ./installers/lib/detection.sh on Linux/macOS or .\installers\windows\lib\detection.ps1 on Windows.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →