Understanding the ODS Hardware Tier Mapping System: Automatic LLM Selection for GPU and CPU Classes
The ODS hardware tier mapping system automatically translates detected hardware capabilities into specific LLM configurations by mapping symbolic tier identifiers like NV_ULTRA or ARC to concrete model artifacts, download URLs, and runtime parameters through a declarative shell library.
The Osmantic/ODS installer uses a sophisticated hardware tier mapping system to bridge the gap between raw hardware detection and optimal LLM deployment. This system, implemented primarily in ods/installers/lib/tier-map.sh, eliminates manual configuration by deterministically selecting quantized models that match your machine's memory constraints and compute class. Whether you are running on Intel Arc GPUs, NVIDIA Ultra cards, or CPU-only environments, the mapping system ensures the correct GGUF model is downloaded, verified, and configured with appropriate context windows and GPU backend settings.
How the Hardware Tier Mapping System Works
The ODS hardware tier mapping system operates as a deterministic translation layer between hardware detection and runtime configuration. It converts symbolic tier identifiers into concrete deployment parameters, ensuring that each machine class receives an appropriately sized quantized model with verified checksums and optimized runtime settings.
Hardware Detection and Tier Assignment
The mapping process begins during the detection phase in ods/installers/phases/02-detection.sh. This phase probes the host machine's GPU and CPU capabilities, then sets the environment variable TIER to a symbolic value representing the detected hardware class. Valid tier identifiers include discrete GPU classes like ARC (Intel Arc), NV_ULTRA (NVIDIA Ultra 90GB+), and SH_LARGE, as well as numeric fallback tiers like 0 or 1 for CPU-centric or minimal configurations.
The Core Mapping Logic in tier-map.sh
At the heart of the system lies ods/installers/lib/tier-map.sh, which exports the resolve_tier_config() function. This pure-function script translates the raw TIER variable into a comprehensive set of runtime parameters without side effects, making it safe to source and invoke multiple times during the installation process.
The function populates the following key variables:
| Variable | Purpose |
|---|---|
| TIER_NAME | Human-readable description (e.g., "Intel Arc", "NVIDIA Ultra (90GB+)") |
| LLM_MODEL | Model identifier (e.g., qwen3.5-9b, gemma-4-31b-it) |
| GGUF_FILE | Filename of the quantized GGUF artifact |
| GGUF_URL | Direct download URL from Hugging Face |
| GGUF_SHA256 | Checksum for integrity verification |
| MAX_CONTEXT | Maximum token context window |
| LLM_MODEL_SIZE_MB | Approximate download size for UI hints |
| GPU_BACKEND | Optional GPU acceleration backend specification |
| N_GPU_LAYERS | Number of layers to offload to GPU (SYCL/CUDA) |
The Two-Stage Resolution Process
The resolve_tier_config() function implements a two-stage resolution pipeline that separates model family selection from tier-specific parameter assignment.
Stage 1: Model Profile Resolution
Before selecting specific model weights, the system determines the effective model profile through normalize_model_profile() and effective_model_profile(). The installer respects the MODEL_PROFILE environment variable, which can be set to qwen, gemma4, or auto.
When the profile is set to auto, the system applies a platform-aware default: Cloud tiers default to the Qwen profile, while all other hardware tiers (including Intel Arc and NVIDIA variants) default to the Gemma-4 profile. This logic ensures that cloud environments receive models optimized for server-side deployment, while local hardware receives models tuned for consumer and workstation GPUs.
Stage 2: Tier-Specific Configuration
Once the effective profile is determined, resolve_tier_config() delegates to specialized configuration functions:
set_qwen_tier_config()populates parameters for the Qwen model family, used as the default for most GPU tiers when the profile explicitly requests Qwen or when running on Cloud tiers with auto-detection.set_gemma4_tier_config()configures the Gemma-4 model family, serving as the default for local GPU tiers when the profile is set togemma4orauto.
After setting the model-specific parameters, the system calls configure_llama_runtime_defaults() to adjust Docker and llama.cpp runtime configurations. This includes selecting newer llama.cpp container images for Gemma-4 models or setting appropriate GPU backend flags for SYCL and CUDA acceleration.
Reverse Tier Lookup for UI Operations
Beyond the standard configuration flow, the mapping system supports reverse resolution via the tier_to_model() function. This utility translates a tier identifier back to its corresponding model name without modifying global state, enabling "model swap" operations in the ODS UI and allowing administrators to query which model would deploy on hypothetical hardware tiers.
Practical Implementation Examples
The following examples demonstrate how to interact with the ODS hardware tier mapping system in shell scripts:
# Resolve the complete tier configuration for the current machine
source ods/installers/lib/tier-map.sh
# TIER is set by the detection phase; e.g., TIER=NV_ULTRA
resolve_tier_config
# Access the resolved variables
echo "$TIER_NAME" # Output: NVIDIA Ultra (90GB+)
echo "$LLM_MODEL" # Output: qwen3-coder-next
echo "$GGUF_URL" # Output: https://huggingface.co/unsloth/Qwen3-Coder-Next-GGUF/...
echo "$MAX_CONTEXT" # Output: 131072
# Query the model name for a specific tier without affecting global variables
source ods/installers/lib/tier-map.sh
model=$(tier_to_model NV_ULTRA)
echo "$model" # Output: qwen3-coder-next (or gemma-4-31b-it if profile=gemma4)
# Force the automatic model profile selection
export MODEL_PROFILE=auto
source ods/installers/lib/tier-map.sh
resolve_tier_config
# For non-cloud GPU tiers, this selects the Gemma-4 profile
# For Cloud tiers, this selects the Qwen profile
Summary
- The ODS hardware tier mapping system lives in
ods/installers/lib/tier-map.shand translates symbolic tier identifiers into concrete LLM deployment parameters. - Hardware detection occurs in
ods/installers/phases/02-detection.sh, which sets theTIERenvironment variable to values likeARC,NV_ULTRA, orSH_LARGE. - The
resolve_tier_config()function provides deterministic mapping to model identifiers, GGUF download URLs, SHA256 checksums, and context window sizes. - Model profile selection supports
qwen,gemma4, andautomodes, with Cloud tiers defaulting to Qwen and local hardware defaulting to Gemma-4. - The
tier_to_model()function enables reverse lookup capabilities for UI-driven model swap operations.
Frequently Asked Questions
How does ODS determine which hardware tier to assign to my machine?
The ODS installer executes a detection phase via ods/installers/phases/02-detection.sh, which probes your system's GPU and CPU capabilities using helper functions from ods/installers/lib/detection.sh. Based on available VRAM, GPU vendor (Intel, NVIDIA), and compute class, it assigns a symbolic tier identifier such as ARC for Intel Arc GPUs or NV_ULTRA for high-memory NVIDIA cards.
Can I override the automatic model selection in the ODS hardware tier mapping system?
Yes, you can override the default selection by setting the MODEL_PROFILE environment variable to qwen or gemma4 before running the installer. When set to auto or left unspecified, the system automatically selects Gemma-4 for local GPU tiers and Qwen for Cloud tiers according to the logic in effective_model_profile().
What is the difference between the Qwen and Gemma-4 model profiles in ODS?
The Qwen profile configures models from the Qwen family (such as qwen3.5-9b), which the system defaults to for Cloud tiers. The Gemma-4 profile configures Google's Gemma-4 models (such as gemma-4-31b-it), which serve as the default for local GPU tiers like Intel Arc and NVIDIA Ultra. Each profile triggers distinct GGUF_URL values and may invoke different runtime defaults via configure_llama_runtime_defaults().
How does the tier mapping system ensure model integrity during download?
The resolve_tier_config() function populates the GGUF_SHA256 variable with a cryptographic checksum for each tier-specific model. When the ODS bootstrap process downloads the GGUF artifact specified in GGUF_URL, it verifies the file against this SHA256 hash before proceeding with deployment, ensuring the quantized model has not been corrupted or tampered with during transfer.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →