# What Models Are Available for Each ODS Tier? Complete LLM Hardware Mapping

> Discover which LLM models map to each ODS tier from cloud APIs to local GGUF. Find the perfect quantized LLM for your hardware setup and explore the official tier mappings.

- Repository: [Osmantic/ODS](https://github.com/Osmantic/ODS)
- Tags: api-reference
- Published: 2026-09-02

---

**ODS assigns specific quantized LLMs to eleven distinct installation tiers ranging from cloud-based APIs to local GGUF models, with tier-to-model mappings defined in [`ods/installers/lib/tier-map.sh`](https://github.com/Osmantic/ODS/blob/main/ods/installers/lib/tier-map.sh) and validated by the test suite in [`ods/tests/test-tier-map.sh`](https://github.com/Osmantic/ODS/blob/main/ods/tests/test-tier-map.sh).**

The Osmantic/ODS framework automatically selects appropriate large language models based on detected hardware capabilities. Each tier maps to a specific **LLM_MODEL** identifier and **GGUF_URL**, ensuring optimal inference performance across devices from entry-level CPUs to high-end NVIDIA GPUs.

## Standard Numeric Tiers (0-4)

The foundational tier system covers cloud fallback and local Qwen-based deployments. These mappings are hardcoded in [`ods/installers/lib/tier-map.sh`](https://github.com/Osmantic/ODS/blob/main/ods/installers/lib/tier-map.sh) alongside the `MAX_CONTEXT` parameters for each configuration.

### Tier 0: Automatic Cloud Fallback

**Tier 0** serves as the automatic detection mode. It maps to `anthropic/claude-sonnet-4-5-20250514` and provides **no GGUF URL**, functioning as a cloud API fallback when local inference is unavailable or explicitly disabled.

### Tier 1: High-Performance Local (Qwen 3.5 9B)

Designed for capable consumer hardware, Tier 1 deploys the **Qwen 3.5 9B** model:

- **Model identifier**: `qwen3.5-9b`
- **GGUF URL**: `https://huggingface.co/unsloth/Qwen3.5-9B-GGUF/resolve/main/Qwen3.5-9B-Q4_K_M.gguf`

### Tier 2: Balanced Efficiency (Qwen 3.5 4B)

**Tier 2** optimizes for memory-constrained environments using the 4-billion parameter variant:

- **Model identifier**: `qwen3.5-4b`
- **GGUF URL**: `https://huggingface.co/unsloth/Qwen3.5-4B-GGUF/resolve/main/Qwen3.5-4B-Q4_K_M.gguf`

### Tier 3: Code-Optimized (Qwen 3 Coder Next)

Specialized for programming tasks, **Tier 3** utilizes the code-specific Qwen variant:

- **Model identifier**: `qwen3-coder-next`
- **GGUF URL**: `https://huggingface.co/unsloth/Qwen3-Coder-Next-GGUF/resolve/main/Qwen3-Coder-Next-Q4_K_M.gguf`

### Tier 4: Large Context Local (Qwen 3.6 35B A3B)

**Tier 4** supports advanced local inference with larger parameter counts:

- **Model identifier**: `qwen3.6-35b-a3b`
- **GGUF URL**: `https://huggingface.co/unsloth/Qwen3.6-35B-A3B-GGUF/resolve/main/Qwen3.6-35B-A3B-UD-Q4_K_M.gguf`

## Premium Hardware-Specific Tiers

Higher-performance tiers leverage Google's Gemma 4 architecture, mapped based on specific hardware profiles detected by [`ods/scripts/detect-hardware.sh`](https://github.com/Osmantic/ODS/blob/main/ods/scripts/detect-hardware.sh).

### NV ULTRA: NVIDIA 90GB+ Configuration

The **NV ULTRA** tier targets high-end NVIDIA GPUs with substantial VRAM:

- **Model**: `gemma-4-31b-it`
- **GGUF URL**: `https://huggingface.co/ggml-org/gemma-4-31B-it-GGUF/resolve/main/gemma-4-31B-it-Q4_K_M.gguf`

### SH LARGE: Apple Silicon Large

**SH LARGE** shares the same 31B Gemma model as NV ULTRA, optimized for high-end Apple Silicon:

- **Model**: `gemma-4-31b-it`
- **GGUF URL**: *Same as NV ULTRA*

### SH COMPACT: Apple Silicon Compact

For resource-constrained Apple devices, **SH COMPACT** uses the 26B parameter variant with activation compression:

- **Model**: `gemma-4-26b-a4b-it`
- **GGUF URL**: `https://huggingface.co/ggml-org/gemma-4-26B-A4B-it-GGUF/resolve/main/gemma-4-26B-A4B-it-Q4_K_M.gguf`

### ARC: Efficient Edge Deployment

The **ARC** tier utilizes the E4B (4-bit embedded) Gemma variant for efficient inference:

- **Model**: `gemma-4-e4b-it`
- **GGUF URL**: `https://huggingface.co/unsloth/gemma-4-E4B-it-GGUF/resolve/bfc15c382204943c3a8fff0c750b94ae2364d7a3/gemma-4-E4B-it-Q4_K_M.gguf`

### ARC LITE: Minimal Resource Mode

**ARC LITE** provides the lightest local option using the E2B variant:

- **Model**: `gemma-4-e2b-it`
- **GGUF URL**: `https://huggingface.co/unsloth/gemma-4-E2B-it-GGUF/resolve/0314792d7f1f7e229411f620751375812bb9faf2/gemma-4-E2B-it-Q4_K_M.gguf`

### CLOUD Mode

When operating in **CLOUD** mode, ODS uses the same model identifier as the locally selected tier, with the actual endpoint determined at runtime rather than via GGUF download.

## Querying Tier Mappings via CLI

The `ods` command-line interface provides direct access to the tier-map logic without inspecting source files.

To check the model assignment for a specific tier:

```bash
ods list-model --tier 2

```

To display the complete tier-to-model mapping table:

```bash
ods tier-info

```

These commands wrap the logic in [`ods/installers/lib/tier-map.sh`](https://github.com/Osmantic/ODS/blob/main/ods/installers/lib/tier-map.sh) and parse the same `LLM_MODEL` and `GGUF_URL` variables used during installation.

## Hardware Detection and Tier Assignment

The assignment process begins in [`ods/installers/phases/02-detection.sh`](https://github.com/Osmantic/ODS/blob/main/ods/installers/phases/02-detection.sh), which sets the `TIER` environment variable based on output from [`ods/scripts/detect-hardware.sh`](https://github.com/Osmantic/ODS/blob/main/ods/scripts/detect-hardware.sh). This detection script evaluates available GPU VRAM, Apple Silicon memory bandwidth, and CPU capabilities to select the appropriate tier from the mapping table.

The unit tests in [`ods/tests/test-tier-map.sh`](https://github.com/Osmantic/ODS/blob/main/ods/tests/test-tier-map.sh) verify that every tier resolves to an expected `LLM_MODEL` value and that valid `GGUF_URL` entries exist for downloadable models (excluding Tier 0 and CLOUD modes).

## Summary

- **Eleven distinct tiers** map hardware capabilities to specific LLMs, from cloud-based Claude to quantized Gemma and Qwen models.
- **Tier 0** provides API-based fallback without local GGUF files, while **Tiers 1-4** use progressively larger Qwen models.
- **Premium tiers** (NV ULTRA, SH LARGE, SH COMPACT, ARC, ARC LITE) deploy Gemma 4 architectures with parameters ranging from 2B to 31B.
- **Configuration files**: [`ods/installers/lib/tier-map.sh`](https://github.com/Osmantic/ODS/blob/main/ods/installers/lib/tier-map.sh) contains the core mappings, while [`ods/scripts/detect-hardware.sh`](https://github.com/Osmantic/ODS/blob/main/ods/scripts/detect-hardware.sh) drives automatic tier selection.
- **Validation**: The [`ods/tests/test-tier-map.sh`](https://github.com/Osmantic/ODS/blob/main/ods/tests/test-tier-map.sh) test suite ensures every tier resolves to valid model identifiers and download URLs.

## Frequently Asked Questions

### How does ODS automatically detect which tier to use?

The framework executes [`ods/scripts/detect-hardware.sh`](https://github.com/Osmantic/ODS/blob/main/ods/scripts/detect-hardware.sh) during the installation phase defined in [`ods/installers/phases/02-detection.sh`](https://github.com/Osmantic/ODS/blob/main/ods/installers/phases/02-detection.sh). This script analyzes GPU VRAM capacity, Apple Silicon memory configuration, and available compute resources to set the `TIER` environment variable, which the mapper in [`ods/installers/lib/tier-map.sh`](https://github.com/Osmantic/ODS/blob/main/ods/installers/lib/tier-map.sh) then translates into a specific model download.

### Why does Tier 0 not have a GGUF download URL?

**Tier 0** operates in "auto" mode and defaults to the cloud-based `anthropic/claude-sonnet-4-5-20250514` endpoint rather than local inference. Because it relies on API access rather than quantized local execution, it requires no GGUF file, distinguishing it from Tiers 1-4 and the premium hardware tiers that download specific Q4_K_M quantized models from HuggingFace repositories.

### What is the difference between the ARC and ARC LITE tiers?

**ARC** deploys the `gemma-4-e4b-it` model (4-bit embedded, 4B parameters), while **ARC LITE** uses the smaller `gemma-4-e2b-it` variant (2-bit embedded, 2B parameters). Both use specialized GGUF files from the Unsloth repository, but ARC LITE targets minimal resource environments with stricter memory constraints, whereas ARC provides a balance between efficiency and capability for edge deployments.

### Can I override the default model assigned to my hardware tier?

While [`ods/installers/lib/tier-map.sh`](https://github.com/Osmantic/ODS/blob/main/ods/installers/lib/tier-map.sh) defines default mappings, you can bypass automatic selection by explicitly setting the `TIER` variable or using the `--tier` flag with CLI commands like `ods list-model --tier 3`. However, modifying the underlying `LLM_MODEL` or `GGUF_URL` assignments requires editing the tier-map script directly, as the hardware detection logic in [`ods/installers/phases/02-detection.sh`](https://github.com/Osmantic/ODS/blob/main/ods/installers/phases/02-detection.sh) is designed to enforce validated configurations for stability.