What Models Are Available for Each ODS Tier? Complete LLM Hardware Mapping

ODS assigns specific quantized LLMs to eleven distinct installation tiers ranging from cloud-based APIs to local GGUF models, with tier-to-model mappings defined in ods/installers/lib/tier-map.sh and validated by the test suite in ods/tests/test-tier-map.sh.

The Osmantic/ODS framework automatically selects appropriate large language models based on detected hardware capabilities. Each tier maps to a specific LLM_MODEL identifier and GGUF_URL, ensuring optimal inference performance across devices from entry-level CPUs to high-end NVIDIA GPUs.

Standard Numeric Tiers (0-4)

The foundational tier system covers cloud fallback and local Qwen-based deployments. These mappings are hardcoded in ods/installers/lib/tier-map.sh alongside the MAX_CONTEXT parameters for each configuration.

Tier 0: Automatic Cloud Fallback

Tier 0 serves as the automatic detection mode. It maps to anthropic/claude-sonnet-4-5-20250514 and provides no GGUF URL, functioning as a cloud API fallback when local inference is unavailable or explicitly disabled.

Tier 1: High-Performance Local (Qwen 3.5 9B)

Designed for capable consumer hardware, Tier 1 deploys the Qwen 3.5 9B model:

  • Model identifier: qwen3.5-9b
  • GGUF URL: https://huggingface.co/unsloth/Qwen3.5-9B-GGUF/resolve/main/Qwen3.5-9B-Q4_K_M.gguf

Tier 2: Balanced Efficiency (Qwen 3.5 4B)

Tier 2 optimizes for memory-constrained environments using the 4-billion parameter variant:

  • Model identifier: qwen3.5-4b
  • GGUF URL: https://huggingface.co/unsloth/Qwen3.5-4B-GGUF/resolve/main/Qwen3.5-4B-Q4_K_M.gguf

Tier 3: Code-Optimized (Qwen 3 Coder Next)

Specialized for programming tasks, Tier 3 utilizes the code-specific Qwen variant:

  • Model identifier: qwen3-coder-next
  • GGUF URL: https://huggingface.co/unsloth/Qwen3-Coder-Next-GGUF/resolve/main/Qwen3-Coder-Next-Q4_K_M.gguf

Tier 4: Large Context Local (Qwen 3.6 35B A3B)

Tier 4 supports advanced local inference with larger parameter counts:

  • Model identifier: qwen3.6-35b-a3b
  • GGUF URL: https://huggingface.co/unsloth/Qwen3.6-35B-A3B-GGUF/resolve/main/Qwen3.6-35B-A3B-UD-Q4_K_M.gguf

Premium Hardware-Specific Tiers

Higher-performance tiers leverage Google's Gemma 4 architecture, mapped based on specific hardware profiles detected by ods/scripts/detect-hardware.sh.

NV ULTRA: NVIDIA 90GB+ Configuration

The NV ULTRA tier targets high-end NVIDIA GPUs with substantial VRAM:

  • Model: gemma-4-31b-it
  • GGUF URL: https://huggingface.co/ggml-org/gemma-4-31B-it-GGUF/resolve/main/gemma-4-31B-it-Q4_K_M.gguf

SH LARGE: Apple Silicon Large

SH LARGE shares the same 31B Gemma model as NV ULTRA, optimized for high-end Apple Silicon:

  • Model: gemma-4-31b-it
  • GGUF URL: Same as NV ULTRA

SH COMPACT: Apple Silicon Compact

For resource-constrained Apple devices, SH COMPACT uses the 26B parameter variant with activation compression:

  • Model: gemma-4-26b-a4b-it
  • GGUF URL: https://huggingface.co/ggml-org/gemma-4-26B-A4B-it-GGUF/resolve/main/gemma-4-26B-A4B-it-Q4_K_M.gguf

ARC: Efficient Edge Deployment

The ARC tier utilizes the E4B (4-bit embedded) Gemma variant for efficient inference:

  • Model: gemma-4-e4b-it
  • GGUF URL: https://huggingface.co/unsloth/gemma-4-E4B-it-GGUF/resolve/bfc15c382204943c3a8fff0c750b94ae2364d7a3/gemma-4-E4B-it-Q4_K_M.gguf

ARC LITE: Minimal Resource Mode

ARC LITE provides the lightest local option using the E2B variant:

  • Model: gemma-4-e2b-it
  • GGUF URL: https://huggingface.co/unsloth/gemma-4-E2B-it-GGUF/resolve/0314792d7f1f7e229411f620751375812bb9faf2/gemma-4-E2B-it-Q4_K_M.gguf

CLOUD Mode

When operating in CLOUD mode, ODS uses the same model identifier as the locally selected tier, with the actual endpoint determined at runtime rather than via GGUF download.

Querying Tier Mappings via CLI

The ods command-line interface provides direct access to the tier-map logic without inspecting source files.

To check the model assignment for a specific tier:

ods list-model --tier 2

To display the complete tier-to-model mapping table:

ods tier-info

These commands wrap the logic in ods/installers/lib/tier-map.sh and parse the same LLM_MODEL and GGUF_URL variables used during installation.

Hardware Detection and Tier Assignment

The assignment process begins in ods/installers/phases/02-detection.sh, which sets the TIER environment variable based on output from ods/scripts/detect-hardware.sh. This detection script evaluates available GPU VRAM, Apple Silicon memory bandwidth, and CPU capabilities to select the appropriate tier from the mapping table.

The unit tests in ods/tests/test-tier-map.sh verify that every tier resolves to an expected LLM_MODEL value and that valid GGUF_URL entries exist for downloadable models (excluding Tier 0 and CLOUD modes).

Summary

  • Eleven distinct tiers map hardware capabilities to specific LLMs, from cloud-based Claude to quantized Gemma and Qwen models.
  • Tier 0 provides API-based fallback without local GGUF files, while Tiers 1-4 use progressively larger Qwen models.
  • Premium tiers (NV ULTRA, SH LARGE, SH COMPACT, ARC, ARC LITE) deploy Gemma 4 architectures with parameters ranging from 2B to 31B.
  • Configuration files: ods/installers/lib/tier-map.sh contains the core mappings, while ods/scripts/detect-hardware.sh drives automatic tier selection.
  • Validation: The ods/tests/test-tier-map.sh test suite ensures every tier resolves to valid model identifiers and download URLs.

Frequently Asked Questions

How does ODS automatically detect which tier to use?

The framework executes ods/scripts/detect-hardware.sh during the installation phase defined in ods/installers/phases/02-detection.sh. This script analyzes GPU VRAM capacity, Apple Silicon memory configuration, and available compute resources to set the TIER environment variable, which the mapper in ods/installers/lib/tier-map.sh then translates into a specific model download.

Why does Tier 0 not have a GGUF download URL?

Tier 0 operates in "auto" mode and defaults to the cloud-based anthropic/claude-sonnet-4-5-20250514 endpoint rather than local inference. Because it relies on API access rather than quantized local execution, it requires no GGUF file, distinguishing it from Tiers 1-4 and the premium hardware tiers that download specific Q4_K_M quantized models from HuggingFace repositories.

What is the difference between the ARC and ARC LITE tiers?

ARC deploys the gemma-4-e4b-it model (4-bit embedded, 4B parameters), while ARC LITE uses the smaller gemma-4-e2b-it variant (2-bit embedded, 2B parameters). Both use specialized GGUF files from the Unsloth repository, but ARC LITE targets minimal resource environments with stricter memory constraints, whereas ARC provides a balance between efficiency and capability for edge deployments.

Can I override the default model assigned to my hardware tier?

While ods/installers/lib/tier-map.sh defines default mappings, you can bypass automatic selection by explicitly setting the TIER variable or using the --tier flag with CLI commands like ods list-model --tier 3. However, modifying the underlying LLM_MODEL or GGUF_URL assignments requires editing the tier-map script directly, as the hardware detection logic in ods/installers/phases/02-detection.sh is designed to enforce validated configurations for stability.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →