# Which LLM Models Does ODS Support by Default? Local and Cloud Providers Explained

> Discover which LLM models ODS supports by default. Explore local Llama-Server, Lemonade, and cloud providers like GPT-4o, Claude Sonnet, and MiniMax with our guide.

- Repository: [Osmantic/ODS](https://github.com/Osmantic/ODS)
- Tags: api-reference
- Published: 2026-08-30

---

**ODS supports local Llama-Server models (CPU/NVIDIA) and Lemonade (AMD) via the `openai/default` alias, while pre-configuring cloud access to Claude Sonnet, GPT-4o, and MiniMax through a unified Litellm routing interface.**

The Osmantic/ODS repository provides an extensible backend system that enables zero-configuration LLM inference. By default, ODS automatically detects your hardware architecture and routes requests to locally-served GGUF models, while maintaining optional fallback paths to cloud APIs through hard-coded configuration files.

## Default Local LLM Providers

ODS ships with three hardware-specific backend configurations that define its default local model support. These backends proxy requests to local inference engines using standardized model aliases, allowing any compatible GGUF file placed in `./data/llama-server/models` to become immediately accessible.

### CPU and NVIDIA GPU Support (Llama-Server)

For x86/ARM CPU and NVIDIA GPU environments, ODS defaults to the **Llama-Server** engine. According to the source configuration in [`ods/config/backends/cpu.json`](https://github.com/Osmantic/ODS/blob/main/ods/config/backends/cpu.json) and [`ods/config/backends/nvidia.json`](https://github.com/Osmantic/ODS/blob/main/ods/config/backends/nvidia.json), both backends use the **backend ID** `cpu` and `nvidia` respectively, map to the `llama-server` engine, and expose the model alias `openai/default`. This alias proxies all requests to the local llama-server instance, enabling out-of-the-box inference without cloud dependencies.

- **CPU Backend**: Defined in [`ods/config/backends/cpu.json`](https://github.com/Osmantic/ODS/blob/main/ods/config/backends/cpu.json), routes to local llama-server on port 8080
- **NVIDIA Backend**: Defined in [`ods/config/backends/nvidia.json`](https://github.com/Osmantic/ODS/blob/main/ods/config/backends/nvidia.json), utilizes GPU acceleration through the same llama-server engine
- **Model Discovery**: Any GGUF model placed in `./data/llama-server/models` is automatically discoverable under the `openai/default` alias

### AMD GPU Support (Lemonade)

On AMD hardware, ODS defaults to the **Lemonade** engine rather than llama-server. The configuration in [`ods/config/backends/amd.json`](https://github.com/Osmantic/ODS/blob/main/ods/config/backends/amd.json) specifies backend ID `amd`, engine `lemonade`, and maintains the same `openai/default` alias for API compatibility. Lemonade serves GGUF models through an OpenAI-compatible proxy, ensuring consistent request formatting across hardware platforms.

## Litellm Routing and Model Aliases

ODS implements a **Litellm switchboard** that unifies local and cloud access behind single model names. The configuration in [`ods/config/litellm/local.yaml`](https://github.com/Osmantic/ODS/blob/main/ods/config/litellm/local.yaml) defines the primary alias `ods/current`, which resolves to `openai/default` and points to whichever local engine is active (llama-server or lemonade). When no cloud API keys are present, the switchboard automatically falls back to this local engine, providing seamless operation without manual model selection.

Key routing behaviors include:

1. **Default Local Routing**: `ods/current` → `openai/default` → local engine
2. **Automatic Fallback**: Absence of cloud credentials triggers instant local engine usage
3. **Consistent Interface**: Single model name works across CPU, NVIDIA, and AMD deployments

## Pre-Configured Cloud Models

While local inference is the default, ODS includes pre-configured cloud providers in [`ods/config/litellm/cloud.yaml`](https://github.com/Osmantic/ODS/blob/main/ods/config/litellm/cloud.yaml) for users who provide API credentials. These configurations load automatically but remain dormant until valid keys are detected.

**Default Cloud Model**:
- `anthropic/claude-sonnet-4-5-20250514` (primary cloud default)

**Additional Pre-Configured Models**:
- `openai/gpt-4o`
- `openai/MiniMax-M2.7`
- `openai/MiniMax-M2.7-highspeed`

These models are selectable via the `model_name` field in Litellm requests, allowing immediate switching between local GGUF inference and commercial APIs without configuration changes.

## How to Verify Your Default Model Configuration

You can confirm which LLM models are active in your ODS deployment using the following validation commands.

Start ODS and initialize the default local server:

```bash
ods install

```

Verify the local LLM endpoint responds using the default alias:

```bash
curl -X POST http://127.0.0.1:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"openai/default","messages":[{"role":"user","content":"Hello"}]}'

```

Test the Litellm router to confirm unified routing:

```bash
curl -X POST http://127.0.0.1:4000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"ods/current","messages":[{"role":"user","content":"Explain ODS"}]}'

```

## Summary

- **ODS defaults to local inference** using `llama-server` for CPU/NVIDIA systems and `lemonade` for AMD GPUs, configured in `ods/config/backends/`.
- **The `openai/default` alias** provides a consistent API endpoint for all local GGUF models placed in the llama-server models directory.
- **Litellm unifies access** through the `ods/current` alias defined in [`ods/config/litellm/local.yaml`](https://github.com/Osmantic/ODS/blob/main/ods/config/litellm/local.yaml), automatically routing to local engines when cloud keys are absent.
- **Cloud models** including Claude Sonnet, GPT-4o, and MiniMax are pre-configured in [`ods/config/litellm/cloud.yaml`](https://github.com/Osmantic/ODS/blob/main/ods/config/litellm/cloud.yaml) but require API keys to activate.
- All default behaviors are hard-coded in JSON backend descriptors and YAML configuration files, enabling immediate operation after installation.

## Frequently Asked Questions

### Does ODS require an API key to work out of the box?

No. ODS is designed to function immediately after installation using local GGUF models. The default configurations in [`ods/config/backends/cpu.json`](https://github.com/Osmantic/ODS/blob/main/ods/config/backends/cpu.json), [`nvidia.json`](https://github.com/Osmantic/ODS/blob/main/nvidia.json), and [`amd.json`](https://github.com/Osmantic/ODS/blob/main/amd.json) route requests to local inference engines without requiring cloud credentials. API keys are only necessary if you wish to use the pre-configured cloud models defined in [`ods/config/litellm/cloud.yaml`](https://github.com/Osmantic/ODS/blob/main/ods/config/litellm/cloud.yaml).

### How do I add my own GGUF models to ODS?

Place your GGUF model files into the `./data/llama-server/models` directory. The local llama-server engine automatically discovers these files, making them accessible through the `openai/default` alias. No configuration file modifications are required; the models appear immediately in the local endpoint upon restarting the service.

### Can I use ODS with both local and cloud models simultaneously?

Yes. The Litellm switchboard in [`ods/config/litellm/local.yaml`](https://github.com/Osmantic/ODS/blob/main/ods/config/litellm/local.yaml) manages routing between local and cloud providers. With valid API keys configured, you can specify cloud model names like `anthropic/claude-sonnet-4-5-20250514` or `openai/gpt-4o` in your requests while maintaining access to local models via `ods/current` or `openai/default`.

### Where is the default model configuration stored in the ODS repository?

Default model providers are defined in two locations: backend-specific configurations reside in `ods/config/backends/` (cpu.json, nvidia.json, amd.json), while routing logic and cloud defaults are stored in `ods/config/litellm/` (local.yaml, cloud.yaml). These files contain the hard-coded model aliases and engine mappings that determine which LLM models ODS supports by default.