Which LLM Models Does ODS Support by Default? Local and Cloud Providers Explained

ODS supports local Llama-Server models (CPU/NVIDIA) and Lemonade (AMD) via the openai/default alias, while pre-configuring cloud access to Claude Sonnet, GPT-4o, and MiniMax through a unified Litellm routing interface.

The Osmantic/ODS repository provides an extensible backend system that enables zero-configuration LLM inference. By default, ODS automatically detects your hardware architecture and routes requests to locally-served GGUF models, while maintaining optional fallback paths to cloud APIs through hard-coded configuration files.

Default Local LLM Providers

ODS ships with three hardware-specific backend configurations that define its default local model support. These backends proxy requests to local inference engines using standardized model aliases, allowing any compatible GGUF file placed in ./data/llama-server/models to become immediately accessible.

CPU and NVIDIA GPU Support (Llama-Server)

For x86/ARM CPU and NVIDIA GPU environments, ODS defaults to the Llama-Server engine. According to the source configuration in ods/config/backends/cpu.json and ods/config/backends/nvidia.json, both backends use the backend ID cpu and nvidia respectively, map to the llama-server engine, and expose the model alias openai/default. This alias proxies all requests to the local llama-server instance, enabling out-of-the-box inference without cloud dependencies.

  • CPU Backend: Defined in ods/config/backends/cpu.json, routes to local llama-server on port 8080
  • NVIDIA Backend: Defined in ods/config/backends/nvidia.json, utilizes GPU acceleration through the same llama-server engine
  • Model Discovery: Any GGUF model placed in ./data/llama-server/models is automatically discoverable under the openai/default alias

AMD GPU Support (Lemonade)

On AMD hardware, ODS defaults to the Lemonade engine rather than llama-server. The configuration in ods/config/backends/amd.json specifies backend ID amd, engine lemonade, and maintains the same openai/default alias for API compatibility. Lemonade serves GGUF models through an OpenAI-compatible proxy, ensuring consistent request formatting across hardware platforms.

Litellm Routing and Model Aliases

ODS implements a Litellm switchboard that unifies local and cloud access behind single model names. The configuration in ods/config/litellm/local.yaml defines the primary alias ods/current, which resolves to openai/default and points to whichever local engine is active (llama-server or lemonade). When no cloud API keys are present, the switchboard automatically falls back to this local engine, providing seamless operation without manual model selection.

Key routing behaviors include:

  1. Default Local Routing: ods/current → openai/default → local engine
  2. Automatic Fallback: Absence of cloud credentials triggers instant local engine usage
  3. Consistent Interface: Single model name works across CPU, NVIDIA, and AMD deployments

Pre-Configured Cloud Models

While local inference is the default, ODS includes pre-configured cloud providers in ods/config/litellm/cloud.yaml for users who provide API credentials. These configurations load automatically but remain dormant until valid keys are detected.

Default Cloud Model:

  • anthropic/claude-sonnet-4-5-20250514 (primary cloud default)

Additional Pre-Configured Models:

  • openai/gpt-4o
  • openai/MiniMax-M2.7
  • openai/MiniMax-M2.7-highspeed

These models are selectable via the model_name field in Litellm requests, allowing immediate switching between local GGUF inference and commercial APIs without configuration changes.

How to Verify Your Default Model Configuration

You can confirm which LLM models are active in your ODS deployment using the following validation commands.

Start ODS and initialize the default local server:

ods install

Verify the local LLM endpoint responds using the default alias:

curl -X POST http://127.0.0.1:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"openai/default","messages":[{"role":"user","content":"Hello"}]}'

Test the Litellm router to confirm unified routing:

curl -X POST http://127.0.0.1:4000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"ods/current","messages":[{"role":"user","content":"Explain ODS"}]}'

Summary

  • ODS defaults to local inference using llama-server for CPU/NVIDIA systems and lemonade for AMD GPUs, configured in ods/config/backends/.
  • The openai/default alias provides a consistent API endpoint for all local GGUF models placed in the llama-server models directory.
  • Litellm unifies access through the ods/current alias defined in ods/config/litellm/local.yaml, automatically routing to local engines when cloud keys are absent.
  • Cloud models including Claude Sonnet, GPT-4o, and MiniMax are pre-configured in ods/config/litellm/cloud.yaml but require API keys to activate.
  • All default behaviors are hard-coded in JSON backend descriptors and YAML configuration files, enabling immediate operation after installation.

Frequently Asked Questions

Does ODS require an API key to work out of the box?

No. ODS is designed to function immediately after installation using local GGUF models. The default configurations in ods/config/backends/cpu.json, nvidia.json, and amd.json route requests to local inference engines without requiring cloud credentials. API keys are only necessary if you wish to use the pre-configured cloud models defined in ods/config/litellm/cloud.yaml.

How do I add my own GGUF models to ODS?

Place your GGUF model files into the ./data/llama-server/models directory. The local llama-server engine automatically discovers these files, making them accessible through the openai/default alias. No configuration file modifications are required; the models appear immediately in the local endpoint upon restarting the service.

Can I use ODS with both local and cloud models simultaneously?

Yes. The Litellm switchboard in ods/config/litellm/local.yaml manages routing between local and cloud providers. With valid API keys configured, you can specify cloud model names like anthropic/claude-sonnet-4-5-20250514 or openai/gpt-4o in your requests while maintaining access to local models via ods/current or openai/default.

Where is the default model configuration stored in the ODS repository?

Default model providers are defined in two locations: backend-specific configurations reside in ods/config/backends/ (cpu.json, nvidia.json, amd.json), while routing logic and cloud defaults are stored in ods/config/litellm/ (local.yaml, cloud.yaml). These files contain the hard-coded model aliases and engine mappings that determine which LLM models ODS supports by default.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →