How to Configure ODS for Cloud APIs and Local LLM Servers

Configure ODS to route LLM requests to cloud providers via LiteLLM or to local inference servers by setting ODS_MODE and LLM_BACKEND in your .env file, then rerun the installer to regenerate the Docker Compose stack.

The Osmantic/ODS (Open Data Science) platform provides flexible LLM backend configuration, allowing you to switch seamlessly between cloud API providers and local model servers. This guide explains how to configure ODS to use cloud APIs or local LLM servers using environment variables defined in .env.example and the LiteLLM routing system located in ods/config/litellm/.

Understanding ODS LLM Backend Architecture

ODS operates through three primary modes controlled by the ODS_MODE environment variable. The architecture routes requests either through a local inference container or via the LiteLLM proxy service.

  • local: Routes all requests to a local inference server such as llama-server or an external endpoint like Ollama
  • cloud: Routes requests through the LiteLLM proxy to commercial providers (Anthropic, OpenAI, etc.)
  • hybrid: Uses local inference as primary with automatic cloud fallback

The concrete service implementation is selected via LLM_BACKEND, with valid values including llama-server, lemonade (for AMD GPUs), litellm, and external.

Configuring Cloud APIs via LiteLLM

To route requests to cloud providers like Anthropic or OpenAI, you must configure the LiteLLM proxy service and provide the necessary API credentials.

Setting Environment Variables

Edit your .env file (copied from ods/.env.example) to enable cloud mode:

ODS_MODE=cloud
LLM_BACKEND=litellm

As defined in .env.example (lines 76–79), these variables instruct the ODS installer to mount the LiteLLM configuration and expose the proxy container on port 4000.

Configuring the LiteLLM Routing Profile

The file ods/config/litellm/cloud.yaml defines the model mapping for cloud providers. Each entry specifies the target provider model and the environment variable that supplies the API key.

Example configuration from cloud.yaml:

model_list:
  - model_name: ods/current
    litellm_params:
      model: anthropic/claude-sonnet-4-5-20250514
      api_key: os.environ/ANTHROPIC_API_KEY

The litellm_params block accepts any LiteLLM-compatible provider string (e.g., openai/gpt-4, anthropic/claude-opus) and references environment variables using the os.environ/ prefix.

Exporting Provider API Keys

Add your provider credentials to .env. These values are never committed to the repository and are injected into the LiteLLM container at runtime:

ANTHROPIC_API_KEY=sk-xxxx
OPENAI_API_KEY=sk-xxxx
MINIMAX_API_KEY=xxxx

The LiteLLM container reads these variables to authenticate with upstream providers while presenting a unified OpenAI-compatible endpoint to ODS services.

Pinning Specific Models

To force the Open WebUI client to use a specific cloud model, override the base URL and API key:

OPEN_WEBUI_LLM_BASE_URL=http://litellm:4000/v1
OPEN_WEBUI_LLM_API_KEY=sk-test-litellm

After modifying these values, regenerate the Compose stack to apply changes.

Configuring Local LLM Servers

For air-gapped environments or local inference, configure ODS to use llama-server, lemonade, or external endpoints like Ollama or LM Studio.

Selecting Local Mode and Backend

Set the mode to local and choose your backend implementation:

ODS_MODE=local
LLM_BACKEND=llama-server

Alternative backends include:

  • lemonade: For AMD GPU inference using the Lemonade SDK
  • external: For connecting to existing Ollama, LM Studio, or other OpenAI-compatible endpoints

Specifying External Inference Endpoints

When using LLM_BACKEND=external, provide the endpoint details:

EXTERNAL_LLM_URL=http://127.0.0.1:11434
EXTERNAL_LLM_PROVIDER=ollama
EXTERNAL_LLM_MODEL=qwen3.5:9b

The ODS web UI routes requests to this URL instead of the internal Docker network, allowing you to leverage existing local inference infrastructure.

GPU Backend and Model Selection

Specify your hardware acceleration layer:

GPU_BACKEND=nvidia

The installer automatically selects the optimal runtime (lemonade for AMD, llama-server for NVIDIA/CPU) based on this variable.

Override the default model selection by setting:

LLM_MODEL=qwen3.5-9b
GGUF_FILE=Qwen3.5-9B-Q4_K_M.gguf

These variables are consumed by the llama-server service definition in the generated Compose file.

Switching Between Modes

ODS includes a helper script at ods/extensions/services/litellm/select-config.sh that dynamically selects the appropriate LiteLLM configuration file (cloud.yaml, hybrid.yaml, lemonade.yaml) based on the current ODS_MODE.

When you change ODS_MODE or LLM_BACKEND, you must rerun the installer to regenerate the Docker Compose files:

./install.sh --quiet

# Or if ODS is already installed:

ods update --force

The installer executes the select-config.sh script during phase 06 ("directories") to ensure the correct configuration is mounted into the litellm container at runtime.

Summary

  • Environment variables: Control routing via ODS_MODE (local, cloud, hybrid) and LLM_BACKEND (litellm, llama-server, external)
  • LiteLLM configuration: Define provider mappings and API key references in ods/config/litellm/cloud.yaml
  • Authentication: Store provider keys (e.g., ANTHROPIC_API_KEY) in .env, never in the repository
  • External endpoints: Use LLM_BACKEND=external with EXTERNAL_LLM_URL to connect to Ollama or LM Studio
  • Deployment: Always rerun install.sh or ods update after modifying backend configuration to regenerate the Compose stack

Frequently Asked Questions

What is the difference between ODS_MODE and LLM_BACKEND?

ODS_MODE determines the overall architectural strategy (local-only, cloud-only, or hybrid fallback), while LLM_BACKEND specifies the concrete service implementation that fulfills the requests. For example, setting ODS_MODE=cloud with LLM_BACKEND=litellm routes traffic through the LiteLLM proxy, whereas ODS_MODE=local with LLM_BACKEND=llama-server uses the local GGUF inference container.

How do I add a new cloud provider to the LiteLLM configuration?

Edit ods/config/litellm/cloud.yaml and add a new entry to the model_list array. Specify the provider-specific model identifier in litellm_params.model (e.g., vertex_ai/gemini-pro) and reference the API key environment variable using os.environ/YOUR_KEY_NAME. Add the corresponding key to your .env file, then run ods update to apply the changes.

Can I use both local and cloud models simultaneously?

Yes. Set ODS_MODE=hybrid to enable the hybrid routing configuration. In this mode, ODS attempts local inference first and falls back to the LiteLLM-configured cloud provider if the local service is unavailable or returns errors. The select-config.sh script automatically loads hybrid.yaml which defines the failover logic.

Why are my API key changes not taking effect?

ODS loads environment variables into the containers at stack creation time, not at runtime. After modifying .env or any file in ods/config/litellm/, you must regenerate the Docker Compose configuration by running ods update --force or ./install.sh --quiet. Simply restarting containers without regenerating the stack preserves the previous configuration state.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →