# How to Configure ODS for Cloud APIs and Local LLM Servers

> Configure ODS to use cloud APIs or local LLM servers by setting ODS_MODE and LLM_BACKEND in your .env file. Rerun the installer to update your Docker Compose stack.

- Repository: [Osmantic/ODS](https://github.com/Osmantic/ODS)
- Tags: how-to-guide
- Published: 2026-09-01

---

**Configure ODS to route LLM requests to cloud providers via LiteLLM or to local inference servers by setting `ODS_MODE` and `LLM_BACKEND` in your `.env` file, then rerun the installer to regenerate the Docker Compose stack.**

The Osmantic/ODS (Open Data Science) platform provides flexible LLM backend configuration, allowing you to switch seamlessly between cloud API providers and local model servers. This guide explains how to configure ODS to use cloud APIs or local LLM servers using environment variables defined in `.env.example` and the LiteLLM routing system located in `ods/config/litellm/`.

## Understanding ODS LLM Backend Architecture

ODS operates through three primary modes controlled by the **`ODS_MODE`** environment variable. The architecture routes requests either through a local inference container or via the LiteLLM proxy service.

- **`local`**: Routes all requests to a local inference server such as `llama-server` or an external endpoint like Ollama
- **`cloud`**: Routes requests through the LiteLLM proxy to commercial providers (Anthropic, OpenAI, etc.)
- **`hybrid`**: Uses local inference as primary with automatic cloud fallback

The concrete service implementation is selected via **`LLM_BACKEND`**, with valid values including `llama-server`, `lemonade` (for AMD GPUs), `litellm`, and `external`.

## Configuring Cloud APIs via LiteLLM

To route requests to cloud providers like Anthropic or OpenAI, you must configure the LiteLLM proxy service and provide the necessary API credentials.

### Setting Environment Variables

Edit your `.env` file (copied from `ods/.env.example`) to enable cloud mode:

```text
ODS_MODE=cloud
LLM_BACKEND=litellm

```

As defined in `.env.example` (lines 76–79), these variables instruct the ODS installer to mount the LiteLLM configuration and expose the proxy container on port 4000.

### Configuring the LiteLLM Routing Profile

The file [`ods/config/litellm/cloud.yaml`](https://github.com/Osmantic/ODS/blob/main/ods/config/litellm/cloud.yaml) defines the model mapping for cloud providers. Each entry specifies the target provider model and the environment variable that supplies the API key.

Example configuration from [`cloud.yaml`](https://github.com/Osmantic/ODS/blob/main/cloud.yaml):

```yaml
model_list:
  - model_name: ods/current
    litellm_params:
      model: anthropic/claude-sonnet-4-5-20250514
      api_key: os.environ/ANTHROPIC_API_KEY

```

The `litellm_params` block accepts any LiteLLM-compatible provider string (e.g., `openai/gpt-4`, `anthropic/claude-opus`) and references environment variables using the `os.environ/` prefix.

### Exporting Provider API Keys

Add your provider credentials to `.env`. These values are never committed to the repository and are injected into the LiteLLM container at runtime:

```text
ANTHROPIC_API_KEY=sk-xxxx
OPENAI_API_KEY=sk-xxxx
MINIMAX_API_KEY=xxxx

```

The LiteLLM container reads these variables to authenticate with upstream providers while presenting a unified OpenAI-compatible endpoint to ODS services.

### Pinning Specific Models

To force the Open WebUI client to use a specific cloud model, override the base URL and API key:

```text
OPEN_WEBUI_LLM_BASE_URL=http://litellm:4000/v1
OPEN_WEBUI_LLM_API_KEY=sk-test-litellm

```

After modifying these values, regenerate the Compose stack to apply changes.

## Configuring Local LLM Servers

For air-gapped environments or local inference, configure ODS to use `llama-server`, `lemonade`, or external endpoints like Ollama or LM Studio.

### Selecting Local Mode and Backend

Set the mode to local and choose your backend implementation:

```text
ODS_MODE=local
LLM_BACKEND=llama-server

```

Alternative backends include:
- **`lemonade`**: For AMD GPU inference using the Lemonade SDK
- **`external`**: For connecting to existing Ollama, LM Studio, or other OpenAI-compatible endpoints

### Specifying External Inference Endpoints

When using `LLM_BACKEND=external`, provide the endpoint details:

```text
EXTERNAL_LLM_URL=http://127.0.0.1:11434
EXTERNAL_LLM_PROVIDER=ollama
EXTERNAL_LLM_MODEL=qwen3.5:9b

```

The ODS web UI routes requests to this URL instead of the internal Docker network, allowing you to leverage existing local inference infrastructure.

### GPU Backend and Model Selection

Specify your hardware acceleration layer:

```text
GPU_BACKEND=nvidia

```

The installer automatically selects the optimal runtime (`lemonade` for AMD, `llama-server` for NVIDIA/CPU) based on this variable.

Override the default model selection by setting:

```text
LLM_MODEL=qwen3.5-9b
GGUF_FILE=Qwen3.5-9B-Q4_K_M.gguf

```

These variables are consumed by the `llama-server` service definition in the generated Compose file.

## Switching Between Modes

ODS includes a helper script at [`ods/extensions/services/litellm/select-config.sh`](https://github.com/Osmantic/ODS/blob/main/ods/extensions/services/litellm/select-config.sh) that dynamically selects the appropriate LiteLLM configuration file ([`cloud.yaml`](https://github.com/Osmantic/ODS/blob/main/cloud.yaml), [`hybrid.yaml`](https://github.com/Osmantic/ODS/blob/main/hybrid.yaml), [`lemonade.yaml`](https://github.com/Osmantic/ODS/blob/main/lemonade.yaml)) based on the current `ODS_MODE`.

When you change `ODS_MODE` or `LLM_BACKEND`, you must rerun the installer to regenerate the Docker Compose files:

```bash
./install.sh --quiet

# Or if ODS is already installed:

ods update --force

```

The installer executes the [`select-config.sh`](https://github.com/Osmantic/ODS/blob/main/select-config.sh) script during phase 06 ("directories") to ensure the correct configuration is mounted into the `litellm` container at runtime.

## Summary

- **Environment variables**: Control routing via `ODS_MODE` (`local`, `cloud`, `hybrid`) and `LLM_BACKEND` (`litellm`, `llama-server`, `external`)
- **LiteLLM configuration**: Define provider mappings and API key references in [`ods/config/litellm/cloud.yaml`](https://github.com/Osmantic/ODS/blob/main/ods/config/litellm/cloud.yaml)
- **Authentication**: Store provider keys (e.g., `ANTHROPIC_API_KEY`) in `.env`, never in the repository
- **External endpoints**: Use `LLM_BACKEND=external` with `EXTERNAL_LLM_URL` to connect to Ollama or LM Studio
- **Deployment**: Always rerun [`install.sh`](https://github.com/Osmantic/ODS/blob/main/install.sh) or `ods update` after modifying backend configuration to regenerate the Compose stack

## Frequently Asked Questions

### What is the difference between ODS_MODE and LLM_BACKEND?

`ODS_MODE` determines the overall architectural strategy (local-only, cloud-only, or hybrid fallback), while `LLM_BACKEND` specifies the concrete service implementation that fulfills the requests. For example, setting `ODS_MODE=cloud` with `LLM_BACKEND=litellm` routes traffic through the LiteLLM proxy, whereas `ODS_MODE=local` with `LLM_BACKEND=llama-server` uses the local GGUF inference container.

### How do I add a new cloud provider to the LiteLLM configuration?

Edit [`ods/config/litellm/cloud.yaml`](https://github.com/Osmantic/ODS/blob/main/ods/config/litellm/cloud.yaml) and add a new entry to the `model_list` array. Specify the provider-specific model identifier in `litellm_params.model` (e.g., `vertex_ai/gemini-pro`) and reference the API key environment variable using `os.environ/YOUR_KEY_NAME`. Add the corresponding key to your `.env` file, then run `ods update` to apply the changes.

### Can I use both local and cloud models simultaneously?

Yes. Set `ODS_MODE=hybrid` to enable the hybrid routing configuration. In this mode, ODS attempts local inference first and falls back to the LiteLLM-configured cloud provider if the local service is unavailable or returns errors. The [`select-config.sh`](https://github.com/Osmantic/ODS/blob/main/select-config.sh) script automatically loads [`hybrid.yaml`](https://github.com/Osmantic/ODS/blob/main/hybrid.yaml) which defines the failover logic.

### Why are my API key changes not taking effect?

ODS loads environment variables into the containers at stack creation time, not at runtime. After modifying `.env` or any file in `ods/config/litellm/`, you must regenerate the Docker Compose configuration by running `ods update --force` or `./install.sh --quiet`. Simply restarting containers without regenerating the stack preserves the previous configuration state.