How to Configure ODS for Cloud APIs and Local LLM Servers
Configure ODS to route LLM requests to cloud providers via LiteLLM or to local inference servers by setting ODS_MODE and LLM_BACKEND in your .env file, then rerun the installer to regenerate the Docker Compose stack.
The Osmantic/ODS (Open Data Science) platform provides flexible LLM backend configuration, allowing you to switch seamlessly between cloud API providers and local model servers. This guide explains how to configure ODS to use cloud APIs or local LLM servers using environment variables defined in .env.example and the LiteLLM routing system located in ods/config/litellm/.
Understanding ODS LLM Backend Architecture
ODS operates through three primary modes controlled by the ODS_MODE environment variable. The architecture routes requests either through a local inference container or via the LiteLLM proxy service.
local: Routes all requests to a local inference server such asllama-serveror an external endpoint like Ollamacloud: Routes requests through the LiteLLM proxy to commercial providers (Anthropic, OpenAI, etc.)hybrid: Uses local inference as primary with automatic cloud fallback
The concrete service implementation is selected via LLM_BACKEND, with valid values including llama-server, lemonade (for AMD GPUs), litellm, and external.
Configuring Cloud APIs via LiteLLM
To route requests to cloud providers like Anthropic or OpenAI, you must configure the LiteLLM proxy service and provide the necessary API credentials.
Setting Environment Variables
Edit your .env file (copied from ods/.env.example) to enable cloud mode:
ODS_MODE=cloud
LLM_BACKEND=litellm
As defined in .env.example (lines 76–79), these variables instruct the ODS installer to mount the LiteLLM configuration and expose the proxy container on port 4000.
Configuring the LiteLLM Routing Profile
The file ods/config/litellm/cloud.yaml defines the model mapping for cloud providers. Each entry specifies the target provider model and the environment variable that supplies the API key.
Example configuration from cloud.yaml:
model_list:
- model_name: ods/current
litellm_params:
model: anthropic/claude-sonnet-4-5-20250514
api_key: os.environ/ANTHROPIC_API_KEY
The litellm_params block accepts any LiteLLM-compatible provider string (e.g., openai/gpt-4, anthropic/claude-opus) and references environment variables using the os.environ/ prefix.
Exporting Provider API Keys
Add your provider credentials to .env. These values are never committed to the repository and are injected into the LiteLLM container at runtime:
ANTHROPIC_API_KEY=sk-xxxx
OPENAI_API_KEY=sk-xxxx
MINIMAX_API_KEY=xxxx
The LiteLLM container reads these variables to authenticate with upstream providers while presenting a unified OpenAI-compatible endpoint to ODS services.
Pinning Specific Models
To force the Open WebUI client to use a specific cloud model, override the base URL and API key:
OPEN_WEBUI_LLM_BASE_URL=http://litellm:4000/v1
OPEN_WEBUI_LLM_API_KEY=sk-test-litellm
After modifying these values, regenerate the Compose stack to apply changes.
Configuring Local LLM Servers
For air-gapped environments or local inference, configure ODS to use llama-server, lemonade, or external endpoints like Ollama or LM Studio.
Selecting Local Mode and Backend
Set the mode to local and choose your backend implementation:
ODS_MODE=local
LLM_BACKEND=llama-server
Alternative backends include:
lemonade: For AMD GPU inference using the Lemonade SDKexternal: For connecting to existing Ollama, LM Studio, or other OpenAI-compatible endpoints
Specifying External Inference Endpoints
When using LLM_BACKEND=external, provide the endpoint details:
EXTERNAL_LLM_URL=http://127.0.0.1:11434
EXTERNAL_LLM_PROVIDER=ollama
EXTERNAL_LLM_MODEL=qwen3.5:9b
The ODS web UI routes requests to this URL instead of the internal Docker network, allowing you to leverage existing local inference infrastructure.
GPU Backend and Model Selection
Specify your hardware acceleration layer:
GPU_BACKEND=nvidia
The installer automatically selects the optimal runtime (lemonade for AMD, llama-server for NVIDIA/CPU) based on this variable.
Override the default model selection by setting:
LLM_MODEL=qwen3.5-9b
GGUF_FILE=Qwen3.5-9B-Q4_K_M.gguf
These variables are consumed by the llama-server service definition in the generated Compose file.
Switching Between Modes
ODS includes a helper script at ods/extensions/services/litellm/select-config.sh that dynamically selects the appropriate LiteLLM configuration file (cloud.yaml, hybrid.yaml, lemonade.yaml) based on the current ODS_MODE.
When you change ODS_MODE or LLM_BACKEND, you must rerun the installer to regenerate the Docker Compose files:
./install.sh --quiet
# Or if ODS is already installed:
ods update --force
The installer executes the select-config.sh script during phase 06 ("directories") to ensure the correct configuration is mounted into the litellm container at runtime.
Summary
- Environment variables: Control routing via
ODS_MODE(local,cloud,hybrid) andLLM_BACKEND(litellm,llama-server,external) - LiteLLM configuration: Define provider mappings and API key references in
ods/config/litellm/cloud.yaml - Authentication: Store provider keys (e.g.,
ANTHROPIC_API_KEY) in.env, never in the repository - External endpoints: Use
LLM_BACKEND=externalwithEXTERNAL_LLM_URLto connect to Ollama or LM Studio - Deployment: Always rerun
install.shorods updateafter modifying backend configuration to regenerate the Compose stack
Frequently Asked Questions
What is the difference between ODS_MODE and LLM_BACKEND?
ODS_MODE determines the overall architectural strategy (local-only, cloud-only, or hybrid fallback), while LLM_BACKEND specifies the concrete service implementation that fulfills the requests. For example, setting ODS_MODE=cloud with LLM_BACKEND=litellm routes traffic through the LiteLLM proxy, whereas ODS_MODE=local with LLM_BACKEND=llama-server uses the local GGUF inference container.
How do I add a new cloud provider to the LiteLLM configuration?
Edit ods/config/litellm/cloud.yaml and add a new entry to the model_list array. Specify the provider-specific model identifier in litellm_params.model (e.g., vertex_ai/gemini-pro) and reference the API key environment variable using os.environ/YOUR_KEY_NAME. Add the corresponding key to your .env file, then run ods update to apply the changes.
Can I use both local and cloud models simultaneously?
Yes. Set ODS_MODE=hybrid to enable the hybrid routing configuration. In this mode, ODS attempts local inference first and falls back to the LiteLLM-configured cloud provider if the local service is unavailable or returns errors. The select-config.sh script automatically loads hybrid.yaml which defines the failover logic.
Why are my API key changes not taking effect?
ODS loads environment variables into the containers at stack creation time, not at runtime. After modifying .env or any file in ods/config/litellm/, you must regenerate the Docker Compose configuration by running ods update --force or ./install.sh --quiet. Simply restarting containers without regenerating the stack preserves the previous configuration state.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →