How to Configure ODS to Use an External LLM Endpoint (Ollama or LM Studio)

ODS routes all LLM traffic to external OpenAI-compatible endpoints by validating the host connection, rewriting the URL for container networking, and injecting a dedicated Docker Compose overlay that replaces local inference containers.

The Osmantic/ODS repository allows you to replace its built-in inference backend—typically llama-server or Lemonade—with an external LLM service such as Ollama or LM Studio. This configuration leverages the standard OpenAI-compatible API to redirect model requests while maintaining the Dashboard API and Open WebUI functionality.

Prerequisites for External LLM Configuration

Before running the installer, ensure your external LLM service is active and accessible on the host machine. Start Ollama or LM Studio and verify the HTTP API endpoint is responding.

For Ollama on Linux or macOS:

ollama serve &
curl http://127.0.0.1:11434/v1/models

Your service must expose the /v1/models endpoint for ODS to discover available models during installation.

Configuring the External Endpoint via CLI Flags

The installer accepts explicit flags to define the external connection. According to install-core.sh (lines 161–169), the supported flags are:

  • --external-llm-url: The host-facing URL (e.g., http://127.0.0.1:11434)
  • --external-llm-provider: Either ollama or lmstudio
  • --external-llm-model: The exact model name (e.g., qwen3.5:9b)
  • --reuse-external-llm: Skip validation if the endpoint was previously configured
  • --no-external-llm: Revert to managed local inference

Run the installation script with these parameters:

./install.sh \
  --external-llm-url http://127.0.0.1:11434 \
  --external-llm-provider ollama \
  --external-llm-model qwen3.5:9b \
  --non-interactive

The installer stores these values in .env and sets LLM_BACKEND=external for the Dashboard API service.

Validation and Environment Variable Resolution

During phase 02b-external-services.sh (lines 106–116), the installer validates the connection by contacting EXTERNAL_LLM_URL and verifying the specified model exists in the remote catalog. If validation fails, the installation aborts immediately.

Upon success, the installer writes the following variables to .env (documented in .env.example, lines 100–110):

  • EXTERNAL_LLM_URL: The original host-facing URL you provided
  • EXTERNAL_LLM_CONTAINER_URL: Rewritten to host.docker.internal for internal container routing
  • EXTERNAL_LLM_PROVIDER: Provider type determining discovery logic
  • EXTERNAL_LLM_MODEL: The specific model identifier

All services use LLM_API_BASE_PATH=/v1 and ODS_TALK_VISION_URL="${EXTERNAL_LLM_CONTAINER_URL}/v1" to communicate with the external endpoint.

Understanding the Docker Compose Overlay

The docker-compose.external-llm.yml file functions as a composition overlay that removes local inference containers (llama-server, model-router) and injects extra_hosts mappings. As implemented in lines 24–30 of the overlay, this configuration ensures the Dashboard API and Open WebUI reach the external service via host.docker.internal.

The scripts/resolve-compose-stack.sh script automatically detects EXTERNAL_LLM_URL in your environment and appends the overlay to the final Docker Compose configuration. You can verify the merged configuration:

docker compose config | grep -A5 'dashboard-api'

Look for extra_hosts entries mapping host.docker.internal to host-gateway and the absence of llama-server services.

Reverting to Local Inference

To restore the managed inference stack and remove external routing, execute the installer with the reversal flag:

./install.sh --no-external-llm

This clears EXTERNAL_LLM_URL from the environment and removes the docker-compose.external-llm.yml overlay, restoring the default llama-server or Lemonade containers in the next deployment.

Summary

  • ODS supports Ollama and LM Studio via OpenAI-compatible API endpoints by using a Docker Compose overlay that replaces local inference services.
  • The installer validates external connectivity during phase 02b-external-services.sh and rewrites EXTERNAL_LLM_URL to EXTERNAL_LLM_CONTAINER_URL for internal Docker networking.
  • Use CLI flags --external-llm-url, --external-llm-provider, and --external-llm-model to configure the connection during installation.
  • The docker-compose.external-llm.yml overlay removes llama-server containers and adds host.docker.internal routing for the Dashboard API.
  • Revert to local inference anytime by running ./install.sh --no-external-llm.

Frequently Asked Questions

What URL format should I use for the external LLM endpoint?

Provide the base HTTP URL where Ollama or LM Studio listens on your host machine, typically http://127.0.0.1:11434 for Ollama. The installer automatically rewrites this to host.docker.internal:11434 in EXTERNAL_LLM_CONTAINER_URL so containers can reach the host service. Do not include /v1 in the URL; ODS appends the API path internally using LLM_API_BASE_PATH.

How does ODS validate that my external LLM is working?

During the 02b-external-services.sh installation phase, the installer attempts to connect to EXTERNAL_LLM_URL and queries the /v1/models endpoint to confirm the specific model you specified in --external-llm-model exists in the catalog. If the endpoint is unreachable or the model is missing, the installation halts with an error message before modifying your Docker Compose stack.

Can I switch between external and local LLM backends after installation?

Yes. Run ./install.sh --no-external-llm to disable external routing and restore the local llama-server containers. To switch back to external, rerun the installer with the --external-llm-url and related flags. The resolve-compose-stack.sh script dynamically includes or excludes the docker-compose.external-llm.yml overlay based on the presence of EXTERNAL_LLM_URL in your .env file.

Why does the Dashboard API show LLM_BACKEND=external?

When an external endpoint is configured, the installer sets LLM_BACKEND=external in the Dashboard API environment variables (as seen in docker-compose.external-llm.yml). This signals the application to route all inference requests to ODS_TALK_VISION_URL instead of the local model router. The Dashboard UI hook in extensions/services/dashboard/src/hooks/useModels.js detects this state and displays an indicator that an external Ollama or LM Studio backend is active.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →