Understanding the ODS Service Communication Flow: A Complete Architecture Guide

ODS routes all traffic through localhost-bound containers, where browser-based UIs communicate with LLM inference, agent gateways, and auxiliary services via a layered Docker Compose stack orchestrated by the Dashboard API.

The ODS service communication flow defines how the Osmantic/ODS (Open-Source AI Stack) orchestrates interactions between user interfaces, inference engines, and supporting microservices. Every component communicates over HTTP endpoints restricted to 127.0.0.1, creating a secure, single-machine architecture where the Docker Compose layer manages internal routing. Understanding this flow is essential for debugging API calls, extending the stack, or integrating custom agents.

Browser to UI Services: The Entry Points

All user interactions begin in the browser and terminate at four primary UI containers defined in the Core Services block of ARCHITECTURE.md【/cache/repos/github.com/Osmantic/ODS/main/ARCHITECTURE.md#L24-L30】.

  • open-webui (port 3000) – Serves the main chat interface.
  • dashboard (port 3001) – Provides the control center for system management.
  • opencode (port 3003) – Hosts the web-based IDE.
  • hermes-proxy (port 9120) – Acts as the default gateway for agent interactions.

These services expose ports only on localhost, ensuring no direct external network access bypasses the host machine.

UI to LLM Inference: Direct and Proxy Paths

When the chat UI sends a generation request, it reaches the LLM inference layer through two possible routes as shown in the Inference Layer diagram【/cache/repos/github.com/Osmantic/ODS/main/ARCHITECTURE.md#L24-L26】.

Direct Inference Path

The fastest route connects directly to the inference server:

curl -X POST http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"local","messages":[{"role":"user","content":"Hello"}]}'

This path bypasses intermediaries and hits llama-server:8080 immediately.

LiteLLM Gateway Path

For hybrid deployments supporting cloud fallback, requests route through the LiteLLM gateway (litellm:4000):

curl -X POST http://localhost:4000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"local","messages":[{"role":"user","content":"Hello"}]}'

In this configuration, open-webui sends traffic to http://localhost:4000/v1/chat/completions, which then proxies to llama-server or falls back to a cloud provider if hybrid mode is active.

Dashboard to Dashboard-API: System Management

The dashboard UI does not query inference engines directly. Instead, it communicates with the Dashboard-API service (dashboard-api:3002) for runtime metadata and health monitoring.

Feature Discovery and Health Checks

The Dashboard-API exposes /api/features to list enabled capabilities based on service manifests found in extensions/services/*/manifest.yaml. Clients must authenticate using an API key header.

curl http://localhost:3002/api/features \
  -H "x-api-key: $DASHBOARD_API_KEY"

Agent Metrics Endpoints

Agent telemetry is implemented in ods/extensions/services/dashboard-api/routers/agents.py. This module provides dual-format outputs for the same data:

  • JSON metrics (lines 14-18)【/cache/repos/github.com/Osmantic/ODS/main/ods/extensions/services/dashboard-api/routers/agents.py#L14-L18】:
curl http://localhost:3002/api/agents/metrics \
  -H "x-api-key: $DASHBOARD_API_KEY"
  • HTML fragments for HTMX (lines 20-64)【/cache/repos/github.com/Osmantic/ODS/main/ods/extensions/services/dashboard-api/routers/agents.py#L20-L64】:
curl http://localhost:3002/api/agents/metrics.html \
  -H "x-api-key: $DASHBOARD_API_KEY"

Agents and Automation: The Hermes Gateway

Agent-related traffic flows through the Agent Gateway (hermes-proxy:9120) depicted in the Agents & Automation block【/cache/repos/github.com/Osmantic/ODS/main/ARCHITECTURE.md#L45-L49】.

After magic-link authentication (handled in ods/extensions/services/dashboard-api/routers/magic_link.py and verified in ods/extensions/services/dashboard-api/routers/auth.py), requests forward to the Hermes container (hermes:9119). Hermes then consumes the OpenAI-compatible endpoint provided by LiteLLM—or directly llama-server—to execute reasoning tasks.

Auxiliary Services: Voice, Search, and RAG

Several supporting containers are reachable only inside the Docker network but indirectly accessed via UI components or the Dashboard-API:

  • Voice pipeline – whisper (port 9000) and tts/Kokoro (port 8880) handle voice I/O when enabled.
  • Search & Research – searxng (port 8888) and perplexica (port 3004) provide web search and deep-research capabilities.
  • RAG pipeline – qdrant (port 6333) and embeddings (port 8090) manage vector storage and embedding generation.
  • Privacy & Observability – privacy-shield (port 8085), token-spy (port 3005), and langfuse (port 3006) enforce PII protection and usage monitoring.

Network Security: Localhost-Only Binding

All services bind explicitly to 127.0.0.1 (localhost) as enforced by the generated port map in lines 88-99 of ARCHITECTURE.md【/cache/repos/github.com/Osmantic/ODS/main/ARCHITECTURE.md#L88-L99】. The Docker-Compose layering script scripts/resolve-compose-stack.sh merges base compose files with GPU-specific overlays and enabled extensions, ensuring the localhost-only topology persists across different hardware configurations【/cache/repos/github.com/Osmantic/ODS/main/ARCHITECTURE.md#L28-L48】.

Summary

  • ODS service communication flow routes all browser traffic through localhost-bound UI containers (open-webui, dashboard, opencode, hermes-proxy).
  • LLM inference supports both direct access to llama-server:8080 and proxied access via litellm:4000 for hybrid cloud fallback.
  • The Dashboard-API (dashboard-api:3002) centralizes feature discovery and agent metrics, with implementations in ods/extensions/services/dashboard-api/routers/agents.py.
  • Agent workflows authenticate via magic-link flows in auth.py and magic_link.py before reaching the Hermes gateway.
  • Auxiliary services for voice, search, RAG, and privacy run on internal ports (9000, 8888, 6333, etc.) accessible only through the Docker network.
  • The entire stack enforces localhost binding (127.0.0.1) through resolve-compose-stack.sh and the architecture definitions in ARCHITECTURE.md.

Frequently Asked Questions

How does ODS handle LLM inference routing?

ODS provides two inference paths. The direct path sends requests from open-webui straight to llama-server:8080. Alternatively, the LiteLLM gateway (litellm:4000) acts as an OpenAI-compatible proxy that routes to local inference first, then falls back to cloud providers if hybrid mode is enabled, as implemented in ods/extensions/services/dashboard-api/routers/lite_llm.py.

What is the role of the Dashboard API in ODS?

The Dashboard-API (dashboard-api:3002) serves as the system-management layer, exposing endpoints like /api/features for service discovery and /api/agents/metrics for telemetry. It reads service manifests from extensions/services/*/manifest.yaml and returns both JSON and HTML fragment responses for HTMX integration, as defined in ods/extensions/services/dashboard-api/routers/agents.py.

How are agent requests authenticated in ODS?

Agent traffic flowing through hermes-proxy:9120 requires magic-link authentication handled by ods/extensions/services/dashboard-api/routers/magic_link.py, with token verification implemented in ods/extensions/services/dashboard-api/routers/auth.py. Once authenticated, requests forward to the Hermes container (hermes:9119) for processing.

Why does ODS bind all services to localhost only?

ODS enforces localhost-only binding (127.0.0.1) to ensure no external network exposure of sensitive AI services, creating a secure-by-default architecture. This constraint is defined in ARCHITECTURE.md lines 88-99 and maintained through the resolve-compose-stack.sh script, which generates the final Docker Compose configuration with restricted port mappings.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →