# Understanding the ODS Service Communication Flow: A Complete Architecture Guide

> Discover the ODS service communication flow. Learn how ODS routes traffic through localhost containers for seamless UI, LLM, and service interaction within its Docker Compose architecture.

- Repository: [Osmantic/ODS](https://github.com/Osmantic/ODS)
- Tags: architecture
- Published: 2026-09-01

---

**ODS routes all traffic through localhost-bound containers, where browser-based UIs communicate with LLM inference, agent gateways, and auxiliary services via a layered Docker Compose stack orchestrated by the Dashboard API.**

The **ODS service communication flow** defines how the Osmantic/ODS (Open-Source AI Stack) orchestrates interactions between user interfaces, inference engines, and supporting microservices. Every component communicates over HTTP endpoints restricted to `127.0.0.1`, creating a secure, single-machine architecture where the Docker Compose layer manages internal routing. Understanding this flow is essential for debugging API calls, extending the stack, or integrating custom agents.

## Browser to UI Services: The Entry Points

All user interactions begin in the browser and terminate at four primary UI containers defined in the *Core Services* block of [`ARCHITECTURE.md`](https://github.com/Osmantic/ODS/blob/main/ARCHITECTURE.md)【/cache/repos/github.com/Osmantic/ODS/main/ARCHITECTURE.md#L24-L30】.

- **`open-webui`** (port 3000) – Serves the main chat interface.
- **`dashboard`** (port 3001) – Provides the control center for system management.
- **`opencode`** (port 3003) – Hosts the web-based IDE.
- **`hermes-proxy`** (port 9120) – Acts as the default gateway for agent interactions.

These services expose ports only on localhost, ensuring no direct external network access bypasses the host machine.

## UI to LLM Inference: Direct and Proxy Paths

When the chat UI sends a generation request, it reaches the **LLM inference layer** through two possible routes as shown in the *Inference Layer* diagram【/cache/repos/github.com/Osmantic/ODS/main/ARCHITECTURE.md#L24-L26】.

### Direct Inference Path

The fastest route connects directly to the inference server:

```bash
curl -X POST http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"local","messages":[{"role":"user","content":"Hello"}]}'

```

This path bypasses intermediaries and hits `llama-server:8080` immediately.

### LiteLLM Gateway Path

For hybrid deployments supporting cloud fallback, requests route through the **LiteLLM gateway** (`litellm:4000`):

```bash
curl -X POST http://localhost:4000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"local","messages":[{"role":"user","content":"Hello"}]}'

```

In this configuration, `open-webui` sends traffic to `http://localhost:4000/v1/chat/completions`, which then proxies to `llama-server` or falls back to a cloud provider if *hybrid* mode is active.

## Dashboard to Dashboard-API: System Management

The `dashboard` UI does not query inference engines directly. Instead, it communicates with the **Dashboard-API** service (`dashboard-api:3002`) for runtime metadata and health monitoring.

### Feature Discovery and Health Checks

The Dashboard-API exposes `/api/features` to list enabled capabilities based on service manifests found in `extensions/services/*/manifest.yaml`. Clients must authenticate using an API key header.

```bash
curl http://localhost:3002/api/features \
  -H "x-api-key: $DASHBOARD_API_KEY"

```

### Agent Metrics Endpoints

Agent telemetry is implemented in [`ods/extensions/services/dashboard-api/routers/agents.py`](https://github.com/Osmantic/ODS/blob/main/ods/extensions/services/dashboard-api/routers/agents.py). This module provides dual-format outputs for the same data:

- **JSON metrics** (lines 14-18)【/cache/repos/github.com/Osmantic/ODS/main/ods/extensions/services/dashboard-api/routers/agents.py#L14-L18】:

```bash
curl http://localhost:3002/api/agents/metrics \
  -H "x-api-key: $DASHBOARD_API_KEY"

```

- **HTML fragments** for HTMX (lines 20-64)【/cache/repos/github.com/Osmantic/ODS/main/ods/extensions/services/dashboard-api/routers/agents.py#L20-L64】:

```bash
curl http://localhost:3002/api/agents/metrics.html \
  -H "x-api-key: $DASHBOARD_API_KEY"

```

## Agents and Automation: The Hermes Gateway

Agent-related traffic flows through the **Agent Gateway** (`hermes-proxy:9120`) depicted in the *Agents & Automation* block【/cache/repos/github.com/Osmantic/ODS/main/ARCHITECTURE.md#L45-L49】.

After magic-link authentication (handled in [`ods/extensions/services/dashboard-api/routers/magic_link.py`](https://github.com/Osmantic/ODS/blob/main/ods/extensions/services/dashboard-api/routers/magic_link.py) and verified in [`ods/extensions/services/dashboard-api/routers/auth.py`](https://github.com/Osmantic/ODS/blob/main/ods/extensions/services/dashboard-api/routers/auth.py)), requests forward to the Hermes container (`hermes:9119`). Hermes then consumes the OpenAI-compatible endpoint provided by LiteLLM—or directly `llama-server`—to execute reasoning tasks.

## Auxiliary Services: Voice, Search, and RAG

Several supporting containers are reachable only inside the Docker network but indirectly accessed via UI components or the Dashboard-API:

- **Voice pipeline** – `whisper` (port 9000) and `tts/Kokoro` (port 8880) handle voice I/O when enabled.
- **Search & Research** – `searxng` (port 8888) and `perplexica` (port 3004) provide web search and deep-research capabilities.
- **RAG pipeline** – `qdrant` (port 6333) and `embeddings` (port 8090) manage vector storage and embedding generation.
- **Privacy & Observability** – `privacy-shield` (port 8085), `token-spy` (port 3005), and `langfuse` (port 3006) enforce PII protection and usage monitoring.

## Network Security: Localhost-Only Binding

All services bind explicitly to `127.0.0.1` (localhost) as enforced by the generated port map in lines 88-99 of [`ARCHITECTURE.md`](https://github.com/Osmantic/ODS/blob/main/ARCHITECTURE.md)【/cache/repos/github.com/Osmantic/ODS/main/ARCHITECTURE.md#L88-L99】. The **Docker-Compose layering** script [`scripts/resolve-compose-stack.sh`](https://github.com/Osmantic/ODS/blob/main/scripts/resolve-compose-stack.sh) merges base compose files with GPU-specific overlays and enabled extensions, ensuring the localhost-only topology persists across different hardware configurations【/cache/repos/github.com/Osmantic/ODS/main/ARCHITECTURE.md#L28-L48】.

## Summary

- **ODS service communication flow** routes all browser traffic through localhost-bound UI containers (`open-webui`, `dashboard`, `opencode`, `hermes-proxy`).
- LLM inference supports both direct access to `llama-server:8080` and proxied access via `litellm:4000` for hybrid cloud fallback.
- The Dashboard-API (`dashboard-api:3002`) centralizes feature discovery and agent metrics, with implementations in [`ods/extensions/services/dashboard-api/routers/agents.py`](https://github.com/Osmantic/ODS/blob/main/ods/extensions/services/dashboard-api/routers/agents.py).
- Agent workflows authenticate via magic-link flows in [`auth.py`](https://github.com/Osmantic/ODS/blob/main/auth.py) and [`magic_link.py`](https://github.com/Osmantic/ODS/blob/main/magic_link.py) before reaching the Hermes gateway.
- Auxiliary services for voice, search, RAG, and privacy run on internal ports (9000, 8888, 6333, etc.) accessible only through the Docker network.
- The entire stack enforces localhost binding (`127.0.0.1`) through [`resolve-compose-stack.sh`](https://github.com/Osmantic/ODS/blob/main/resolve-compose-stack.sh) and the architecture definitions in [`ARCHITECTURE.md`](https://github.com/Osmantic/ODS/blob/main/ARCHITECTURE.md).

## Frequently Asked Questions

### How does ODS handle LLM inference routing?

ODS provides two inference paths. The direct path sends requests from `open-webui` straight to `llama-server:8080`. Alternatively, the **LiteLLM gateway** (`litellm:4000`) acts as an OpenAI-compatible proxy that routes to local inference first, then falls back to cloud providers if hybrid mode is enabled, as implemented in [`ods/extensions/services/dashboard-api/routers/lite_llm.py`](https://github.com/Osmantic/ODS/blob/main/ods/extensions/services/dashboard-api/routers/lite_llm.py).

### What is the role of the Dashboard API in ODS?

The **Dashboard-API** (`dashboard-api:3002`) serves as the system-management layer, exposing endpoints like `/api/features` for service discovery and `/api/agents/metrics` for telemetry. It reads service manifests from `extensions/services/*/manifest.yaml` and returns both JSON and HTML fragment responses for HTMX integration, as defined in [`ods/extensions/services/dashboard-api/routers/agents.py`](https://github.com/Osmantic/ODS/blob/main/ods/extensions/services/dashboard-api/routers/agents.py).

### How are agent requests authenticated in ODS?

Agent traffic flowing through `hermes-proxy:9120` requires **magic-link authentication** handled by [`ods/extensions/services/dashboard-api/routers/magic_link.py`](https://github.com/Osmantic/ODS/blob/main/ods/extensions/services/dashboard-api/routers/magic_link.py), with token verification implemented in [`ods/extensions/services/dashboard-api/routers/auth.py`](https://github.com/Osmantic/ODS/blob/main/ods/extensions/services/dashboard-api/routers/auth.py). Once authenticated, requests forward to the Hermes container (`hermes:9119`) for processing.

### Why does ODS bind all services to localhost only?

ODS enforces **localhost-only binding** (`127.0.0.1`) to ensure no external network exposure of sensitive AI services, creating a secure-by-default architecture. This constraint is defined in [`ARCHITECTURE.md`](https://github.com/Osmantic/ODS/blob/main/ARCHITECTURE.md) lines 88-99 and maintained through the [`resolve-compose-stack.sh`](https://github.com/Osmantic/ODS/blob/main/resolve-compose-stack.sh) script, which generates the final Docker Compose configuration with restricted port mappings.