# ODS Architecture Overview: Complete Technical Guide to the Open-Source Distributed Stack

> Explore the ODS architecture overview, a complete technical guide to the Open-Source Distributed Stack. Learn about its layered Docker-Compose setup for local end-to-end AI workflows on a single machine.

- Repository: [Osmantic/ODS](https://github.com/Osmantic/ODS)
- Tags: architecture
- Published: 2026-09-01

---

**ODS (Open-Source Distributed Stack) is a fully local AI platform that runs on a single machine and provides end-to-end AI workflows—including LLM inference, RAG pipelines, voice I/O, and autonomous agents—using a layered Docker-Compose architecture with GPU-specific overlays.**

The Osmantic/ODS repository delivers a self-contained AI infrastructure designed for privacy-first deployment on local hardware. Unlike cloud-dependent platforms, ODS combines inference engines, vector databases, and automation tools into a unified stack orchestrated by a 13-phase installer and CLI-driven lifecycle management.

## Two-Layer High-Level Structure

ODS architecture splits logically into an **Outer Wrapper** and a **Core Product**.

The **Outer Wrapper** contains bootstrap resources, installer entry points, and CI definitions. Key files include [`install.sh`](https://github.com/Osmantic/ODS/blob/main/install.sh), [`install-core.sh`](https://github.com/Osmantic/ODS/blob/main/install-core.sh), and `.github/workflows/*`. This layer handles hardware detection, dependency validation, and initial environment preparation.

The **Core Product** resides in the `ods/` directory and houses all deployable services, runtime helpers, and service manifests. This is where the actual AI workloads execute, managed through declarative compose files and the `ods-cli` command interface.

## Core System Components

The platform organizes functionality into logical service groups defined in [`ARCHITECTURE.md`](https://github.com/Osmantic/ODS/blob/main/ARCHITECTURE.md) and visualized via a Mermaid system diagram.

### Inference and Gateway Layer

**`llama-server`** provides the primary LLM inference backend, exposing OpenAI-compatible endpoints for local model execution. GPU acceleration is selected through overlay files such as [`docker-compose.nvidia.yml`](https://github.com/Osmantic/ODS/blob/main/docker-compose.nvidia.yml), [`docker-compose.amd.yml`](https://github.com/Osmantic/ODS/blob/main/docker-compose.amd.yml), [`docker-compose.apple.yml`](https://github.com/Osmantic/ODS/blob/main/docker-compose.apple.yml), or [`docker-compose.cpu.yml`](https://github.com/Osmantic/ODS/blob/main/docker-compose.cpu.yml).

**`litellm`** acts as the gateway proxy (port 8080), routing requests between local models and optional cloud fallbacks while maintaining a unified API interface.

### User Interface Layer

**`open-webui`** (port 3000) delivers the primary chat interface with integrated image generation, web search, and voice interaction capabilities.

**`dashboard`** (port 3001) provides a React/Vite control panel for stack management, while **`dashboard-api`** (port 3002) supplies the FastAPI backend that serves operational metrics and health data.

### Data and Research Layer

**`searxng`** (port 8888) implements privacy-preserving metasearch across multiple engines without telemetry. **`perplexica`** (port 3004) adds LLM-augmented research capabilities with citation tracking.

For retrieval-augmented generation, **`qdrant`** (port 6333) stores vector embeddings, while the **`embeddings`** service (port 8090) generates TEI (Text Embeddings Inference) vectors using dedicated acceleration.

### Automation and Voice Layer

**`hermes`** (port 9120 via proxy) serves as the default autonomous agent, executing tool calls and workflows. **`ape`** enforces policy constraints on agent actions, while **`n8n`** provides visual workflow orchestration for complex automation pipelines.

Voice interaction runs through **`whisper`** (port 9000) for speech-to-text and **`tts/Kokoro`** (port 8880) for text-to-speech, both exposing OpenAI-compatible audio APIs.

### Media and Development Layer

**`comfyui`** (port 8188) handles image generation via SDXL Lightning, enabling local media synthesis without external APIs.

**`opencode`** (port 3003) runs as a host-managed systemd service, providing a web-based IDE for rapid prototyping and stack extension development.

## The 13-Phase Installation Pipeline

ODS employs a deterministic installation pipeline orchestrated by [`install-core.sh`](https://github.com/Osmantic/ODS/blob/main/install-core.sh). The process is divided into thirteen sequential phases located in `ods/installers/phases/`, ranging from [`01-preflight.sh`](https://github.com/Osmantic/ODS/blob/main/01-preflight.sh) through [`13-summary.sh`](https://github.com/Osmantic/ODS/blob/main/13-summary.sh).

Pure-function libraries in `ods/installers/lib/`—including [`detection.sh`](https://github.com/Osmantic/ODS/blob/main/detection.sh) for hardware probing, [`tier-map.sh`](https://github.com/Osmantic/ODS/blob/main/tier-map.sh) for capability classification, and [`compose-select.sh`](https://github.com/Osmantic/ODS/blob/main/compose-select.sh) for backend selection—support the imperative phase scripts. This architecture ensures idempotent hardware detection, tier assignment, service selection, and final health validation.

## Docker-Compose Layering Strategy

The runtime stack uses a merging strategy to accommodate diverse hardware configurations:

- **[`docker-compose.base.yml`](https://github.com/Osmantic/ODS/blob/main/docker-compose.base.yml)** defines core services and network topologies
- **GPU overlays** ([`docker-compose.nvidia.yml`](https://github.com/Osmantic/ODS/blob/main/docker-compose.nvidia.yml), [`docker-compose.amd.yml`](https://github.com/Osmantic/ODS/blob/main/docker-compose.amd.yml), [`docker-compose.apple.yml`](https://github.com/Osmantic/ODS/blob/main/docker-compose.apple.yml), [`docker-compose.cpu.yml`](https://github.com/Osmantic/ODS/blob/main/docker-compose.cpu.yml)) inject accelerator-specific container images
- **Extension manifests** (`ods/extensions/services/*/manifest.yaml`) register optional services like `langfuse` or `privacy-shield`

The resolver script [`ods/scripts/resolve-compose-stack.sh`](https://github.com/Osmantic/ODS/blob/main/ods/scripts/resolve-compose-stack.sh) discovers enabled extensions, applies the appropriate GPU overlay, and generates the final compose configuration before executing `docker compose up -d`.

## Quick Start Commands

Deploy the stack using the official installer and CLI tools:

```bash

# Execute the single-command installer

curl -fsSL https://raw.githubusercontent.com/Osmantic/ODS/main/install.sh | bash

# Start all services via CLI composition resolution

ods-cli start

# Verify service health across all containers

ods-cli health

```

Access the interfaces at:
- `http://localhost:3000` — Chat UI (`open-webui`)
- `http://localhost:3001` — Control Panel (`dashboard`)
- `http://localhost:3003` — Web IDE (`opencode`)

## Summary

- ODS separates concerns into an **Outer Wrapper** (installer/CI) and **Core Product** (`ods/` directory), enabling clean updates and modular development.
- The architecture supports **five GPU backends** (NVIDIA, AMD, Apple Metal, Intel SYCL, CPU) through overlay compose files and runtime detection.
- **Thirteen installer phases** handle everything from hardware detection to health checks, using pure-function libraries in `ods/installers/lib/`.
- Services communicate through standard ports (3000-3004, 6333, 8090, 8880-8888, 9000, 9120) with `litellm` providing unified API proxying.
- The **layered compose model** allows selective activation of RAG, voice, agents, and observability tools without modifying base infrastructure.

## Frequently Asked Questions

### What hardware accelerators does ODS support?

ODS automatically detects and configures NVIDIA CUDA, AMD ROCm, Apple Metal, Intel SYCL, and standard CPU backends. The installer selects the appropriate overlay file (e.g., [`docker-compose.nvidia.yml`](https://github.com/Osmantic/ODS/blob/main/docker-compose.nvidia.yml)) based on hardware probes defined in [`ods/installers/lib/detection.sh`](https://github.com/Osmantic/ODS/blob/main/ods/installers/lib/detection.sh).

### How does the ODS CLI manage service lifecycles?

The `ods-cli` command reads service manifests from `ods/extensions/services/*/manifest.yaml` and delegates to [`ods/scripts/resolve-compose-stack.sh`](https://github.com/Osmantic/ODS/blob/main/ods/scripts/resolve-compose-stack.sh) to merge compose files. It then orchestrates `docker compose` operations and executes health checks defined in [`ods/installers/phases/12-health.sh`](https://github.com/Osmantic/ODS/blob/main/ods/installers/phases/12-health.sh).

### Can I disable specific services like voice or image generation?

Yes. The layered architecture allows selective enablement of extensions. Remove or disable the corresponding entries in the extensions directory before running `ods-cli start`, and the resolver will exclude them from the final compose stack.

### Where does ODS store configuration and model data?

Configuration writers defined in [`ods/docs/INSTALLER-ARCHITECTURE.md`](https://github.com/Osmantic/ODS/blob/main/ods/docs/INSTALLER-ARCHITECTURE.md) persist settings to host directories mounted into containers. The installer creates these paths during phase execution, ensuring data persists across container restarts while keeping sensitive information local to your machine.