ODS Architecture Overview: Complete Technical Guide to the Open-Source Distributed Stack

ODS (Open-Source Distributed Stack) is a fully local AI platform that runs on a single machine and provides end-to-end AI workflows—including LLM inference, RAG pipelines, voice I/O, and autonomous agents—using a layered Docker-Compose architecture with GPU-specific overlays.

The Osmantic/ODS repository delivers a self-contained AI infrastructure designed for privacy-first deployment on local hardware. Unlike cloud-dependent platforms, ODS combines inference engines, vector databases, and automation tools into a unified stack orchestrated by a 13-phase installer and CLI-driven lifecycle management.

Two-Layer High-Level Structure

ODS architecture splits logically into an Outer Wrapper and a Core Product.

The Outer Wrapper contains bootstrap resources, installer entry points, and CI definitions. Key files include install.sh, install-core.sh, and .github/workflows/*. This layer handles hardware detection, dependency validation, and initial environment preparation.

The Core Product resides in the ods/ directory and houses all deployable services, runtime helpers, and service manifests. This is where the actual AI workloads execute, managed through declarative compose files and the ods-cli command interface.

Core System Components

The platform organizes functionality into logical service groups defined in ARCHITECTURE.md and visualized via a Mermaid system diagram.

Inference and Gateway Layer

llama-server provides the primary LLM inference backend, exposing OpenAI-compatible endpoints for local model execution. GPU acceleration is selected through overlay files such as docker-compose.nvidia.yml, docker-compose.amd.yml, docker-compose.apple.yml, or docker-compose.cpu.yml.

litellm acts as the gateway proxy (port 8080), routing requests between local models and optional cloud fallbacks while maintaining a unified API interface.

User Interface Layer

open-webui (port 3000) delivers the primary chat interface with integrated image generation, web search, and voice interaction capabilities.

dashboard (port 3001) provides a React/Vite control panel for stack management, while dashboard-api (port 3002) supplies the FastAPI backend that serves operational metrics and health data.

Data and Research Layer

searxng (port 8888) implements privacy-preserving metasearch across multiple engines without telemetry. perplexica (port 3004) adds LLM-augmented research capabilities with citation tracking.

For retrieval-augmented generation, qdrant (port 6333) stores vector embeddings, while the embeddings service (port 8090) generates TEI (Text Embeddings Inference) vectors using dedicated acceleration.

Automation and Voice Layer

hermes (port 9120 via proxy) serves as the default autonomous agent, executing tool calls and workflows. ape enforces policy constraints on agent actions, while n8n provides visual workflow orchestration for complex automation pipelines.

Voice interaction runs through whisper (port 9000) for speech-to-text and tts/Kokoro (port 8880) for text-to-speech, both exposing OpenAI-compatible audio APIs.

Media and Development Layer

comfyui (port 8188) handles image generation via SDXL Lightning, enabling local media synthesis without external APIs.

opencode (port 3003) runs as a host-managed systemd service, providing a web-based IDE for rapid prototyping and stack extension development.

The 13-Phase Installation Pipeline

ODS employs a deterministic installation pipeline orchestrated by install-core.sh. The process is divided into thirteen sequential phases located in ods/installers/phases/, ranging from 01-preflight.sh through 13-summary.sh.

Pure-function libraries in ods/installers/lib/—including detection.sh for hardware probing, tier-map.sh for capability classification, and compose-select.sh for backend selection—support the imperative phase scripts. This architecture ensures idempotent hardware detection, tier assignment, service selection, and final health validation.

Docker-Compose Layering Strategy

The runtime stack uses a merging strategy to accommodate diverse hardware configurations:

The resolver script ods/scripts/resolve-compose-stack.sh discovers enabled extensions, applies the appropriate GPU overlay, and generates the final compose configuration before executing docker compose up -d.

Quick Start Commands

Deploy the stack using the official installer and CLI tools:


# Execute the single-command installer

curl -fsSL https://raw.githubusercontent.com/Osmantic/ODS/main/install.sh | bash

# Start all services via CLI composition resolution

ods-cli start

# Verify service health across all containers

ods-cli health

Access the interfaces at:

  • http://localhost:3000 — Chat UI (open-webui)
  • http://localhost:3001 — Control Panel (dashboard)
  • http://localhost:3003 — Web IDE (opencode)

Summary

  • ODS separates concerns into an Outer Wrapper (installer/CI) and Core Product (ods/ directory), enabling clean updates and modular development.
  • The architecture supports five GPU backends (NVIDIA, AMD, Apple Metal, Intel SYCL, CPU) through overlay compose files and runtime detection.
  • Thirteen installer phases handle everything from hardware detection to health checks, using pure-function libraries in ods/installers/lib/.
  • Services communicate through standard ports (3000-3004, 6333, 8090, 8880-8888, 9000, 9120) with litellm providing unified API proxying.
  • The layered compose model allows selective activation of RAG, voice, agents, and observability tools without modifying base infrastructure.

Frequently Asked Questions

What hardware accelerators does ODS support?

ODS automatically detects and configures NVIDIA CUDA, AMD ROCm, Apple Metal, Intel SYCL, and standard CPU backends. The installer selects the appropriate overlay file (e.g., docker-compose.nvidia.yml) based on hardware probes defined in ods/installers/lib/detection.sh.

How does the ODS CLI manage service lifecycles?

The ods-cli command reads service manifests from ods/extensions/services/*/manifest.yaml and delegates to ods/scripts/resolve-compose-stack.sh to merge compose files. It then orchestrates docker compose operations and executes health checks defined in ods/installers/phases/12-health.sh.

Can I disable specific services like voice or image generation?

Yes. The layered architecture allows selective enablement of extensions. Remove or disable the corresponding entries in the extensions directory before running ods-cli start, and the resolver will exclude them from the final compose stack.

Where does ODS store configuration and model data?

Configuration writers defined in ods/docs/INSTALLER-ARCHITECTURE.md persist settings to host directories mounted into containers. The installer creates these paths during phase execution, ensuring data persists across container restarts while keeping sensitive information local to your machine.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →