Core Components of the ODS Architecture: A Technical Deep Dive into Osmantic/ODS
TLDR: The Osmantic/ODS (Open-Source Data Stack) is a two-layer system comprising a 13-phase installer wrapper and a core product built on a Layered Compose Model, Core Services mesh, and modular Installer Architecture that deploys GPU-accelerated AI services to localhost.
ODS provides a self-hosted AI infrastructure that binds all services to 127.0.0.1 for security. According to the ARCHITECTURE.md specification in the Osmantic/ODS repository, the system uses Docker Compose overlays to support heterogeneous hardware from NVIDIA GPUs to Apple Silicon. Understanding the core components of the ODS architecture reveals how it maintains modularity across its deterministic installer pipeline and service orchestration.
Layered Compose Model
The Layered Compose Model enables hardware-agnostic deployment through file composition rather than monolithic configuration. At the base, docker-compose.base.yml defines immutable core services including llama-server, open-webui, and dashboard-api.
Hardware acceleration is injected via GPU overlay files (docker-compose.nvidia.yml, docker-compose.amd.yml, docker-compose.apple.yml, docker-compose.cpu.yml). These replace container images and runtime settings based on the detected backend.
Optional capabilities enter through extension compose files located in extensions/services/*/compose*.yaml. Services like comfyui, n8n, and openclaw are added without modifying core definitions.
The resolver script (scripts/resolve-compose-stack.sh) merges these layers into a final deployable stack. It discovers enabled extensions, selects the appropriate GPU overlay, and outputs a unified compose file.
# Manually resolve and inspect the final Docker Compose stack
scripts/resolve-compose-stack.sh > full-stack.yml
docker compose -f full-stack.yml config
Core Services and Localhost Service Mesh
The core services layer exposes twenty-two specialized containers, each binding to 127.0.0.1 to preserve local-only access. These communicate via HTTP/REST or OpenAI-compatible APIs.
Key infrastructure components include:
- llama-server (port 8080): The GPU-accelerated LLM inference engine
- litellm (port 4000): OpenAI-compatible proxy enabling hybrid cloud fallback
- dashboard-api (port 3002): FastAPI backend exposing setup, agents, and privacy endpoints
- qdrant (port 6333): Vector database for RAG pipelines
- embeddings (port 8090): HuggingFace TEI for vector generation
User-facing interfaces include open-webui (port 3000) for chat and dashboard (port 3001) for system control. Voice processing runs through whisper (port 9000) and tts/kokoro (port 8880), while comfyui (port 8188) handles media generation. Additional services like perplexica (port 3004) for deep research, langfuse (port 3006) for observability, and opencode (port 3003) as a web IDE complete the stack.
The gateway (litellm) routes OpenAI-style requests to llama-server and can fallback to cloud APIs when local models are insufficient.
# Verify that core services are healthy after startup
ods-cli health
Installer Architecture
The installer architecture implements a deterministic 13-phase pipeline orchestrated by install-core.sh. It strictly separates pure functions from imperative side effects to ensure testability.
Pure-function libraries reside in installers/lib/ and handle GPU detection, tier mapping, and service registry lookups without side effects. These include detection.sh, tier-map.sh, and service-registry.sh.
Sequential phases live in installers/phases/ as numbered scripts (01-preflight.sh through 13-summary.sh). Each phase consumes state from previous steps via exported environment variables.
Critical phases include:
- Preflight: OS and dependency validation
- Detection: GPU hardware detection and tier assignment
- Features: Interactive selection of voice, RAG, and media components
- Docker: Installation of Docker Engine, Compose, and NVIDIA toolkit
- Images: Pulling container images for the detected hardware tier
- Services: Model file generation, stack launch, and health verification
# Install ODS (runs the full 13-phase installer pipeline)
curl -fsSL https://raw.githubusercontent.com/Osmantic/ODS/main/install.sh -o install.sh
bash install.sh
Summary
- Layered Compose Model: Uses base, GPU overlay, and extension compose files merged by
scripts/resolve-compose-stack.shto support hardware from CPUs to NVIDIA GPUs. - Core Services: Twenty-two localhost-bound services exposing ports 3000-9000+ for LLM inference, RAG, voice, and media generation, coordinated via the
litellmgateway. - Installer Architecture: A 13-phase pipeline in
install-core.shseparating pure logic (installers/lib/) from imperative phases (installers/phases/). - Security Model: All services bind to
127.0.0.1only, ensuring local-only access to sensitive AI infrastructure.
Frequently Asked Questions
What is the ODS (Open-Source Data Stack)?
ODS is a self-hosted AI infrastructure platform maintained by Osmantic that deploys a complete stack of large language model services, RAG components, and automation tools to your local machine. According to the repository's ARCHITECTURE.md, it is designed as a two-layer system combining an installer wrapper with deployable core services that run entirely on localhost.
How does ODS handle different GPU backends?
ODS uses a Layered Compose Model where docker-compose.base.yml defines core services and hardware-specific overlays (docker-compose.nvidia.yml, docker-compose.amd.yml, docker-compose.apple.yml, docker-compose.cpu.yml) modify container runtimes and images. The scripts/resolve-compose-stack.sh resolver automatically selects the appropriate overlay during installation based on the detection logic in installers/lib/detection.sh.
How do I add optional services to the ODS stack?
Optional services are added via the extensions directory at extensions/services/*/. Each extension contains a manifest.yaml service contract and compose files that the resolver merges into the final stack. You enable these during the Features phase of installation or by manually configuring the resolver to include specific extension paths before running docker compose up -d.
Is ODS suitable for production environments?
ODS is architected for local-only deployment with all services binding to 127.0.0.1, making it ideal for development workstations and secure local inference rather than exposed production infrastructure. While the deterministic 13-phase installer and layered compose structure provide reliability, the localhost binding and lack of built-in remote access controls mean additional networking layers would be required for internet-facing production use.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →