How to Deploy LLM Inference Locally with ODS: Complete Setup Guide
ODS (Osmantic Deployment System) provides a turnkey, single-command installer that auto-detects your hardware, selects an optimized GGUF model, and launches llama-server with a full Open WebUI dashboard for local LLM inference.
The Osmantic/ODS repository delivers a production-ready framework for deploying large language models on consumer hardware without cloud dependencies. Whether you are running Linux, macOS, or Windows, ODS handles GPU detection, model curation, and service orchestration through Docker and native backends. This guide walks you through how to deploy LLM inference locally with ODS using the official installer scripts and CLI tools.
ODS Architecture for Local LLM Inference
Hardware Detection and Tier Mapping
The installation orchestration begins in install.sh and install-core.sh, which delegate hardware introspection to installers/lib/detection.sh. This script probes for NVIDIA, AMD, or Apple Silicon GPUs and maps capabilities to predefined performance tiers via installers/lib/tier-map.sh. The installer writes these selections to your .env file, setting variables like LLM_MODEL and GGUF_FILE (e.g., qwen3.5-27b) along with context window parameters such as CTX_SIZE or MAX_CONTEXT.
Inference Engine and Port Configuration
At the heart of the stack is llama-server, a high-performance inference binary based on llama.cpp. On Linux, ODS exposes the server at host port 11434 (mapped to container port 8080), while macOS and Windows native deployments bind directly to port 8080. The service loads the quantized GGUF model specified in config/model-library.json and serves an OpenAI-compatible API endpoint.
Dashboard and API Gateway
The complete stack includes Open WebUI (accessible at http://localhost:3000) for conversational interaction, backed by a LiteLLM gateway on port 4000. The dashboard-api component provides granular controls for model management, GPU metrics monitoring, and agent orchestration for extensions like Hermes and OpenClaw.
Prerequisites
- Linux/macOS:
curl,jq(auto-installed by the script), and Docker Engine ≥ 20.10 - Windows: Docker Desktop with WSL2 backend and PowerShell 5.0 or higher
Installation Steps
Linux and macOS (One-Line Installer)
Execute the bootstrap script to clone the main branch, detect hardware, and initialize services:
curl -fsSL https://install.osmantic.com/ods.sh | bash
The script terminates when all containers are healthy. Navigate to http://localhost:3000 to access the Open WebUI dashboard.
macOS Apple Silicon (Metal Acceleration)
For Apple Silicon Macs, running ./install.sh explicitly triggers the Metal codepath. The installer launches llama-server as a native process for GPU acceleration while keeping auxiliary services containerized.
./install.sh
Windows (PowerShell and WSL2)
On Windows, download and expand the repository archive, then execute the PowerShell installer:
$odsSrc = Join-Path $env:TEMP ("ods-install-" + [guid]::NewGuid().ToString("N"))
$odsZip = Join-Path $odsSrc "ods-main.zip"
New-Item -ItemType Directory -Path $odsSrc | Out-Null
Invoke-WebRequest "https://github.com/Osmantic/ODS/archive/refs/heads/main.zip" -OutFile $odsZip
Expand-Archive -LiteralPath $odsZip -DestinationPath $odsSrc -Force
cd (Get-ChildItem -LiteralPath $odsSrc -Directory | Select-Object -First 1).FullName
Set-ExecutionPolicy -Scope Process -ExecutionPolicy Bypass
.\install.ps1
The installer provisions the runtime under $env:USERPROFILE\ods and exposes the native llama-server on port 8080.
Running LLM Inference Locally
Verify that the inference endpoint is healthy using the platform-specific port:
# Linux (Docker host port)
curl http://localhost:11434/health
# macOS/Windows (Native)
curl http://localhost:8080/health
Submit chat completions using the OpenAI API schema:
curl -X POST http://localhost:11434/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"ods/current","messages":[{"role":"user","content":"Explain quantum computing"}]}'
Managing Models in ODS
Dashboard Method (Recommended)
Access the Models pane in Open WebUI to browse the ODS catalog or search Hugging Face registries. Clicking Download followed by Load automatically updates ~/ods/.env, syncs config/llama-server/models.ini, and restarts the inference container without manual CLI intervention.
CLI Method
Use the ods binary for headless administration:
# Display active model configuration
ods model current
# View hardware-tier mappings
ods model list
# Swap to tier 4 (higher capability model)
ods model swap T4
These commands execute the same activation transaction as the Dashboard, ensuring .env and service configs remain consistent.
Manual Configuration
If the UI is inaccessible, copy a compatible GGUF file into ~/ods/data/models/, then edit .env or run ods config edit to update LLM_MODEL paths. Modify config/llama-server/models.ini if adding custom context sizes, and apply changes with ods restart llama-server.
Example Deployment Workflow
# 1. Install ODS on Ubuntu
curl -fsSL https://install.osmantic.com/ods.sh | bash
# 2. Verify GPU detection
ods status | grep GPU
# 3. Select high-performance tier
ods model swap T4
# 4. Validate running model
ods model current
curl http://localhost:11434/v1/models
Key Source Files and Configuration
Understanding these paths aids in troubleshooting and customization:
install.sh– Top-level entry point that orchestrates the 13-phase installationinstallers/lib/detection.sh– GPU/CPU capability detection logicinstallers/lib/tier-map.sh– Maps hardware profiles to model tiersdocker-compose.base.yml– Core service definitions for llama-server and Open WebUIconfig/model-library.json– Curated GGUF model metadata and download URLsscripts/resolve-compose-stack.sh– Merges extension manifests into the final Docker Compose configuration
Summary
- ODS automates hardware detection and GGUF model selection through scripts in
installers/lib/. - llama-server provides the inference engine, exposed on port 11434 (Linux) or 8080 (macOS/Windows).
- The Open WebUI dashboard on port 3000 offers model management, while the
odsCLI supports headless operations. - Installation requires a single command on Linux/macOS or a PowerShell script on Windows WSL2.
- Model swaps update
.envandmodels.iniatomically, ensuring service consistency.
Frequently Asked Questions
What hardware does ODS support for local LLM inference?
ODS supports NVIDIA CUDA GPUs, AMD ROCm accelerators, and Apple Silicon via Metal. The installers/lib/detection.sh script identifies your hardware during installation and maps it to an appropriate performance tier using tier-map.sh, selecting quantized GGUF models that fit your VRAM or unified memory constraints.
Can I use my own GGUF models instead of the ODS catalog?
Yes. Place your custom GGUF files in ~/ods/data/models/, then update the LLM_MODEL and GGUF_FILE variables in .env or run ods config edit. You must also update config/llama-server/models.ini with the new model metadata and restart the service using ods restart llama-server.
How do I access the LLM API from external applications?
The inference endpoint follows the OpenAI REST specification. On Linux, target http://localhost:11434/v1/chat/completions; on macOS/Windows native deployments, use http://localhost:8080. The LiteLLM gateway on port 4000 provides additional routing and load balancing capabilities for multi-model setups.
Where are ODS configuration files stored?
ODS stores runtime configuration in ~/ods/.env, model definitions in ~/ods/config/model-library.json, and service orchestration files in the repository root. The ods CLI provides shortcuts like ods config edit to modify these files safely without manual path navigation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →