# How to Deploy LLM Inference Locally with ODS: Complete Setup Guide

> Deploy LLM inference locally with ODS. This guide shows how to use the Osmantic Deployment System for a simple, single-command setup with auto-detected hardware and Open WebUI.

- Repository: [Osmantic/ODS](https://github.com/Osmantic/ODS)
- Tags: how-to-guide
- Published: 2026-09-01

---

**ODS (Osmantic Deployment System) provides a turnkey, single-command installer that auto-detects your hardware, selects an optimized GGUF model, and launches llama-server with a full Open WebUI dashboard for local LLM inference.**

The Osmantic/ODS repository delivers a production-ready framework for deploying large language models on consumer hardware without cloud dependencies. Whether you are running Linux, macOS, or Windows, ODS handles GPU detection, model curation, and service orchestration through Docker and native backends. This guide walks you through how to deploy LLM inference locally with ODS using the official installer scripts and CLI tools.

## ODS Architecture for Local LLM Inference

### Hardware Detection and Tier Mapping

The installation orchestration begins in [`install.sh`](https://github.com/Osmantic/ODS/blob/main/install.sh) and [`install-core.sh`](https://github.com/Osmantic/ODS/blob/main/install-core.sh), which delegate hardware introspection to [`installers/lib/detection.sh`](https://github.com/Osmantic/ODS/blob/main/installers/lib/detection.sh). This script probes for NVIDIA, AMD, or Apple Silicon GPUs and maps capabilities to predefined performance tiers via [`installers/lib/tier-map.sh`](https://github.com/Osmantic/ODS/blob/main/installers/lib/tier-map.sh). The installer writes these selections to your `.env` file, setting variables like `LLM_MODEL` and `GGUF_FILE` (e.g., `qwen3.5-27b`) along with context window parameters such as `CTX_SIZE` or `MAX_CONTEXT`.

### Inference Engine and Port Configuration

At the heart of the stack is **llama-server**, a high-performance inference binary based on llama.cpp. On Linux, ODS exposes the server at **host port 11434** (mapped to container port 8080), while macOS and Windows native deployments bind directly to **port 8080**. The service loads the quantized GGUF model specified in [`config/model-library.json`](https://github.com/Osmantic/ODS/blob/main/config/model-library.json) and serves an OpenAI-compatible API endpoint.

### Dashboard and API Gateway

The complete stack includes **Open WebUI** (accessible at `http://localhost:3000`) for conversational interaction, backed by a LiteLLM gateway on port 4000. The `dashboard-api` component provides granular controls for model management, GPU metrics monitoring, and agent orchestration for extensions like Hermes and OpenClaw.

## Prerequisites

- **Linux/macOS**: `curl`, `jq` (auto-installed by the script), and Docker Engine ≥ 20.10
- **Windows**: Docker Desktop with WSL2 backend and PowerShell 5.0 or higher

## Installation Steps

### Linux and macOS (One-Line Installer)

Execute the bootstrap script to clone the `main` branch, detect hardware, and initialize services:

```bash
curl -fsSL https://install.osmantic.com/ods.sh | bash

```

The script terminates when all containers are healthy. Navigate to **http://localhost:3000** to access the Open WebUI dashboard.

### macOS Apple Silicon (Metal Acceleration)

For Apple Silicon Macs, running [`./install.sh`](https://github.com/Osmantic/ODS/blob/main/./install.sh) explicitly triggers the Metal codepath. The installer launches `llama-server` as a native process for GPU acceleration while keeping auxiliary services containerized.

```bash
./install.sh

```

### Windows (PowerShell and WSL2)

On Windows, download and expand the repository archive, then execute the PowerShell installer:

```powershell
$odsSrc = Join-Path $env:TEMP ("ods-install-" + [guid]::NewGuid().ToString("N"))
$odsZip = Join-Path $odsSrc "ods-main.zip"
New-Item -ItemType Directory -Path $odsSrc | Out-Null
Invoke-WebRequest "https://github.com/Osmantic/ODS/archive/refs/heads/main.zip" -OutFile $odsZip
Expand-Archive -LiteralPath $odsZip -DestinationPath $odsSrc -Force
cd (Get-ChildItem -LiteralPath $odsSrc -Directory | Select-Object -First 1).FullName
Set-ExecutionPolicy -Scope Process -ExecutionPolicy Bypass
.\install.ps1

```

The installer provisions the runtime under `$env:USERPROFILE\ods` and exposes the native `llama-server` on **port 8080**.

## Running LLM Inference Locally

Verify that the inference endpoint is healthy using the platform-specific port:

```bash

# Linux (Docker host port)

curl http://localhost:11434/health

# macOS/Windows (Native)

curl http://localhost:8080/health

```

Submit chat completions using the OpenAI API schema:

```bash
curl -X POST http://localhost:11434/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"ods/current","messages":[{"role":"user","content":"Explain quantum computing"}]}'

```

## Managing Models in ODS

### Dashboard Method (Recommended)

Access the Models pane in Open WebUI to browse the ODS catalog or search Hugging Face registries. Clicking **Download** followed by **Load** automatically updates `~/ods/.env`, syncs [`config/llama-server/models.ini`](https://github.com/Osmantic/ODS/blob/main/config/llama-server/models.ini), and restarts the inference container without manual CLI intervention.

### CLI Method

Use the `ods` binary for headless administration:

```bash

# Display active model configuration

ods model current

# View hardware-tier mappings

ods model list

# Swap to tier 4 (higher capability model)

ods model swap T4

```

These commands execute the same activation transaction as the Dashboard, ensuring `.env` and service configs remain consistent.

### Manual Configuration

If the UI is inaccessible, copy a compatible GGUF file into `~/ods/data/models/`, then edit `.env` or run `ods config edit` to update `LLM_MODEL` paths. Modify [`config/llama-server/models.ini`](https://github.com/Osmantic/ODS/blob/main/config/llama-server/models.ini) if adding custom context sizes, and apply changes with `ods restart llama-server`.

## Example Deployment Workflow

```bash

# 1. Install ODS on Ubuntu

curl -fsSL https://install.osmantic.com/ods.sh | bash

# 2. Verify GPU detection

ods status | grep GPU

# 3. Select high-performance tier

ods model swap T4

# 4. Validate running model

ods model current
curl http://localhost:11434/v1/models

```

## Key Source Files and Configuration

Understanding these paths aids in troubleshooting and customization:

- [`install.sh`](https://github.com/Osmantic/ODS/blob/main/install.sh) – Top-level entry point that orchestrates the 13-phase installation
- [`installers/lib/detection.sh`](https://github.com/Osmantic/ODS/blob/main/installers/lib/detection.sh) – GPU/CPU capability detection logic
- [`installers/lib/tier-map.sh`](https://github.com/Osmantic/ODS/blob/main/installers/lib/tier-map.sh) – Maps hardware profiles to model tiers
- [`docker-compose.base.yml`](https://github.com/Osmantic/ODS/blob/main/docker-compose.base.yml) – Core service definitions for llama-server and Open WebUI
- [`config/model-library.json`](https://github.com/Osmantic/ODS/blob/main/config/model-library.json) – Curated GGUF model metadata and download URLs
- [`scripts/resolve-compose-stack.sh`](https://github.com/Osmantic/ODS/blob/main/scripts/resolve-compose-stack.sh) – Merges extension manifests into the final Docker Compose configuration

## Summary

- **ODS** automates hardware detection and GGUF model selection through scripts in `installers/lib/`.
- **llama-server** provides the inference engine, exposed on port 11434 (Linux) or 8080 (macOS/Windows).
- The **Open WebUI** dashboard on port 3000 offers model management, while the `ods` CLI supports headless operations.
- Installation requires a single command on Linux/macOS or a PowerShell script on Windows WSL2.
- Model swaps update `.env` and [`models.ini`](https://github.com/Osmantic/ODS/blob/main/models.ini) atomically, ensuring service consistency.

## Frequently Asked Questions

### What hardware does ODS support for local LLM inference?

ODS supports NVIDIA CUDA GPUs, AMD ROCm accelerators, and Apple Silicon via Metal. The [`installers/lib/detection.sh`](https://github.com/Osmantic/ODS/blob/main/installers/lib/detection.sh) script identifies your hardware during installation and maps it to an appropriate performance tier using [`tier-map.sh`](https://github.com/Osmantic/ODS/blob/main/tier-map.sh), selecting quantized GGUF models that fit your VRAM or unified memory constraints.

### Can I use my own GGUF models instead of the ODS catalog?

Yes. Place your custom GGUF files in `~/ods/data/models/`, then update the `LLM_MODEL` and `GGUF_FILE` variables in `.env` or run `ods config edit`. You must also update [`config/llama-server/models.ini`](https://github.com/Osmantic/ODS/blob/main/config/llama-server/models.ini) with the new model metadata and restart the service using `ods restart llama-server`.

### How do I access the LLM API from external applications?

The inference endpoint follows the OpenAI REST specification. On Linux, target `http://localhost:11434/v1/chat/completions`; on macOS/Windows native deployments, use `http://localhost:8080`. The LiteLLM gateway on port 4000 provides additional routing and load balancing capabilities for multi-model setups.

### Where are ODS configuration files stored?

ODS stores runtime configuration in `~/ods/.env`, model definitions in `~/ods/config/model-library.json`, and service orchestration files in the repository root. The `ods` CLI provides shortcuts like `ods config edit` to modify these files safely without manual path navigation.