# Colibri Deployment Strategies: 5 Production-Ready Methods for Inference Workloads

> Discover 5 production-ready Colibri deployment strategies for inference workloads. Explore native CLI to Kubernetes orchestration for your C++ inference engine.

- Repository: [Vincenzo Fornaro/colibri](https://github.com/JustVugg/colibri)
- Tags: best-practices
- Published: 2026-09-12

---

**Colibri supports five distinct deployment strategies ranging from native CLI execution to Kubernetes orchestration, each leveraging the architectural separation between the C-based inference engine and the protocol-speaking server components.**

Colibri is a flexible, high-performance inference engine designed for diverse deployment scenarios. Its architecture separates the **engine** (native C code performing token generation) from the **server** (process speaking a line-oriented protocol over stdin/stdout), enabling deployment across environments from local laptops to cloud-scale clusters. According to the JustVugg/colibri source code, the engine compiles for multiple GPU backends including CUDA, Metal, and CPU-only configurations, making it adaptable to various infrastructure requirements.

## Native CLI Deployment

The simplest deployment strategy runs the `colibri` binary directly on the host system. This approach minimizes overhead and eliminates container runtime dependencies, making it ideal for rapid prototyping and development workflows.

The native CLI supports two primary commands defined in [`colibri/cli.py`](https://github.com/JustVugg/colibri/blob/main/colibri/cli.py):

- **`colibri serve`** – Starts the server using the modern **mux protocol** for multiplexed request handling
- **`colibri chat`** – Uses the legacy protocol for backward compatibility

The mux protocol, documented in [`docs/serve_protocol.md`](https://github.com/JustVugg/colibri/blob/main/docs/serve_protocol.md), enables continuous batching and concurrent request processing through line-oriented messages over stdin/stdout.

```bash

# Start the engine in mux mode for continuous batching

colibri serve --batch

# Submit a request via stdin (example format)

echo -e "SUBMIT 1 0 12 64 0.7 0.9\nHello world!" | colibri serve --batch

```

## Docker Container Deployment

For production environments requiring isolation and reproducibility, the Docker strategy packages the compiled engine with a minimal runtime. The official image exposes both the HTTP API at `/v1/chat/completions` and a single-page application frontend.

The Dockerfile located at `docker/Dockerfile` handles the multi-stage build process, compiling the C sources from [`c/colibri.c`](https://github.com/JustVugg/colibri/blob/main/c/colibri.c) and bundling the Tauri desktop build for GUI-capable environments.

```bash

# Pull the official image

docker pull justvugg/colibri:latest

# Run with GPU support and model mounting

docker run -d -p 8000:8000 \
  -e COLI_MODEL=/models/colibri-7b.q4_K_M.bin \
  -v ./models:/models \
  justvugg/colibri:latest

# Test the OpenAI-compatible endpoint

curl -X POST http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"colibri","messages":[{"role":"user","content":"Hello"}]}'

```

## Docker Compose and Kubernetes Orchestration

Scalable deployments utilize [`docker/docker-compose.yml`](https://github.com/JustVugg/colibri/blob/main/docker/docker-compose.yml) to orchestrate multiple services including reverse proxies and persistent storage. This strategy supports horizontal scaling and high-availability configurations through health checks and service dependencies.

The compose file includes configurations for:
- Persistent volume mounting for the `.coli_usage` telemetry file
- Nginx reverse proxy for load balancing
- Environment variable management for GPU backend selection

```yaml

# docker/docker-compose.yml

version: "3.9"
services:
  colibri:
    image: justvugg/colibri:latest
    ports:
      - "8000:8000"
    environment:
      - COLI_MODEL=/models/colibri-7b.q4_K_M.bin
    volumes:
      - ./models:/models
      - ./telemetry:/app/.coli_usage
    restart: unless-stopped

  nginx:
    image: nginx:alpine
    ports:
      - "80:80"
    volumes:
      - ./nginx.conf:/etc/nginx/nginx.conf:ro
    depends_on:
      - colibri

```

```bash

# Deploy the stack

docker compose -f docker/docker-compose.yml up -d

```

For Kubernetes environments, the compose file can be converted to Helm charts using kompose or manual translation, enabling auto-scaling and rolling updates.

## Tauri Desktop Application

End-users requiring a graphical interface without Docker can deploy the **Tauri Desktop App**, which embeds the inference engine within a native Windows or macOS binary. The desktop build compiles the React/TypeScript frontend into a native application that exposes the same HTTP endpoint on `localhost`.

Build instructions in [`desktop/README.md`](https://github.com/JustVugg/colibri/blob/main/desktop/README.md) specify the Rust-based compilation process:

```bash

# Build the desktop client (requires Rust toolchain)

cargo tauri build --release

# Execute the bundled binary

./target/release/ColibriDesktop

```

The desktop application communicates with the embedded engine via the mux protocol internally, requiring no additional configuration from the end user.

## OpenAI-Compatible API Server

Integration-focused deployments utilize `colibri/server` (implemented in Python) to wrap the mux protocol inside a Flask-compatible server. This strategy exposes a drop-in replacement for the OpenAI chat completions API, documented in [`docs/api.md`](https://github.com/JustVugg/colibri/blob/main/docs/api.md).

The server supports WSGI gateways like gunicorn for production workloads or can deploy as serverless functions on Azure or AWS Lambda.

```bash

# Install Python dependencies

pip install -r requirements.txt

# Start the OpenAI-compatible server

python -m colibri.server --host 0.0.0.0 --port 8000

```

```bash

# Query using standard OpenAI client patterns

curl -X POST http://localhost:8000/v1/chat/completions \
  -H "Authorization: Bearer sk-no-key-needed" \
  -H "Content-Type: application/json" \
  -d '{"model":"colibri","messages":[{"role":"user","content":"Explain quantum entanglement"}]}'

```

## Summary

- **Native CLI** provides minimal-overhead deployment for development and scripting via `colibri serve` and the mux protocol defined in [`docs/serve_protocol.md`](https://github.com/JustVugg/colibri/blob/main/docs/serve_protocol.md).
- **Docker containers** offer isolated, reproducible production environments with HTTP API access, configured through `docker/Dockerfile`.
- **Docker Compose and Kubernetes** enable orchestrated deployments with persistent telemetry storage and load balancing using the compose file at [`docker/docker-compose.yml`](https://github.com/JustVugg/colibri/blob/main/docker/docker-compose.yml).
- **Tauri Desktop** delivers native GUI applications for Windows and macOS without container runtimes, as detailed in [`desktop/README.md`](https://github.com/JustVugg/colibri/blob/main/desktop/README.md).
- **OpenAI-Compatible API** allows seamless integration into existing AI pipelines using standard HTTP clients and the Python server module.

## Frequently Asked Questions

### What GPU backends does Colibri support for deployment?

Colibri supports **CUDA**, **Metal**, and **CPU-only** backends. The build system in `c/` directory selects the appropriate backend during compilation, with supported configurations documented in [`GPU_BACKENDS.md`](https://github.com/JustVugg/colibri/blob/main/GPU_BACKENDS.md). Docker images can be built with specific backends by adjusting build arguments in `docker/Dockerfile`.

### What is the difference between the mux protocol and the legacy protocol?

The **mux protocol** supports multiplexed requests and continuous batching, enabling concurrent processing of multiple prompts through a single engine instance. The **legacy protocol** handles sequential requests without batching. Both protocols use line-oriented communication over stdin/stdout, with full specifications available in [`docs/serve_protocol.md`](https://github.com/JustVugg/colibri/blob/main/docs/serve_protocol.md).

### Can Colibri run on macOS without Docker?

Yes, the **Tauri Desktop App** strategy supports native macOS deployment without Docker. The application bundles the compiled engine from [`c/colibri.c`](https://github.com/JustVugg/colibri/blob/main/c/colibri.c) with a React frontend into a standalone `.app` binary. Users can also compile the native CLI directly for macOS using the Metal backend for GPU acceleration.

### How does Colibri handle persistent state and telemetry across deployments?

The engine emits telemetry lines (`HWINFO`, `TIERS`, `EMAP`) that servers can persist to the `.coli_usage` file. In Docker and Kubernetes deployments, this file should be mounted as a persistent volume to survive container restarts. The native CLI writes this file to the working directory by default.