Colibri Deployment Strategies: 5 Production-Ready Methods for Inference Workloads

Colibri supports five distinct deployment strategies ranging from native CLI execution to Kubernetes orchestration, each leveraging the architectural separation between the C-based inference engine and the protocol-speaking server components.

Colibri is a flexible, high-performance inference engine designed for diverse deployment scenarios. Its architecture separates the engine (native C code performing token generation) from the server (process speaking a line-oriented protocol over stdin/stdout), enabling deployment across environments from local laptops to cloud-scale clusters. According to the JustVugg/colibri source code, the engine compiles for multiple GPU backends including CUDA, Metal, and CPU-only configurations, making it adaptable to various infrastructure requirements.

Native CLI Deployment

The simplest deployment strategy runs the colibri binary directly on the host system. This approach minimizes overhead and eliminates container runtime dependencies, making it ideal for rapid prototyping and development workflows.

The native CLI supports two primary commands defined in colibri/cli.py:

  • colibri serve – Starts the server using the modern mux protocol for multiplexed request handling
  • colibri chat – Uses the legacy protocol for backward compatibility

The mux protocol, documented in docs/serve_protocol.md, enables continuous batching and concurrent request processing through line-oriented messages over stdin/stdout.


# Start the engine in mux mode for continuous batching

colibri serve --batch

# Submit a request via stdin (example format)

echo -e "SUBMIT 1 0 12 64 0.7 0.9\nHello world!" | colibri serve --batch

Docker Container Deployment

For production environments requiring isolation and reproducibility, the Docker strategy packages the compiled engine with a minimal runtime. The official image exposes both the HTTP API at /v1/chat/completions and a single-page application frontend.

The Dockerfile located at docker/Dockerfile handles the multi-stage build process, compiling the C sources from c/colibri.c and bundling the Tauri desktop build for GUI-capable environments.


# Pull the official image

docker pull justvugg/colibri:latest

# Run with GPU support and model mounting

docker run -d -p 8000:8000 \
  -e COLI_MODEL=/models/colibri-7b.q4_K_M.bin \
  -v ./models:/models \
  justvugg/colibri:latest

# Test the OpenAI-compatible endpoint

curl -X POST http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"colibri","messages":[{"role":"user","content":"Hello"}]}'

Docker Compose and Kubernetes Orchestration

Scalable deployments utilize docker/docker-compose.yml to orchestrate multiple services including reverse proxies and persistent storage. This strategy supports horizontal scaling and high-availability configurations through health checks and service dependencies.

The compose file includes configurations for:

  • Persistent volume mounting for the .coli_usage telemetry file
  • Nginx reverse proxy for load balancing
  • Environment variable management for GPU backend selection

# docker/docker-compose.yml

version: "3.9"
services:
  colibri:
    image: justvugg/colibri:latest
    ports:
      - "8000:8000"
    environment:
      - COLI_MODEL=/models/colibri-7b.q4_K_M.bin
    volumes:
      - ./models:/models
      - ./telemetry:/app/.coli_usage
    restart: unless-stopped

  nginx:
    image: nginx:alpine
    ports:
      - "80:80"
    volumes:
      - ./nginx.conf:/etc/nginx/nginx.conf:ro
    depends_on:
      - colibri

# Deploy the stack

docker compose -f docker/docker-compose.yml up -d

For Kubernetes environments, the compose file can be converted to Helm charts using kompose or manual translation, enabling auto-scaling and rolling updates.

Tauri Desktop Application

End-users requiring a graphical interface without Docker can deploy the Tauri Desktop App, which embeds the inference engine within a native Windows or macOS binary. The desktop build compiles the React/TypeScript frontend into a native application that exposes the same HTTP endpoint on localhost.

Build instructions in desktop/README.md specify the Rust-based compilation process:


# Build the desktop client (requires Rust toolchain)

cargo tauri build --release

# Execute the bundled binary

./target/release/ColibriDesktop

The desktop application communicates with the embedded engine via the mux protocol internally, requiring no additional configuration from the end user.

OpenAI-Compatible API Server

Integration-focused deployments utilize colibri/server (implemented in Python) to wrap the mux protocol inside a Flask-compatible server. This strategy exposes a drop-in replacement for the OpenAI chat completions API, documented in docs/api.md.

The server supports WSGI gateways like gunicorn for production workloads or can deploy as serverless functions on Azure or AWS Lambda.


# Install Python dependencies

pip install -r requirements.txt

# Start the OpenAI-compatible server

python -m colibri.server --host 0.0.0.0 --port 8000

# Query using standard OpenAI client patterns

curl -X POST http://localhost:8000/v1/chat/completions \
  -H "Authorization: Bearer sk-no-key-needed" \
  -H "Content-Type: application/json" \
  -d '{"model":"colibri","messages":[{"role":"user","content":"Explain quantum entanglement"}]}'

Summary

  • Native CLI provides minimal-overhead deployment for development and scripting via colibri serve and the mux protocol defined in docs/serve_protocol.md.
  • Docker containers offer isolated, reproducible production environments with HTTP API access, configured through docker/Dockerfile.
  • Docker Compose and Kubernetes enable orchestrated deployments with persistent telemetry storage and load balancing using the compose file at docker/docker-compose.yml.
  • Tauri Desktop delivers native GUI applications for Windows and macOS without container runtimes, as detailed in desktop/README.md.
  • OpenAI-Compatible API allows seamless integration into existing AI pipelines using standard HTTP clients and the Python server module.

Frequently Asked Questions

What GPU backends does Colibri support for deployment?

Colibri supports CUDA, Metal, and CPU-only backends. The build system in c/ directory selects the appropriate backend during compilation, with supported configurations documented in GPU_BACKENDS.md. Docker images can be built with specific backends by adjusting build arguments in docker/Dockerfile.

What is the difference between the mux protocol and the legacy protocol?

The mux protocol supports multiplexed requests and continuous batching, enabling concurrent processing of multiple prompts through a single engine instance. The legacy protocol handles sequential requests without batching. Both protocols use line-oriented communication over stdin/stdout, with full specifications available in docs/serve_protocol.md.

Can Colibri run on macOS without Docker?

Yes, the Tauri Desktop App strategy supports native macOS deployment without Docker. The application bundles the compiled engine from c/colibri.c with a React frontend into a standalone .app binary. Users can also compile the native CLI directly for macOS using the Metal backend for GPU acceleration.

How does Colibri handle persistent state and telemetry across deployments?

The engine emits telemetry lines (HWINFO, TIERS, EMAP) that servers can persist to the .coli_usage file. In Docker and Kubernetes deployments, this file should be mounted as a persistent volume to survive container restarts. The native CLI writes this file to the working directory by default.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →