Colibri Deployment Strategies: 5 Production-Ready Methods for Inference Workloads
Colibri supports five distinct deployment strategies ranging from native CLI execution to Kubernetes orchestration, each leveraging the architectural separation between the C-based inference engine and the protocol-speaking server components.
Colibri is a flexible, high-performance inference engine designed for diverse deployment scenarios. Its architecture separates the engine (native C code performing token generation) from the server (process speaking a line-oriented protocol over stdin/stdout), enabling deployment across environments from local laptops to cloud-scale clusters. According to the JustVugg/colibri source code, the engine compiles for multiple GPU backends including CUDA, Metal, and CPU-only configurations, making it adaptable to various infrastructure requirements.
Native CLI Deployment
The simplest deployment strategy runs the colibri binary directly on the host system. This approach minimizes overhead and eliminates container runtime dependencies, making it ideal for rapid prototyping and development workflows.
The native CLI supports two primary commands defined in colibri/cli.py:
colibri serve– Starts the server using the modern mux protocol for multiplexed request handlingcolibri chat– Uses the legacy protocol for backward compatibility
The mux protocol, documented in docs/serve_protocol.md, enables continuous batching and concurrent request processing through line-oriented messages over stdin/stdout.
# Start the engine in mux mode for continuous batching
colibri serve --batch
# Submit a request via stdin (example format)
echo -e "SUBMIT 1 0 12 64 0.7 0.9\nHello world!" | colibri serve --batch
Docker Container Deployment
For production environments requiring isolation and reproducibility, the Docker strategy packages the compiled engine with a minimal runtime. The official image exposes both the HTTP API at /v1/chat/completions and a single-page application frontend.
The Dockerfile located at docker/Dockerfile handles the multi-stage build process, compiling the C sources from c/colibri.c and bundling the Tauri desktop build for GUI-capable environments.
# Pull the official image
docker pull justvugg/colibri:latest
# Run with GPU support and model mounting
docker run -d -p 8000:8000 \
-e COLI_MODEL=/models/colibri-7b.q4_K_M.bin \
-v ./models:/models \
justvugg/colibri:latest
# Test the OpenAI-compatible endpoint
curl -X POST http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"colibri","messages":[{"role":"user","content":"Hello"}]}'
Docker Compose and Kubernetes Orchestration
Scalable deployments utilize docker/docker-compose.yml to orchestrate multiple services including reverse proxies and persistent storage. This strategy supports horizontal scaling and high-availability configurations through health checks and service dependencies.
The compose file includes configurations for:
- Persistent volume mounting for the
.coli_usagetelemetry file - Nginx reverse proxy for load balancing
- Environment variable management for GPU backend selection
# docker/docker-compose.yml
version: "3.9"
services:
colibri:
image: justvugg/colibri:latest
ports:
- "8000:8000"
environment:
- COLI_MODEL=/models/colibri-7b.q4_K_M.bin
volumes:
- ./models:/models
- ./telemetry:/app/.coli_usage
restart: unless-stopped
nginx:
image: nginx:alpine
ports:
- "80:80"
volumes:
- ./nginx.conf:/etc/nginx/nginx.conf:ro
depends_on:
- colibri
# Deploy the stack
docker compose -f docker/docker-compose.yml up -d
For Kubernetes environments, the compose file can be converted to Helm charts using kompose or manual translation, enabling auto-scaling and rolling updates.
Tauri Desktop Application
End-users requiring a graphical interface without Docker can deploy the Tauri Desktop App, which embeds the inference engine within a native Windows or macOS binary. The desktop build compiles the React/TypeScript frontend into a native application that exposes the same HTTP endpoint on localhost.
Build instructions in desktop/README.md specify the Rust-based compilation process:
# Build the desktop client (requires Rust toolchain)
cargo tauri build --release
# Execute the bundled binary
./target/release/ColibriDesktop
The desktop application communicates with the embedded engine via the mux protocol internally, requiring no additional configuration from the end user.
OpenAI-Compatible API Server
Integration-focused deployments utilize colibri/server (implemented in Python) to wrap the mux protocol inside a Flask-compatible server. This strategy exposes a drop-in replacement for the OpenAI chat completions API, documented in docs/api.md.
The server supports WSGI gateways like gunicorn for production workloads or can deploy as serverless functions on Azure or AWS Lambda.
# Install Python dependencies
pip install -r requirements.txt
# Start the OpenAI-compatible server
python -m colibri.server --host 0.0.0.0 --port 8000
# Query using standard OpenAI client patterns
curl -X POST http://localhost:8000/v1/chat/completions \
-H "Authorization: Bearer sk-no-key-needed" \
-H "Content-Type: application/json" \
-d '{"model":"colibri","messages":[{"role":"user","content":"Explain quantum entanglement"}]}'
Summary
- Native CLI provides minimal-overhead deployment for development and scripting via
colibri serveand the mux protocol defined indocs/serve_protocol.md. - Docker containers offer isolated, reproducible production environments with HTTP API access, configured through
docker/Dockerfile. - Docker Compose and Kubernetes enable orchestrated deployments with persistent telemetry storage and load balancing using the compose file at
docker/docker-compose.yml. - Tauri Desktop delivers native GUI applications for Windows and macOS without container runtimes, as detailed in
desktop/README.md. - OpenAI-Compatible API allows seamless integration into existing AI pipelines using standard HTTP clients and the Python server module.
Frequently Asked Questions
What GPU backends does Colibri support for deployment?
Colibri supports CUDA, Metal, and CPU-only backends. The build system in c/ directory selects the appropriate backend during compilation, with supported configurations documented in GPU_BACKENDS.md. Docker images can be built with specific backends by adjusting build arguments in docker/Dockerfile.
What is the difference between the mux protocol and the legacy protocol?
The mux protocol supports multiplexed requests and continuous batching, enabling concurrent processing of multiple prompts through a single engine instance. The legacy protocol handles sequential requests without batching. Both protocols use line-oriented communication over stdin/stdout, with full specifications available in docs/serve_protocol.md.
Can Colibri run on macOS without Docker?
Yes, the Tauri Desktop App strategy supports native macOS deployment without Docker. The application bundles the compiled engine from c/colibri.c with a React frontend into a standalone .app binary. Users can also compile the native CLI directly for macOS using the Metal backend for GPU acceleration.
How does Colibri handle persistent state and telemetry across deployments?
The engine emits telemetry lines (HWINFO, TIERS, EMAP) that servers can persist to the .coli_usage file. In Docker and Kubernetes deployments, this file should be mounted as a persistent volume to survive container restarts. The native CLI writes this file to the working directory by default.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →