How to Integrate Colibri with Other Services: OpenAI-Compatible API Guide
Colibri ships with a dependency-free OpenAI-compatible HTTP gateway that exposes standard REST endpoints, allowing any OpenAI client to connect by simply redirecting the base URL to your Colibri instance.
The JustVugg/colibri repository provides a drop-in replacement for commercial LLM APIs through its lightweight gateway. When you integrate Colibri with other services, you leverage a pure-Python server that translates standard OpenAI protocol requests into optimized local inference calls.
Understanding the Colibri Gateway Architecture
The integration backbone resides in c/openai_server.py, which implements a full HTTP gateway without external dependencies.
Core Server Components
default_engine(): Defined at lines 34-44, this function detects the requested model family and launches the compiled engine binary adjacent to the server process.GenerationScheduler: Implemented at lines 97-112, this class manages inference capacity through a bounded FIFO queue. Theadmit()context manager (lines 21-73) enforcesmax_queuelimits,queue_timeoutthresholds, andcapacityconstraints, returning HTTP 429 or 503 errors when the system is overloaded.
Model Family Registration
The c/family_registry.py file maintains the family_by_id mapping (lines 1-25) that associates architecture identifiers—such as "glm", "inkling", "kimi_k3", or "deepseek_v4"—with their respective compiled engine binaries. When the gateway receives a request, it queries this registry to locate the correct backend for the specified --arch flag.
Starting the OpenAI-Compatible Server
Launch the gateway using Python module execution syntax:
python -m c.openai_server --model /path/to/converted/model --arch glm --api-key my-secret-key
The server accepts several critical parameters:
--model: Filesystem path to the converted model weights (required).--arch: Architecture family identifier (defaults to GLM; must match an entry infamily_registry.py).--api-key: Optional bearer token for request authentication and per-request throttling.--port: Listening port (defaults to 8000 on 127.0.0.1).
Upon startup, the server binds to http://127.0.0.1:8000 and exposes the standard OpenAI endpoint hierarchy: /v1/models, /v1/chat/completions, /v1/completions, and /v1/health.
Integration Methods for External Services
Any client capable of posting JSON to an HTTP endpoint can integrate with Colibri. The following examples demonstrate connections from Python, JavaScript, and shell environments.
Python Integration
Use the requests library to stream chat completions from the Colibri gateway:
import requests
import json
BASE_URL = "http://localhost:8000"
API_KEY = "my-secret-key"
headers = {
"Authorization": f"Bearer {API_KEY}",
"Content-Type": "application/json",
}
payload = {
"model": "glm",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Explain the multitier memory architecture."}
],
"max_tokens": 256,
"temperature": 0.7,
}
resp = requests.post(
f"{BASE_URL}/v1/chat/completions",
headers=headers,
json=payload,
stream=True
)
for line in resp.iter_lines():
if line:
data = json.loads(line.decode())
print(data["choices"][0]["delta"]["content"], end="")
The server returns SSE-formatted JSON lines, which the example decodes incrementally to reconstruct the streaming response.
JavaScript Integration
For browser or Node.js environments, use the fetch API as demonstrated in the reference client at web/src/lib/api.ts:
const baseUrl = "http://localhost:8000";
const apiKey = "my-secret-key";
async function streamChat(prompt) {
const response = await fetch(`${baseUrl}/v1/chat/completions`, {
method: "POST",
headers: {
"Authorization": `Bearer ${apiKey}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "glm",
messages: [
{role: "system", content: "You are a helpful assistant."},
{role: "user", content: prompt}
],
max_tokens: 256,
temperature: 0.7,
})
});
const reader = response.body.getReader();
const decoder = new TextDecoder();
while (true) {
const {value, done} = await reader.read();
if (done) break;
const lines = decoder.decode(value).split("\n").filter(l => l);
for (const line of lines) {
const data = JSON.parse(line);
process.stdout.write(data.choices[0].delta.content);
}
}
}
streamChat("What is the advantage of multitenant VRAM management?");
Command Line Integration
Test connectivity directly using cURL:
curl -X POST http://localhost:8000/v1/chat/completions \
-H "Authorization: Bearer my-secret-key" \
-H "Content-Type: application/json" \
-d '{
"model": "glm",
"messages": [{"role":"user", "content":"Summarize Colibri architecture."}],
"max_tokens": 128
}'
Authentication and Load Management
When you specify --api-key during server startup, the gateway validates the Bearer token in the Authorization header and forwards it to the underlying engine for per-request throttling.
The GenerationScheduler class protects the inference engine from overload through configurable constraints:
capacity: Maximum concurrent KV contexts allowed.max_queue: Bounded FIFO queue length for pending requests.queue_timeout: Maximum wait time before returning HTTP 503.
These parameters ensure that external service integrations fail gracefully with standard HTTP status codes rather than degrading system performance.
Summary
- Colibri exposes OpenAI-compatible endpoints through
c/openai_server.py, enabling drop-in replacement for commercial APIs. - Zero dependencies are required for the gateway; it runs as a pure-Python HTTP server.
- Model families are registered in
c/family_registry.pyand selected via the--archflag. - Streaming responses follow Server-Sent Events format, compatible with standard OpenAI client libraries.
- Built-in scheduling via
GenerationSchedulerprovides enterprise-grade rate limiting and queue management.
Frequently Asked Questions
Can I use existing OpenAI SDKs with Colibri?
Yes. Any SDK or client library that allows configuring the base URL can connect to Colibri. Simply point the client to http://localhost:8000 (or your configured host/port) instead of the standard OpenAI API endpoint. The gateway implements /v1/models, /v1/chat/completions, and other standard routes.
How does Colibri handle high-traffic scenarios from multiple services?
The GenerationScheduler in c/openai_server.py implements a bounded FIFO queue with configurable capacity, max_queue, and queue_timeout parameters. When limits are exceeded, the gateway returns HTTP 429 (Too Many Requests) or 503 (Service Unavailable) errors, preventing engine overload and ensuring predictable performance.
Is the API key authentication mandatory for integration?
No. The --api-key flag is optional. When omitted, the server accepts anonymous requests. When enabled, the gateway validates the Bearer token and passes it to the engine for per-request tracking and throttling, which is useful for multi-tenant deployments.
Which model architectures does the gateway support?
The gateway supports any architecture registered in c/family_registry.py, including GLM, Inkling, Kimi K3, and DeepSeek V4. The --arch flag selects the appropriate engine binary, while the family_by_id function (lines 1-25) resolves model identifiers to their respective inference engines.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →