# How to Integrate Colibri with Other Services: OpenAI-Compatible API Guide

> Easily integrate Colibri with other services using its OpenAI-compatible API. Connect any OpenAI client by redirecting the base URL to your Colibri instance for seamless integration.

- Repository: [Vincenzo Fornaro/colibri](https://github.com/JustVugg/colibri)
- Tags: how-to-guide
- Published: 2026-09-12

---

**Colibri ships with a dependency-free OpenAI-compatible HTTP gateway that exposes standard REST endpoints, allowing any OpenAI client to connect by simply redirecting the base URL to your Colibri instance.**

The JustVugg/colibri repository provides a drop-in replacement for commercial LLM APIs through its lightweight gateway. When you integrate Colibri with other services, you leverage a pure-Python server that translates standard OpenAI protocol requests into optimized local inference calls.

## Understanding the Colibri Gateway Architecture

The integration backbone resides in [`c/openai_server.py`](https://github.com/JustVugg/colibri/blob/main/c/openai_server.py), which implements a full HTTP gateway without external dependencies.

### Core Server Components

- **`default_engine()`**: Defined at lines 34-44, this function detects the requested model family and launches the compiled engine binary adjacent to the server process.
- **`GenerationScheduler`**: Implemented at lines 97-112, this class manages inference capacity through a bounded FIFO queue. The `admit()` context manager (lines 21-73) enforces `max_queue` limits, `queue_timeout` thresholds, and `capacity` constraints, returning HTTP 429 or 503 errors when the system is overloaded.

### Model Family Registration

The [`c/family_registry.py`](https://github.com/JustVugg/colibri/blob/main/c/family_registry.py) file maintains the `family_by_id` mapping (lines 1-25) that associates architecture identifiers—such as "glm", "inkling", "kimi_k3", or "deepseek_v4"—with their respective compiled engine binaries. When the gateway receives a request, it queries this registry to locate the correct backend for the specified `--arch` flag.

## Starting the OpenAI-Compatible Server

Launch the gateway using Python module execution syntax:

```bash
python -m c.openai_server --model /path/to/converted/model --arch glm --api-key my-secret-key

```

The server accepts several critical parameters:

1. **`--model`**: Filesystem path to the converted model weights (required).
2. **`--arch`**: Architecture family identifier (defaults to GLM; must match an entry in [`family_registry.py`](https://github.com/JustVugg/colibri/blob/main/family_registry.py)).
3. **`--api-key`**: Optional bearer token for request authentication and per-request throttling.
4. **`--port`**: Listening port (defaults to 8000 on 127.0.0.1).

Upon startup, the server binds to `http://127.0.0.1:8000` and exposes the standard OpenAI endpoint hierarchy: `/v1/models`, `/v1/chat/completions`, `/v1/completions`, and `/v1/health`.

## Integration Methods for External Services

Any client capable of posting JSON to an HTTP endpoint can integrate with Colibri. The following examples demonstrate connections from Python, JavaScript, and shell environments.

### Python Integration

Use the `requests` library to stream chat completions from the Colibri gateway:

```python
import requests
import json

BASE_URL = "http://localhost:8000"
API_KEY = "my-secret-key"

headers = {
    "Authorization": f"Bearer {API_KEY}",
    "Content-Type": "application/json",
}

payload = {
    "model": "glm",
    "messages": [
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Explain the multitier memory architecture."}
    ],
    "max_tokens": 256,
    "temperature": 0.7,
}

resp = requests.post(
    f"{BASE_URL}/v1/chat/completions",
    headers=headers,
    json=payload,
    stream=True
)

for line in resp.iter_lines():
    if line:
        data = json.loads(line.decode())
        print(data["choices"][0]["delta"]["content"], end="")

```

The server returns SSE-formatted JSON lines, which the example decodes incrementally to reconstruct the streaming response.

### JavaScript Integration

For browser or Node.js environments, use the fetch API as demonstrated in the reference client at [`web/src/lib/api.ts`](https://github.com/JustVugg/colibri/blob/main/web/src/lib/api.ts):

```javascript
const baseUrl = "http://localhost:8000";
const apiKey = "my-secret-key";

async function streamChat(prompt) {
  const response = await fetch(`${baseUrl}/v1/chat/completions`, {
    method: "POST",
    headers: {
      "Authorization": `Bearer ${apiKey}`,
      "Content-Type": "application/json",
    },
    body: JSON.stringify({
      model: "glm",
      messages: [
        {role: "system", content: "You are a helpful assistant."},
        {role: "user", content: prompt}
      ],
      max_tokens: 256,
      temperature: 0.7,
    })
  });

  const reader = response.body.getReader();
  const decoder = new TextDecoder();

  while (true) {
    const {value, done} = await reader.read();
    if (done) break;
    
    const lines = decoder.decode(value).split("\n").filter(l => l);
    for (const line of lines) {
      const data = JSON.parse(line);
      process.stdout.write(data.choices[0].delta.content);
    }
  }
}

streamChat("What is the advantage of multitenant VRAM management?");

```

### Command Line Integration

Test connectivity directly using cURL:

```bash
curl -X POST http://localhost:8000/v1/chat/completions \
  -H "Authorization: Bearer my-secret-key" \
  -H "Content-Type: application/json" \
  -d '{
        "model": "glm",
        "messages": [{"role":"user", "content":"Summarize Colibri architecture."}],
        "max_tokens": 128
      }'

```

## Authentication and Load Management

When you specify `--api-key` during server startup, the gateway validates the Bearer token in the Authorization header and forwards it to the underlying engine for per-request throttling.

The **GenerationScheduler** class protects the inference engine from overload through configurable constraints:

- **`capacity`**: Maximum concurrent KV contexts allowed.
- **`max_queue`**: Bounded FIFO queue length for pending requests.
- **`queue_timeout`**: Maximum wait time before returning HTTP 503.

These parameters ensure that external service integrations fail gracefully with standard HTTP status codes rather than degrading system performance.

## Summary

- **Colibri exposes OpenAI-compatible endpoints** through [`c/openai_server.py`](https://github.com/JustVugg/colibri/blob/main/c/openai_server.py), enabling drop-in replacement for commercial APIs.
- **Zero dependencies** are required for the gateway; it runs as a pure-Python HTTP server.
- **Model families** are registered in [`c/family_registry.py`](https://github.com/JustVugg/colibri/blob/main/c/family_registry.py) and selected via the `--arch` flag.
- **Streaming responses** follow Server-Sent Events format, compatible with standard OpenAI client libraries.
- **Built-in scheduling** via `GenerationScheduler` provides enterprise-grade rate limiting and queue management.

## Frequently Asked Questions

### Can I use existing OpenAI SDKs with Colibri?

Yes. Any SDK or client library that allows configuring the base URL can connect to Colibri. Simply point the client to `http://localhost:8000` (or your configured host/port) instead of the standard OpenAI API endpoint. The gateway implements `/v1/models`, `/v1/chat/completions`, and other standard routes.

### How does Colibri handle high-traffic scenarios from multiple services?

The `GenerationScheduler` in [`c/openai_server.py`](https://github.com/JustVugg/colibri/blob/main/c/openai_server.py) implements a bounded FIFO queue with configurable `capacity`, `max_queue`, and `queue_timeout` parameters. When limits are exceeded, the gateway returns HTTP 429 (Too Many Requests) or 503 (Service Unavailable) errors, preventing engine overload and ensuring predictable performance.

### Is the API key authentication mandatory for integration?

No. The `--api-key` flag is optional. When omitted, the server accepts anonymous requests. When enabled, the gateway validates the Bearer token and passes it to the engine for per-request tracking and throttling, which is useful for multi-tenant deployments.

### Which model architectures does the gateway support?

The gateway supports any architecture registered in [`c/family_registry.py`](https://github.com/JustVugg/colibri/blob/main/c/family_registry.py), including GLM, Inkling, Kimi K3, and DeepSeek V4. The `--arch` flag selects the appropriate engine binary, while the `family_by_id` function (lines 1-25) resolves model identifiers to their respective inference engines.