# Colibri API Structure: OpenAI-Compatible Endpoints and Implementation Details

> Explore the Colibri API structure with OpenAI-compatible endpoints. Implement chat completions, tool-calling, and more using a lightweight Python gateway with zero external dependencies.

- Repository: [Vincenzo Fornaro/colibri](https://github.com/JustVugg/colibri)
- Tags: api-reference
- Published: 2026-09-12

---

**Colibri exposes a pure OpenAI-compatible HTTP API through a lightweight Python gateway that runs on zero external dependencies, offering chat completions, legacy completions, and tool-calling support via standard REST endpoints.**

The JustVugg/colibri repository implements a self-contained inference engine with a thin Python wrapper that presents a familiar REST interface. Understanding the Colibri API structure reveals how the gateway in [`c/openai_server.py`](https://github.com/JustVugg/colibri/blob/main/c/openai_server.py) translates standard OpenAI SDK calls into efficient native execution without requiring external libraries. This architecture separates protocol handling from inference, keeping the core engine dependency-free while maintaining full compatibility with existing client libraries.

## Core HTTP Endpoints

The API surface follows the OpenAI specification with additional Anthropic compatibility. All routes share the base URL `http://127.0.0.1:8000/v1` and are implemented in [`c/openai_server.py`](https://github.com/JustVugg/colibri/blob/main/c/openai_server.py) using only Python standard library components.

### Model Management and Health

- **GET `/v1/models`** — Lists available model IDs via `handle_models` (lines 24‑31)
- **GET `/v1/models/{model}`** — Retrieves metadata for a specific model by parsing the URL segment
- **GET `/health`** — Exposes queue counters (`active`, `queued`, `completed`) via `handle_health` (lines 610‑630)

### Completion Endpoints

- **POST `/v1/chat/completions`** — Primary chat interface supporting streaming, temperature, top‑p, and tool calls via `handle_chat_completion` (lines 350‑420)
- **POST `/v1/completions`** — Legacy completion endpoint for older clients via `handle_completion` (lines 440‑500)
- **POST `/v1/messages`** — Anthropic Messages API compatibility via `handle_anthropic_messages` (lines 520‑580)

## Request Format and Streaming Protocol

Requests use standard JSON bodies following the OpenAI schema. The `model` field accepts identifiers like `glm-5.2-colibri`, while `messages` expects arrays of `{role, content}` objects.

Optional parameters include `max_tokens`, `temperature`, `top_p`, `stop`, `stream`, and `tool_choice`.

Streaming responses use Server-Sent Events (SSE) terminated by the special marker defined as `END = b"\x01\x01END\x01\x01\n"` in [`openai_server.py`](https://github.com/JustVugg/colibri/blob/main/openai_server.py). This marker signals the end of generation to the client.

## Tool-Calling Architecture

Tool support varies by engine. The gateway parses native engine formats through `parse_arch_tool_calls` (lines 620‑640) and converts them into OpenAI-compatible JSON responses.

### Engine-Specific Tool Support

| Engine | OpenAI `tools` | Anthropic `tool_use` | Native Format |
|--------|----------------|----------------------|---------------|
| **GLM-5.2** | ✅ | ✅ | `<tool_call>` XML tags |
| **DeepSeek V4** | ✅ | ✅ | DSML blocks (`<｜DSML｜tool_calls>`) |
| **Kimi K3** | ✅ | ✅ | XTML (`<|open|>tools<|sep|>`) |
| **Inkling, Qwen 3.8, OLMoE** | ❌ | ❌ | Returns HTTP 400 |

When quantized models produce malformed tool calls, setting `COLI_TOOL_SALVAGE=1` activates recovery logic using the `_SALVAGE` flag (lines 44‑48).

## KV Context Slots and Request Scheduling

Colibri maintains up to 16 independent KV cache contexts controlled by `--kv-slots` or the `COLI_KV_SLOTS` environment variable. Clients target specific slots using the `cache_slot` JSON field to reuse cached prefixes across multi-turn conversations, reducing latency.

The `GenerationScheduler` (lines 97‑180) manages admission via a bounded FIFO queue. Configure limits through:

- **COLI_MAX_QUEUE**: Maximum concurrent requests
- **COLI_QUEUE_TIMEOUT**: Milliseconds before queue rejection

The `/health` endpoint and `x-colibri-queue-wait-ms` response header expose current queue statistics.

## Authentication and CORS Configuration

Authentication is optional. When `COLI_API_KEY` is set, the server requires `Authorization: Bearer <key>` headers; otherwise any non-empty string (commonly `"local"`) suffices.

CORS defaults to common local origins defined in `DEFAULT_CORS_ORIGINS` (lines 48‑55). Additional allowed origins are configured via `--cors-origin` or the `COLI_ALLOWED_HOSTS` environment variable.

## Environment Variables and Runtime Flags

Configuration occurs through environment variables parsed in [`colibri/cli.py`](https://github.com/JustVugg/colibri/blob/main/colibri/cli.py) and passed to the gateway:

| Variable | Purpose |
|----------|---------|
| `COLI_MODEL` | Path to model checkpoint (e.g., `/nvme/glm52_i4`) |
| `COLI_API_KEY` | Bearer token for external exposure |
| `COLI_MAX_QUEUE` / `COLI_QUEUE_TIMEOUT` | Scheduler admission controls |
| `COLI_TOOL_SALVAGE` | Enable malformed tool call recovery (set to `1`) |
| `COLI_ALLOWED_HOSTS` | Comma-separated reverse proxy hostnames |
| `COLI_KV_SLOTS` | Number of KV contexts (default 8) |

## Practical Code Examples

### Basic Health Check

```bash
curl http://127.0.0.1:8000/health

```

### Streaming Chat Completion

```bash
curl http://127.0.0.1:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
        "model": "glm-5.2-colibri",
        "messages": [{"role":"user","content":"Hello"}],
        "stream": true
      }'

```

### OpenAI Python SDK Integration

```python
import openai

openai.api_base = "http://localhost:8000/v1"
openai.api_key = "local"

resp = openai.ChatCompletion.create(
    model="glm-5.2-colibri",
    messages=[{"role": "user", "content": "What is the capital of France?"}]
)
print(resp.choices[0].message.content)

```

### Tool Calling with GLM-5.2

```python
import openai
import json

tools = [{
    "type": "function",
    "function": {
        "name": "get_weather",
        "description": "Lookup current weather",
        "parameters": {
            "type": "object",
            "properties": {"city": {"type": "string"}},
            "required": ["city"]
        }
    }
}]

resp = openai.ChatCompletion.create(
    model="glm-5.2-colibri",
    messages=[{"role": "user", "content": "Weather in Berlin?"}],
    tools=tools,
    tool_choice="auto"
)

if resp.choices[0].finish_reason == "tool_calls":
    call = resp.choices[0].message.tool_calls[0]
    print(f"Tool invoked: {call.function.name}")

```

### Utilizing KV Cache Slots

```bash
curl http://127.0.0.1:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
        "model": "glm-5.2-colibri",
        "messages": [{"role": "user", "content": "Context setting prompt"}],
        "cache_slot": 1
      }'

```

Subsequent requests with `"cache_slot": 1` reuse the cached KV prefix.

## Summary

- **Colibri's API structure** centers on [`c/openai_server.py`](https://github.com/JustVugg/colibri/blob/main/c/openai_server.py), a zero-dependency Python gateway implementing OpenAI-compatible REST endpoints including `/v1/chat/completions` and `/v1/models`.
- **Streaming and tools** are fully supported via Server-Sent Events terminated by `\x01\x01END\x01\x01`, with tool-calling available for GLM-5.2, DeepSeek V4, and Kimi K3 through native syntax parsing in `parse_arch_tool_calls`.
- **KV slot management** allows up to 16 independent cache contexts via the `cache_slot` parameter, managed by the `GenerationScheduler` with configurable queue limits.
- **Security and deployment** require no external packages; authentication is optional via `COLI_API_KEY`, and CORS is configurable through environment variables.

## Frequently Asked Questions

### Does Colibri require external dependencies to run the API server?

No. The gateway in [`c/openai_server.py`](https://github.com/JustVugg/colibri/blob/main/c/openai_server.py) uses only the Python standard library. The inference engine in the `c/` directory is a standalone binary with zero external dependencies, making deployment entirely self-contained.

### How does Colibri handle streaming responses?

Streaming uses Server-Sent Events (SSE) with a custom termination marker `\x01\x01END\x01\x01` defined in [`openai_server.py`](https://github.com/JustVugg/colibri/blob/main/openai_server.py). The `handle_chat_completion` function (lines 350‑420) manages the SSE stream, chunking tokens as they generate from the C engine subprocess and flushing them to the client until the end marker is reached.

### Can I use the standard OpenAI Python SDK with Colibri?

Yes. Point the SDK to `http://localhost:8000/v1` and use any non-empty API key (such as `"local"`). The gateway accepts standard parameters like `temperature`, `max_tokens`, and `tools`, forwarding them to the native engine without modification.

### What happens when tool-calling is requested on an unsupported engine?

Engines without native tool support (Inkling, Qwen 3.8, OLMoE) return HTTP 400 Bad Request when the `tools` parameter is present. For supported engines, the gateway parses native formats like `<tool_call>` or DSML blocks in `parse_arch_tool_calls` (lines 620‑640) and converts them to OpenAI-compatible JSON structured responses.