Colibri API Structure: OpenAI-Compatible Endpoints and Implementation Details

Colibri exposes a pure OpenAI-compatible HTTP API through a lightweight Python gateway that runs on zero external dependencies, offering chat completions, legacy completions, and tool-calling support via standard REST endpoints.

The JustVugg/colibri repository implements a self-contained inference engine with a thin Python wrapper that presents a familiar REST interface. Understanding the Colibri API structure reveals how the gateway in c/openai_server.py translates standard OpenAI SDK calls into efficient native execution without requiring external libraries. This architecture separates protocol handling from inference, keeping the core engine dependency-free while maintaining full compatibility with existing client libraries.

Core HTTP Endpoints

The API surface follows the OpenAI specification with additional Anthropic compatibility. All routes share the base URL http://127.0.0.1:8000/v1 and are implemented in c/openai_server.py using only Python standard library components.

Model Management and Health

  • GET /v1/models — Lists available model IDs via handle_models (lines 24‑31)
  • GET /v1/models/{model} — Retrieves metadata for a specific model by parsing the URL segment
  • GET /health — Exposes queue counters (active, queued, completed) via handle_health (lines 610‑630)

Completion Endpoints

  • POST /v1/chat/completions — Primary chat interface supporting streaming, temperature, top‑p, and tool calls via handle_chat_completion (lines 350‑420)
  • POST /v1/completions — Legacy completion endpoint for older clients via handle_completion (lines 440‑500)
  • POST /v1/messages — Anthropic Messages API compatibility via handle_anthropic_messages (lines 520‑580)

Request Format and Streaming Protocol

Requests use standard JSON bodies following the OpenAI schema. The model field accepts identifiers like glm-5.2-colibri, while messages expects arrays of {role, content} objects.

Optional parameters include max_tokens, temperature, top_p, stop, stream, and tool_choice.

Streaming responses use Server-Sent Events (SSE) terminated by the special marker defined as END = b"\x01\x01END\x01\x01\n" in openai_server.py. This marker signals the end of generation to the client.

Tool-Calling Architecture

Tool support varies by engine. The gateway parses native engine formats through parse_arch_tool_calls (lines 620‑640) and converts them into OpenAI-compatible JSON responses.

Engine-Specific Tool Support

Engine OpenAI tools Anthropic tool_use Native Format
GLM-5.2 ✅ ✅ <tool_call> XML tags
DeepSeek V4 ✅ ✅ DSML blocks (<|DSML|tool_calls>)
Kimi K3 ✅ ✅ XTML (`<
Inkling, Qwen 3.8, OLMoE ❌ ❌ Returns HTTP 400

When quantized models produce malformed tool calls, setting COLI_TOOL_SALVAGE=1 activates recovery logic using the _SALVAGE flag (lines 44‑48).

KV Context Slots and Request Scheduling

Colibri maintains up to 16 independent KV cache contexts controlled by --kv-slots or the COLI_KV_SLOTS environment variable. Clients target specific slots using the cache_slot JSON field to reuse cached prefixes across multi-turn conversations, reducing latency.

The GenerationScheduler (lines 97‑180) manages admission via a bounded FIFO queue. Configure limits through:

  • COLI_MAX_QUEUE: Maximum concurrent requests
  • COLI_QUEUE_TIMEOUT: Milliseconds before queue rejection

The /health endpoint and x-colibri-queue-wait-ms response header expose current queue statistics.

Authentication and CORS Configuration

Authentication is optional. When COLI_API_KEY is set, the server requires Authorization: Bearer <key> headers; otherwise any non-empty string (commonly "local") suffices.

CORS defaults to common local origins defined in DEFAULT_CORS_ORIGINS (lines 48‑55). Additional allowed origins are configured via --cors-origin or the COLI_ALLOWED_HOSTS environment variable.

Environment Variables and Runtime Flags

Configuration occurs through environment variables parsed in colibri/cli.py and passed to the gateway:

Variable Purpose
COLI_MODEL Path to model checkpoint (e.g., /nvme/glm52_i4)
COLI_API_KEY Bearer token for external exposure
COLI_MAX_QUEUE / COLI_QUEUE_TIMEOUT Scheduler admission controls
COLI_TOOL_SALVAGE Enable malformed tool call recovery (set to 1)
COLI_ALLOWED_HOSTS Comma-separated reverse proxy hostnames
COLI_KV_SLOTS Number of KV contexts (default 8)

Practical Code Examples

Basic Health Check

curl http://127.0.0.1:8000/health

Streaming Chat Completion

curl http://127.0.0.1:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
        "model": "glm-5.2-colibri",
        "messages": [{"role":"user","content":"Hello"}],
        "stream": true
      }'

OpenAI Python SDK Integration

import openai

openai.api_base = "http://localhost:8000/v1"
openai.api_key = "local"

resp = openai.ChatCompletion.create(
    model="glm-5.2-colibri",
    messages=[{"role": "user", "content": "What is the capital of France?"}]
)
print(resp.choices[0].message.content)

Tool Calling with GLM-5.2

import openai
import json

tools = [{
    "type": "function",
    "function": {
        "name": "get_weather",
        "description": "Lookup current weather",
        "parameters": {
            "type": "object",
            "properties": {"city": {"type": "string"}},
            "required": ["city"]
        }
    }
}]

resp = openai.ChatCompletion.create(
    model="glm-5.2-colibri",
    messages=[{"role": "user", "content": "Weather in Berlin?"}],
    tools=tools,
    tool_choice="auto"
)

if resp.choices[0].finish_reason == "tool_calls":
    call = resp.choices[0].message.tool_calls[0]
    print(f"Tool invoked: {call.function.name}")

Utilizing KV Cache Slots

curl http://127.0.0.1:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
        "model": "glm-5.2-colibri",
        "messages": [{"role": "user", "content": "Context setting prompt"}],
        "cache_slot": 1
      }'

Subsequent requests with "cache_slot": 1 reuse the cached KV prefix.

Summary

  • Colibri's API structure centers on c/openai_server.py, a zero-dependency Python gateway implementing OpenAI-compatible REST endpoints including /v1/chat/completions and /v1/models.
  • Streaming and tools are fully supported via Server-Sent Events terminated by \x01\x01END\x01\x01, with tool-calling available for GLM-5.2, DeepSeek V4, and Kimi K3 through native syntax parsing in parse_arch_tool_calls.
  • KV slot management allows up to 16 independent cache contexts via the cache_slot parameter, managed by the GenerationScheduler with configurable queue limits.
  • Security and deployment require no external packages; authentication is optional via COLI_API_KEY, and CORS is configurable through environment variables.

Frequently Asked Questions

Does Colibri require external dependencies to run the API server?

No. The gateway in c/openai_server.py uses only the Python standard library. The inference engine in the c/ directory is a standalone binary with zero external dependencies, making deployment entirely self-contained.

How does Colibri handle streaming responses?

Streaming uses Server-Sent Events (SSE) with a custom termination marker \x01\x01END\x01\x01 defined in openai_server.py. The handle_chat_completion function (lines 350‑420) manages the SSE stream, chunking tokens as they generate from the C engine subprocess and flushing them to the client until the end marker is reached.

Can I use the standard OpenAI Python SDK with Colibri?

Yes. Point the SDK to http://localhost:8000/v1 and use any non-empty API key (such as "local"). The gateway accepts standard parameters like temperature, max_tokens, and tools, forwarding them to the native engine without modification.

What happens when tool-calling is requested on an unsupported engine?

Engines without native tool support (Inkling, Qwen 3.8, OLMoE) return HTTP 400 Bad Request when the tools parameter is present. For supported engines, the gateway parses native formats like <tool_call> or DSML blocks in parse_arch_tool_calls (lines 620‑640) and converts them to OpenAI-compatible JSON structured responses.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →