Colibri API Structure: OpenAI-Compatible Endpoints and Implementation Details
Colibri exposes a pure OpenAI-compatible HTTP API through a lightweight Python gateway that runs on zero external dependencies, offering chat completions, legacy completions, and tool-calling support via standard REST endpoints.
The JustVugg/colibri repository implements a self-contained inference engine with a thin Python wrapper that presents a familiar REST interface. Understanding the Colibri API structure reveals how the gateway in c/openai_server.py translates standard OpenAI SDK calls into efficient native execution without requiring external libraries. This architecture separates protocol handling from inference, keeping the core engine dependency-free while maintaining full compatibility with existing client libraries.
Core HTTP Endpoints
The API surface follows the OpenAI specification with additional Anthropic compatibility. All routes share the base URL http://127.0.0.1:8000/v1 and are implemented in c/openai_server.py using only Python standard library components.
Model Management and Health
- GET
/v1/models— Lists available model IDs viahandle_models(lines 24‑31) - GET
/v1/models/{model}— Retrieves metadata for a specific model by parsing the URL segment - GET
/health— Exposes queue counters (active,queued,completed) viahandle_health(lines 610‑630)
Completion Endpoints
- POST
/v1/chat/completions— Primary chat interface supporting streaming, temperature, top‑p, and tool calls viahandle_chat_completion(lines 350‑420) - POST
/v1/completions— Legacy completion endpoint for older clients viahandle_completion(lines 440‑500) - POST
/v1/messages— Anthropic Messages API compatibility viahandle_anthropic_messages(lines 520‑580)
Request Format and Streaming Protocol
Requests use standard JSON bodies following the OpenAI schema. The model field accepts identifiers like glm-5.2-colibri, while messages expects arrays of {role, content} objects.
Optional parameters include max_tokens, temperature, top_p, stop, stream, and tool_choice.
Streaming responses use Server-Sent Events (SSE) terminated by the special marker defined as END = b"\x01\x01END\x01\x01\n" in openai_server.py. This marker signals the end of generation to the client.
Tool-Calling Architecture
Tool support varies by engine. The gateway parses native engine formats through parse_arch_tool_calls (lines 620‑640) and converts them into OpenAI-compatible JSON responses.
Engine-Specific Tool Support
| Engine | OpenAI tools |
Anthropic tool_use |
Native Format |
|---|---|---|---|
| GLM-5.2 | ✅ | ✅ | <tool_call> XML tags |
| DeepSeek V4 | ✅ | ✅ | DSML blocks (<|DSML|tool_calls>) |
| Kimi K3 | ✅ | ✅ | XTML (`< |
| Inkling, Qwen 3.8, OLMoE | ❌ | ❌ | Returns HTTP 400 |
When quantized models produce malformed tool calls, setting COLI_TOOL_SALVAGE=1 activates recovery logic using the _SALVAGE flag (lines 44‑48).
KV Context Slots and Request Scheduling
Colibri maintains up to 16 independent KV cache contexts controlled by --kv-slots or the COLI_KV_SLOTS environment variable. Clients target specific slots using the cache_slot JSON field to reuse cached prefixes across multi-turn conversations, reducing latency.
The GenerationScheduler (lines 97‑180) manages admission via a bounded FIFO queue. Configure limits through:
- COLI_MAX_QUEUE: Maximum concurrent requests
- COLI_QUEUE_TIMEOUT: Milliseconds before queue rejection
The /health endpoint and x-colibri-queue-wait-ms response header expose current queue statistics.
Authentication and CORS Configuration
Authentication is optional. When COLI_API_KEY is set, the server requires Authorization: Bearer <key> headers; otherwise any non-empty string (commonly "local") suffices.
CORS defaults to common local origins defined in DEFAULT_CORS_ORIGINS (lines 48‑55). Additional allowed origins are configured via --cors-origin or the COLI_ALLOWED_HOSTS environment variable.
Environment Variables and Runtime Flags
Configuration occurs through environment variables parsed in colibri/cli.py and passed to the gateway:
| Variable | Purpose |
|---|---|
COLI_MODEL |
Path to model checkpoint (e.g., /nvme/glm52_i4) |
COLI_API_KEY |
Bearer token for external exposure |
COLI_MAX_QUEUE / COLI_QUEUE_TIMEOUT |
Scheduler admission controls |
COLI_TOOL_SALVAGE |
Enable malformed tool call recovery (set to 1) |
COLI_ALLOWED_HOSTS |
Comma-separated reverse proxy hostnames |
COLI_KV_SLOTS |
Number of KV contexts (default 8) |
Practical Code Examples
Basic Health Check
curl http://127.0.0.1:8000/health
Streaming Chat Completion
curl http://127.0.0.1:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "glm-5.2-colibri",
"messages": [{"role":"user","content":"Hello"}],
"stream": true
}'
OpenAI Python SDK Integration
import openai
openai.api_base = "http://localhost:8000/v1"
openai.api_key = "local"
resp = openai.ChatCompletion.create(
model="glm-5.2-colibri",
messages=[{"role": "user", "content": "What is the capital of France?"}]
)
print(resp.choices[0].message.content)
Tool Calling with GLM-5.2
import openai
import json
tools = [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Lookup current weather",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"]
}
}
}]
resp = openai.ChatCompletion.create(
model="glm-5.2-colibri",
messages=[{"role": "user", "content": "Weather in Berlin?"}],
tools=tools,
tool_choice="auto"
)
if resp.choices[0].finish_reason == "tool_calls":
call = resp.choices[0].message.tool_calls[0]
print(f"Tool invoked: {call.function.name}")
Utilizing KV Cache Slots
curl http://127.0.0.1:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "glm-5.2-colibri",
"messages": [{"role": "user", "content": "Context setting prompt"}],
"cache_slot": 1
}'
Subsequent requests with "cache_slot": 1 reuse the cached KV prefix.
Summary
- Colibri's API structure centers on
c/openai_server.py, a zero-dependency Python gateway implementing OpenAI-compatible REST endpoints including/v1/chat/completionsand/v1/models. - Streaming and tools are fully supported via Server-Sent Events terminated by
\x01\x01END\x01\x01, with tool-calling available for GLM-5.2, DeepSeek V4, and Kimi K3 through native syntax parsing inparse_arch_tool_calls. - KV slot management allows up to 16 independent cache contexts via the
cache_slotparameter, managed by theGenerationSchedulerwith configurable queue limits. - Security and deployment require no external packages; authentication is optional via
COLI_API_KEY, and CORS is configurable through environment variables.
Frequently Asked Questions
Does Colibri require external dependencies to run the API server?
No. The gateway in c/openai_server.py uses only the Python standard library. The inference engine in the c/ directory is a standalone binary with zero external dependencies, making deployment entirely self-contained.
How does Colibri handle streaming responses?
Streaming uses Server-Sent Events (SSE) with a custom termination marker \x01\x01END\x01\x01 defined in openai_server.py. The handle_chat_completion function (lines 350‑420) manages the SSE stream, chunking tokens as they generate from the C engine subprocess and flushing them to the client until the end marker is reached.
Can I use the standard OpenAI Python SDK with Colibri?
Yes. Point the SDK to http://localhost:8000/v1 and use any non-empty API key (such as "local"). The gateway accepts standard parameters like temperature, max_tokens, and tools, forwarding them to the native engine without modification.
What happens when tool-calling is requested on an unsupported engine?
Engines without native tool support (Inkling, Qwen 3.8, OLMoE) return HTTP 400 Bad Request when the tools parameter is present. For supported engines, the gateway parses native formats like <tool_call> or DSML blocks in parse_arch_tool_calls (lines 620‑640) and converts them to OpenAI-compatible JSON structured responses.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →