How to Monitor Colibri Performance: Real-Time Telemetry and Metrics Guide

Colibri streams detailed performance telemetry—including latency, token throughput, GPU utilization, and tier statistics—through its API response protocol and server logs, enabling real-time monitoring via the --log-telemetry flag and custom client parsers.

Colibri is an open-source inference engine that exposes granular performance data through a structured telemetry protocol. Understanding how to monitor Colibri performance is critical for optimizing production deployments and diagnosing latency bottlenecks. The system emits metrics directly from the inference engine through the OpenAI-compatible server implementation found in c/openai_server.py.

Core Telemetry Architecture

Colibri's monitoring system operates through three integrated layers: engine-level telemetry blocks, API aggregation, and routing-cache instrumentation.

Engine Telemetry Blocks

After each generation step, the inference engine creates a telemetry block containing timing data, token counts, and hardware usage statistics. According to docs/serve_protocol.md, this block is appended to the response stream immediately before the DONE marker, allowing clients to receive performance data alongside generated content.

These blocks include:

  • Per-turn timing: Milliseconds spent on each generation step
  • Token counters: Input and output token volumes
  • Hardware metrics: GPU memory and compute utilization percentages

API-Level Aggregation

The c/openai_server.py file implements the telemetry aggregation layer. As the server receives per-step telemetry from the engine, it aggregates these blocks into human-readable summaries displayed in the UI's "Metrics" sidebar. This aggregation produces:

  • Tier bar visualization: Distribution of inference work across model tiers
  • Per-turn breakdown: Step-by-step latency analysis
  • Aggregate throughput: Overall tokens-per-second (tok/s) calculations

Routing Telemetry

For deployments utilizing the routing cache, docs/routing-telemetry.md specifies additional telemetry records that track cache efficiency. These records describe:

  • Cache hits/misses: Routing decision effectiveness
  • Tier path tracing: The specific route taken through the tier system for each request

This data helps diagnose latency spikes caused by cache fragmentation or suboptimal routing paths.

Enabling Performance Monitoring

Colibri provides two primary mechanisms for accessing telemetry data: command-line logging and programmatic diagnostics.

Server Logging Configuration

Enable detailed telemetry output when starting the server using the --log-telemetry flag. As documented in docs/serve_protocol.md, this writes each telemetry block to standard output, allowing redirection to files or external monitoring systems like Prometheus or Grafana.


# Start server with telemetry logging enabled

colibri serve --log-telemetry

When enabled, the server emits structured log entries containing the same telemetry blocks sent to API clients, providing a persistent record for historical analysis.

Diagnostic Tools

The c/doctor.py module provides diagnostic commands that dump telemetry logs on demand. Use this utility to:

  • Extract telemetry from running processes
  • Validate telemetry stream integrity
  • Generate performance reports without interrupting active inference

Consuming Telemetry Data

Applications can parse telemetry directly from the streaming API response or integrate with dashboard tools.

Python Stream Parser

The following Python client demonstrates how to extract telemetry blocks from the streaming response, as implemented in the protocol specification:

import json
import requests

def stream_completion(prompt):
    url = "http://localhost:8000/v1/chat/completions"
    headers = {"Content-Type": "application/json"}
    payload = {
        "model": "qwen38", 
        "messages": [{"role": "user", "content": prompt}], 
        "stream": True
    }
    resp = requests.post(url, headers=headers, json=payload, stream=True)

    telemetry = []
    for line in resp.iter_lines():
        if line.startswith(b"data: "):
            data = json.loads(line[6:])
            if "choices" in data and data["choices"][0]["finish_reason"] == "stop":
                break
            if "telemetry" in data:
                telemetry.append(data["telemetry"])

    return telemetry

metrics = stream_completion("Explain quantum entanglement.")
print(json.dumps(metrics, indent=2))

This parser inspects Server-Sent Events (SSE) lines beginning with data: , extracting the telemetry key from each block while ignoring standard completion chunks.

JavaScript Dashboard Integration

For browser-based monitoring, consume the telemetry stream using EventSource and render metrics in real-time:

const evtSource = new EventSource(
  '/v1/chat/completions/stream?model=qwen38&prompt=hello'
);

evtSource.onmessage = (e) => {
  const data = JSON.parse(e.data);
  if (data.telemetry) {
    console.table(data.telemetry);
    // Update dashboard widgets with:
    // data.telemetry.timing, data.telemetry.tokens, data.telemetry.gpu_usage
  }
};

This approach enables live visualization of GPU utilization and throughput metrics without polling the server.

Interpreting Key Metrics

According to docs/api.md and model-specific documentation such as docs/qwen38.md, the following metrics indicate system health:

  • Per-turn time: Duration of individual generation steps, measured in milliseconds. Spikes indicate memory pressure or tier switching overhead.

  • Tokens per second: Aggregate throughput calculated from total output tokens divided by generation time. Baseline values vary by model—consult model documentation for expected ranges.

  • GPU utilization: Percentage of VRAM and compute units active during inference. Sustained values below 80% suggest CPU bottlenecks or inefficient batching.

  • Tier bar distribution: Workload balance across model tiers. Uneven distribution indicates routing table misconfiguration.

  • Cache hit ratios: Percentage of requests served from cached routes. Values below 60% in docs/routing-telemetry.md suggest cache size limitations or query pattern changes.

Summary

  • Colibri emits telemetry blocks after every generation step via the streaming protocol defined in docs/serve_protocol.md.
  • Enable --log-telemetry to persist metrics to standard output for external monitoring integration.
  • Parse telemetry in Python or JavaScript by inspecting the telemetry key in SSE data blocks from c/openai_server.py.
  • Monitor routing cache efficiency using specialized records in docs/routing-telemetry.md to diagnose latency spikes.
  • Compare observed metrics against baselines in model-specific documentation like docs/qwen38.md to identify performance degradation.

Frequently Asked Questions

How do I access historical telemetry data from a Colibri server?

When starting the server with the --log-telemetry flag, Colibri writes each telemetry block to standard output. Redirect this output to a file or log aggregation service to maintain historical records. The c/doctor.py utility can also extract telemetry from running processes for post-hoc analysis.

What does the "tier bar" metric indicate in Colibri's telemetry?

The tier bar shows the distribution of inference work across different model tiers in the routing system. According to docs/api.md, this visualization helps identify whether requests are hitting the appropriate cache tiers or falling back to slower paths, directly impacting latency and throughput.

Why are my tokens-per-second metrics lower than the documentation suggests?

Compare your observed telemetry against the baseline ranges specified in your model's documentation (e.g., docs/qwen38.md). Discrepancies usually indicate GPU underutilization, thermal throttling, or inefficient batching configurations. Check the GPU utilization field in the telemetry blocks to distinguish between hardware limits and software bottlenecks.

How can I distinguish cache hits from misses in the telemetry stream?

Routing telemetry records, detailed in docs/routing-telemetry.md, include specific fields indicating whether a request was served from cache or required full inference. These records appear alongside standard telemetry blocks when the routing system is active, allowing you to calculate hit ratios and identify cache effectiveness issues.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →