# How to Debug Colibri Applications: Engine, CLI, and UI Techniques

> Debug Colibri applications using environment variables for the C inference engine, stderr monitoring via the CLI, and React Profiling for per-turn timings. Enhance your debugging workflow.

- Repository: [Vincenzo Fornaro/colibri](https://github.com/JustVugg/colibri)
- Tags: how-to-guide
- Published: 2026-09-12

---

**Use environment variables to instrument the C inference engine, monitor stderr output through the CLI wrapper, and inspect per-turn timings via the React Profiling panel to debug Colibri applications effectively.**

Colibri is a high-performance inference engine developed by JustVugg/colibri that separates low-level computation in C from a React-based web interface. When you debug Colibri applications, you interact with three distinct layers: the **engine core** that writes diagnostics to stderr, the **CLI wrapper** that propagates these logs, and the **Web UI** that visualizes performance metrics and catches runtime errors.

## Engine-Level Debugging with Environment Variables

The C core in [`c/colibri.c`](https://github.com/JustVugg/colibri/blob/main/c/colibri.c) reads a rich set of environment variables early in its initialization loop (around line 8423) to configure diagnostic output. These flags control everything from token-level logging to GPU failure simulation.

### Enabling Verbose Engine Output

To surface detailed diagnostics from the inference engine, set the model-specific verbose flag. For the GLM-5.3 engine implemented in [`c/glm53.c`](https://github.com/JustVugg/colibri/blob/main/c/glm53.c), the check appears at lines 1245-1247:

```bash
GLM53_VERBOSE=1 ./coli chat "Your prompt here"

```

This directs the engine to print detailed state information to stderr, which the CLI wrapper forwards directly to your terminal without filtering.

### Dumping Per-Token Logits

For debugging model output discrepancies between builds, you can log the top-5 logits at every generation step. According to [`docs/ENVIRONMENT.md`](https://github.com/JustVugg/colibri/blob/main/docs/ENVIRONMENT.md) (lines 299-301), use either:

```bash
DEBUG_LOGITS=1 ./coli chat "What is the capital of France?"

```

Or:

```bash
COLI_LOGIT_DUMP=1 ./coli chat "Explain quantum computing"

```

These variables trigger the engine to output `id:logit` pairs for each token generated, enabling precise A/B comparisons between different model versions or hardware configurations.

### Simulating GPU Failures

To test the CPU fallback path without waiting for hardware faults, Colibri provides an injection mechanism:

```bash
COLI_GPU_FAIL_AFTER=10 ./coli chat "Run a long generation"

```

This forces a GPU failure after 10 calls, exercising the error handling and CPU fallback logic implemented in the engine core.

### Tracing Expert Swapping

For MoE (Mixture of Experts) models, you can monitor when experts are swapped between RAM and VRAM. As implemented in [`c/colibri.c`](https://github.com/JustVugg/colibri/blob/main/c/colibri.c) (lines 8423-8426), set:

```bash
REPIN_VERBOSE=1 ./coli chat "Your prompt"

```

This prints timestamped logs each time an expert is loaded or evicted, helping diagnose latency spikes caused by memory pressure.

## Debugging via the CLI Wrapper

The Python entry point at [`colibri/cli.py`](https://github.com/JustVugg/colibri/blob/main/colibri/cli.py) acts as a thin wrapper around the C engine script (`c/coli`). It does not buffer or filter stderr, meaning any `fprintf(stderr, …)` calls from the engine appear immediately in your terminal.

When running the CLI, simply prepend your chosen environment variables to the command:

```bash
COLI_MODEL=/nvme/glm52_i4 \
COLI_LOGIT_DUMP=1 \
GLM53_VERBOSE=1 \
./coli chat "Explain the difference between RAM and VRAM swapping"

```

The wrapper propagates the environment context to the engine process and returns exit codes unchanged, preserving the full debugging interface.

## Debugging the React Web UI

Colibri’s frontend provides two primary debugging mechanisms: the **Profiling** panel for performance analysis and **ErrorBoundary** components for runtime error isolation.

### Visualizing Performance with the Profiling Panel

The [`web/src/Profiling.tsx`](https://github.com/JustVugg/colibri/blob/main/web/src/Profiling.tsx) component aggregates per-turn timing data and renders a stacked bar chart showing execution phases. It polls the backend endpoint `/api/profile` every 2 seconds via the `getProfile` function (lines 82-95) and displays metrics including:

- `wall_s`: Total wall-clock time per turn
- `expert_wait_s`: Time waiting for expert loading
- `attention_s`: Computation time for attention layers
- `other_s`: Remainder calculated as wall time minus tracked phases (lines 17-20)

To use this for debugging:

1. Start the server: `COLI_MODEL=/nvme/glm52_i4 ./coli serve &`
2. Open `http://localhost:5173`
3. Navigate to the **Profiling** tab
4. Trigger generation via API or UI

The component updates every 2 seconds, showing exactly where time is spent across I/O wait, matrix multiplication, attention, and LM head operations.

### Catching Runtime Errors with Error Boundaries

When React components crash during rendering, the `ErrorBoundary` component in [`web/src/ErrorBoundary.tsx`](https://github.com/JustVugg/colibri/blob/main/web/src/ErrorBoundary.tsx) (lines 16-27) intercepts the error. It displays the stack trace and provides a retry button that re-mounts the component without refreshing the entire page.

In development, you can wrap problematic components explicitly:

```tsx
import { ErrorBoundary } from './ErrorBoundary';

<ErrorBoundary>
  <ProblematicComponent />
</ErrorBoundary>

```

This prevents a single component failure from crashing the entire application state, allowing you to inspect the error context and attempt recovery.

## Practical Debugging Workflows

Combine these layers to diagnose complex issues:

**Example 1: Full-stack latency investigation**

```bash

# Terminal 1: Start server with verbose engine logging

COLI_MODEL=/nvme/glm52_i4 \
GLM53_VERBOSE=1 \
REPIN_VERBOSE=1 \
./coli serve &

# Terminal 2: Monitor logs

tail -f /var/log/colibri.log

# Terminal 3: Trigger generation and watch UI profiling

curl -X POST http://127.0.0.1:8000/v1/chat/completions \
  -H 'Authorization: Bearer local' \
  -H 'Content-Type: application/json' \
  -d '{"model":"glm-5.2-colibri","messages":[{"role":"user","content":"Debug me"}]}'

```

**Example 2: Logit comparison between model versions**

```bash

# Build A

COLI_LOGIT_DUMP=1 COLI_MODEL=/models/glm52_v1 ./coli chat "Test" > build_a.log

# Build B  

COLI_LOGIT_DUMP=1 COLI_MODEL=/models/glm52_v2 ./coli chat "Test" > build_b.log

# Compare top-5 logits at each step

diff build_a.log build_b.log

```

## Summary

- **Engine level**: Set environment variables like `GLM53_VERBOSE`, `COLI_LOGIT_DUMP`, and `REPIN_VERBOSE` to instrument the C core; changes take effect immediately as variables are read in [`c/colibri.c`](https://github.com/JustVugg/colibri/blob/main/c/colibri.c) during initialization.
- **CLI level**: The Python wrapper at [`colibri/cli.py`](https://github.com/JustVugg/colibri/blob/main/colibri/cli.py) forwards stderr directly, preserving all engine diagnostics in your terminal.
- **UI level**: Use the **Profiling** panel in [`web/src/Profiling.tsx`](https://github.com/JustVugg/colibri/blob/main/web/src/Profiling.tsx) to visualize per-turn timing breakdowns, and wrap components with `ErrorBoundary` to isolate React runtime crashes.
- **Documentation**: Refer to [`docs/ENVIRONMENT.md`](https://github.com/JustVugg/colibri/blob/main/docs/ENVIRONMENT.md) (lines 262-306) for the canonical list of all debugging knobs available in the current release.

## Frequently Asked Questions

### How do I enable debug logging for the Colibri inference engine?

Set the environment variable corresponding to your active engine (e.g., `GLM53_VERBOSE=1` for GLM-5.3) or use generic flags like `COLI_LOGIT_DUMP=1`. The engine reads these variables early in [`c/colibri.c`](https://github.com/JustVugg/colibri/blob/main/c/colibri.c) and writes diagnostics to stderr, which the CLI wrapper forwards to your console unfiltered.

### Where can I view per-token logits to debug generation issues?

Enable `DEBUG_LOGITS=1` or `COLI_LOGIT_DUMP=1` before running the `coli` command. The engine will print the top-5 `id:logit` pairs for every generation step to stderr, allowing you to trace exactly how token probabilities evolve during inference.

### What is the best way to test CPU fallback without hardware failures?

Use the injection variable `COLI_GPU_FAIL_AFTER=N` to force a GPU error after N calls. This exercises the CPU fallback path in the engine without requiring actual hardware faults, letting you verify error handling logic deterministically.

### How do I capture React errors in the Colibri web interface?

The `ErrorBoundary` component in [`web/src/ErrorBoundary.tsx`](https://github.com/JustVugg/colibri/blob/main/web/src/ErrorBoundary.tsx) catches uncaught render exceptions and displays a fallback UI with the error message and a retry button. In development, wrap suspect components with `<ErrorBoundary>` to isolate crashes without refreshing the entire application.