How to Debug Colibri Applications: Engine, CLI, and UI Techniques
Use environment variables to instrument the C inference engine, monitor stderr output through the CLI wrapper, and inspect per-turn timings via the React Profiling panel to debug Colibri applications effectively.
Colibri is a high-performance inference engine developed by JustVugg/colibri that separates low-level computation in C from a React-based web interface. When you debug Colibri applications, you interact with three distinct layers: the engine core that writes diagnostics to stderr, the CLI wrapper that propagates these logs, and the Web UI that visualizes performance metrics and catches runtime errors.
Engine-Level Debugging with Environment Variables
The C core in c/colibri.c reads a rich set of environment variables early in its initialization loop (around line 8423) to configure diagnostic output. These flags control everything from token-level logging to GPU failure simulation.
Enabling Verbose Engine Output
To surface detailed diagnostics from the inference engine, set the model-specific verbose flag. For the GLM-5.3 engine implemented in c/glm53.c, the check appears at lines 1245-1247:
GLM53_VERBOSE=1 ./coli chat "Your prompt here"
This directs the engine to print detailed state information to stderr, which the CLI wrapper forwards directly to your terminal without filtering.
Dumping Per-Token Logits
For debugging model output discrepancies between builds, you can log the top-5 logits at every generation step. According to docs/ENVIRONMENT.md (lines 299-301), use either:
DEBUG_LOGITS=1 ./coli chat "What is the capital of France?"
Or:
COLI_LOGIT_DUMP=1 ./coli chat "Explain quantum computing"
These variables trigger the engine to output id:logit pairs for each token generated, enabling precise A/B comparisons between different model versions or hardware configurations.
Simulating GPU Failures
To test the CPU fallback path without waiting for hardware faults, Colibri provides an injection mechanism:
COLI_GPU_FAIL_AFTER=10 ./coli chat "Run a long generation"
This forces a GPU failure after 10 calls, exercising the error handling and CPU fallback logic implemented in the engine core.
Tracing Expert Swapping
For MoE (Mixture of Experts) models, you can monitor when experts are swapped between RAM and VRAM. As implemented in c/colibri.c (lines 8423-8426), set:
REPIN_VERBOSE=1 ./coli chat "Your prompt"
This prints timestamped logs each time an expert is loaded or evicted, helping diagnose latency spikes caused by memory pressure.
Debugging via the CLI Wrapper
The Python entry point at colibri/cli.py acts as a thin wrapper around the C engine script (c/coli). It does not buffer or filter stderr, meaning any fprintf(stderr, …) calls from the engine appear immediately in your terminal.
When running the CLI, simply prepend your chosen environment variables to the command:
COLI_MODEL=/nvme/glm52_i4 \
COLI_LOGIT_DUMP=1 \
GLM53_VERBOSE=1 \
./coli chat "Explain the difference between RAM and VRAM swapping"
The wrapper propagates the environment context to the engine process and returns exit codes unchanged, preserving the full debugging interface.
Debugging the React Web UI
Colibri’s frontend provides two primary debugging mechanisms: the Profiling panel for performance analysis and ErrorBoundary components for runtime error isolation.
Visualizing Performance with the Profiling Panel
The web/src/Profiling.tsx component aggregates per-turn timing data and renders a stacked bar chart showing execution phases. It polls the backend endpoint /api/profile every 2 seconds via the getProfile function (lines 82-95) and displays metrics including:
wall_s: Total wall-clock time per turnexpert_wait_s: Time waiting for expert loadingattention_s: Computation time for attention layersother_s: Remainder calculated as wall time minus tracked phases (lines 17-20)
To use this for debugging:
- Start the server:
COLI_MODEL=/nvme/glm52_i4 ./coli serve & - Open
http://localhost:5173 - Navigate to the Profiling tab
- Trigger generation via API or UI
The component updates every 2 seconds, showing exactly where time is spent across I/O wait, matrix multiplication, attention, and LM head operations.
Catching Runtime Errors with Error Boundaries
When React components crash during rendering, the ErrorBoundary component in web/src/ErrorBoundary.tsx (lines 16-27) intercepts the error. It displays the stack trace and provides a retry button that re-mounts the component without refreshing the entire page.
In development, you can wrap problematic components explicitly:
import { ErrorBoundary } from './ErrorBoundary';
<ErrorBoundary>
<ProblematicComponent />
</ErrorBoundary>
This prevents a single component failure from crashing the entire application state, allowing you to inspect the error context and attempt recovery.
Practical Debugging Workflows
Combine these layers to diagnose complex issues:
Example 1: Full-stack latency investigation
# Terminal 1: Start server with verbose engine logging
COLI_MODEL=/nvme/glm52_i4 \
GLM53_VERBOSE=1 \
REPIN_VERBOSE=1 \
./coli serve &
# Terminal 2: Monitor logs
tail -f /var/log/colibri.log
# Terminal 3: Trigger generation and watch UI profiling
curl -X POST http://127.0.0.1:8000/v1/chat/completions \
-H 'Authorization: Bearer local' \
-H 'Content-Type: application/json' \
-d '{"model":"glm-5.2-colibri","messages":[{"role":"user","content":"Debug me"}]}'
Example 2: Logit comparison between model versions
# Build A
COLI_LOGIT_DUMP=1 COLI_MODEL=/models/glm52_v1 ./coli chat "Test" > build_a.log
# Build B
COLI_LOGIT_DUMP=1 COLI_MODEL=/models/glm52_v2 ./coli chat "Test" > build_b.log
# Compare top-5 logits at each step
diff build_a.log build_b.log
Summary
- Engine level: Set environment variables like
GLM53_VERBOSE,COLI_LOGIT_DUMP, andREPIN_VERBOSEto instrument the C core; changes take effect immediately as variables are read inc/colibri.cduring initialization. - CLI level: The Python wrapper at
colibri/cli.pyforwards stderr directly, preserving all engine diagnostics in your terminal. - UI level: Use the Profiling panel in
web/src/Profiling.tsxto visualize per-turn timing breakdowns, and wrap components withErrorBoundaryto isolate React runtime crashes. - Documentation: Refer to
docs/ENVIRONMENT.md(lines 262-306) for the canonical list of all debugging knobs available in the current release.
Frequently Asked Questions
How do I enable debug logging for the Colibri inference engine?
Set the environment variable corresponding to your active engine (e.g., GLM53_VERBOSE=1 for GLM-5.3) or use generic flags like COLI_LOGIT_DUMP=1. The engine reads these variables early in c/colibri.c and writes diagnostics to stderr, which the CLI wrapper forwards to your console unfiltered.
Where can I view per-token logits to debug generation issues?
Enable DEBUG_LOGITS=1 or COLI_LOGIT_DUMP=1 before running the coli command. The engine will print the top-5 id:logit pairs for every generation step to stderr, allowing you to trace exactly how token probabilities evolve during inference.
What is the best way to test CPU fallback without hardware failures?
Use the injection variable COLI_GPU_FAIL_AFTER=N to force a GPU error after N calls. This exercises the CPU fallback path in the engine without requiring actual hardware faults, letting you verify error handling logic deterministically.
How do I capture React errors in the Colibri web interface?
The ErrorBoundary component in web/src/ErrorBoundary.tsx catches uncaught render exceptions and displays a fallback UI with the error message and a retry button. In development, wrap suspect components with <ErrorBoundary> to isolate crashes without refreshing the entire application.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →