# How to Contribute to the Colibri Project: A Complete Guide for MoE Inference Engine Developers

> Learn how to contribute to the Colibri MoE Inference Engine. Submit pull requests, run make check, and include benchmark reports for performance changes. Get involved with JustVugg/colibri today.

- Repository: [Vincenzo Fornaro/colibri](https://github.com/JustVugg/colibri)
- Tags: how-to-guide
- Published: 2026-09-12

---

**Submit pull requests to the `dev` branch, run `make check` to enforce zero compiler warnings and passing tests, and include detailed benchmark reports for any performance-related changes.**

Colibri is a lightweight, pure-C inference engine that streams massive Mixture-of-Experts (MoE) models across a unified VRAM/RAM/NVMe hierarchy. Contributing to this minimal codebase requires understanding its unique architecture where each model family resides in a dedicated C source file while sharing common headers for I/O, tokenization, and expert caching. This guide walks you through the complete contribution workflow as implemented in the `JustVugg/colibri` repository.

## Repository Structure and Architecture

The Colibri repository follows an intentionally minimal layout designed to keep the default CPU path dependency-free. According to the source code, the top-level structure contains:

- `c/` – Core engine sources and shared headers
- `web/` – Browser UI and OpenAI-compatible API gateway
- `desktop/` – Tauri desktop wrapper
- `docs/` – Detailed design docs, benchmarks, and hardware guides
- `Makefile` – Top-level build entry point

Each supported model family lives in its own single-source C file, such as [`c/colibri.c`](https://github.com/JustVugg/colibri/blob/main/c/colibri.c) for the GLM-5.2 family, [`c/inkling.c`](https://github.com/JustVugg/colibri/blob/main/c/inkling.c), or [`c/kimi_k3.c`](https://github.com/JustVugg/colibri/blob/main/c/kimi_k3.c). These files share common headers including [`c/st.h`](https://github.com/JustVugg/colibri/blob/main/c/st.h) (safetensors index), [`c/quant.h`](https://github.com/JustVugg/colibri/blob/main/c/quant.h) (container decoders), [`c/tok.h`](https://github.com/JustVugg/colibri/blob/main/c/tok.h) (tokenizer), and [`c/expert_store.h`](https://github.com/JustVugg/colibri/blob/main/c/expert_store.h) (streaming cache). This architecture ensures that capabilities implemented in shared headers automatically propagate to all model families.

## The Six-Step Contribution Workflow

Colibri enforces a strict workflow defined in [`CONTRIBUTING.md`](https://github.com/JustVugg/colibri/blob/main/CONTRIBUTING.md) to maintain code quality and performance consistency.

### Step 1: Target the Dev Branch

Always open pull requests against the `dev` branch. The maintainers fast-forward batches of tested PRs from `dev` to `main` only after integration testing passes. Do not submit PRs directly to `main`.

### Step 2: Execute Local Checks

Run `make check` from the repository root to execute the portable CPU build, C unit tests, and Python standard-library tests. This command validates that your changes compile and function correctly on CPU-only systems.

### Step 3: Verify CUDA Changes

If your contribution modifies GPU code, run `make -C c cuda-test CUDA_ARCH=native` on a CUDA-enabled host. This executes the CUDA test suite to ensure GPU backend integrity.

### Step 4: Enforce Zero Warnings and Oracle Verification

All code must compile with **0 warnings** and preserve token-exact oracle verification. Validate this using the oracle test command: `SNAP=./glm_tiny TF=1 ./colibri 64 16 16`. Any regression in output exactness will block merging.

### Step 5: Update Documentation

Add or update markdown files under `docs/` if your change impacts public APIs, environment variables, or usage workflows. The `docs/` directory contains the project's reference documentation, and maintaining it ensures users understand new capabilities.

### Step 6: Include Benchmark Reports

For performance-related changes, provide comprehensive benchmark data including commit hash, exact command line, hardware specifications, storage specs, warm-up policy, run count, and median throughput. This data belongs in your PR description to justify optimization changes.

## Working with Shared Headers and Model Files

The engine encourages code reuse through shared headers. When adding new capabilities like novel compression schemes or routing heuristics, implement them once in a shared header such as [`c/quant.h`](https://github.com/JustVugg/colibri/blob/main/c/quant.h) or [`c/expert_store.h`](https://github.com/JustVugg/colibri/blob/main/c/expert_store.h), then reference them from each model's source file. This guarantees consistency across all supported families and minimizes regression risks.

GPU backend contributions must respect the "Rule for contributors" documented in [`GPU_BACKENDS.md`](https://github.com/JustVugg/colibri/blob/main/GPU_BACKENDS.md). Files such as `c/backend_cuda.cu`, `c/backend_metal.*`, and `c/backend_vulkan.*` contain optional GPU implementations that must maintain compatibility with the core engine's streaming architecture.

## Complete Contribution Workflow Example

Follow these commands to set up your environment and submit a contribution:

```bash

# Clone the repo and switch to the dev branch

git clone https://github.com/JustVugg/colibri.git
cd colibri
git checkout dev

# Build the default CPU engine (GLM-5.2 example)

make -C c glm          # compiles c/colibri.c

./c/colibri --help     # verify CLI functionality

# Run the full test suite locally

make check             # portable CPU build + C + Python tests

# If you modified CUDA code, test on a GPU host

make -C c cuda-test CUDA_ARCH=native

# Generate a benchmark report for performance changes

COLI_MODEL=/path/to/glm52_i4 ./coli chat --model /path/to/glm52_i4 > bench.log

# Add hardware, storage, and throughput details to the log before PR

# Create and push your feature branch

git checkout -b feature/my-new-heuristic

# ... edit files (e.g., c/expert_store.h, c/colibri.c) ...

git commit -am "Add learning-cache heuristic for expert placement"
git push origin feature/my-new-heuristic

# Open a PR on GitHub targeting the `dev` branch

```

## Summary

- **Branch Strategy**: Submit all pull requests to the `dev` branch, never directly to `main`.
- **Testing**: Run `make check` for CPU validation and `make -C c cuda-test CUDA_ARCH=native` for GPU changes.
- **Code Quality**: Maintain **0 compiler warnings** and pass token-exact oracle verification using `SNAP=./glm_tiny TF=1 ./colibri 64 16 16`.
- **Architecture**: Implement shared features in headers like [`c/st.h`](https://github.com/JustVugg/colibri/blob/main/c/st.h) or [`c/quant.h`](https://github.com/JustVugg/colibri/blob/main/c/quant.h) to propagate across all model families.
- **Documentation**: Update `docs/` files when changing public APIs or workflows.
- **Benchmarking**: Include commit hash, hardware specs, storage details, and median throughput for performance PRs.

## Frequently Asked Questions

### Which branch should I target for pull requests?

Always target the `dev` branch. The maintainers batch and fast-forward tested PRs from `dev` to `main` after integration testing completes. Submitting directly to `main` will result in your PR being rejected or redirected.

### How do I verify my changes meet the code quality standards?

Run `make check` to ensure zero compiler warnings and passing tests. For the strict token-exact verification requirement, execute `SNAP=./glm_tiny TF=1 ./colibri 64 16 16` and confirm the output matches the oracle exactly. Any compiler warnings or output deviations will block merging.

### What documentation is required for new features?

If your change affects public APIs, environment variables, or user workflows, you must add or update markdown files in the `docs/` directory. Reference the existing structure in [`docs/README.md`](https://github.com/JustVugg/colibri/blob/main/docs/README.md) to maintain consistency with the project's technical documentation standards.

### How do I test GPU backend changes locally?

Navigate to the `c/` directory and run `make cuda-test CUDA_ARCH=native` on a CUDA-enabled machine. This command compiles and executes the CUDA test suite against your changes. For Metal or Vulkan backends, consult the specific build targets and validation steps outlined in [`GPU_BACKENDS.md`](https://github.com/JustVugg/colibri/blob/main/GPU_BACKENDS.md) and the backend-specific source files in `c/`.