How to Contribute to the Colibri Project: A Complete Guide for MoE Inference Engine Developers
Submit pull requests to the dev branch, run make check to enforce zero compiler warnings and passing tests, and include detailed benchmark reports for any performance-related changes.
Colibri is a lightweight, pure-C inference engine that streams massive Mixture-of-Experts (MoE) models across a unified VRAM/RAM/NVMe hierarchy. Contributing to this minimal codebase requires understanding its unique architecture where each model family resides in a dedicated C source file while sharing common headers for I/O, tokenization, and expert caching. This guide walks you through the complete contribution workflow as implemented in the JustVugg/colibri repository.
Repository Structure and Architecture
The Colibri repository follows an intentionally minimal layout designed to keep the default CPU path dependency-free. According to the source code, the top-level structure contains:
c/– Core engine sources and shared headersweb/– Browser UI and OpenAI-compatible API gatewaydesktop/– Tauri desktop wrapperdocs/– Detailed design docs, benchmarks, and hardware guidesMakefile– Top-level build entry point
Each supported model family lives in its own single-source C file, such as c/colibri.c for the GLM-5.2 family, c/inkling.c, or c/kimi_k3.c. These files share common headers including c/st.h (safetensors index), c/quant.h (container decoders), c/tok.h (tokenizer), and c/expert_store.h (streaming cache). This architecture ensures that capabilities implemented in shared headers automatically propagate to all model families.
The Six-Step Contribution Workflow
Colibri enforces a strict workflow defined in CONTRIBUTING.md to maintain code quality and performance consistency.
Step 1: Target the Dev Branch
Always open pull requests against the dev branch. The maintainers fast-forward batches of tested PRs from dev to main only after integration testing passes. Do not submit PRs directly to main.
Step 2: Execute Local Checks
Run make check from the repository root to execute the portable CPU build, C unit tests, and Python standard-library tests. This command validates that your changes compile and function correctly on CPU-only systems.
Step 3: Verify CUDA Changes
If your contribution modifies GPU code, run make -C c cuda-test CUDA_ARCH=native on a CUDA-enabled host. This executes the CUDA test suite to ensure GPU backend integrity.
Step 4: Enforce Zero Warnings and Oracle Verification
All code must compile with 0 warnings and preserve token-exact oracle verification. Validate this using the oracle test command: SNAP=./glm_tiny TF=1 ./colibri 64 16 16. Any regression in output exactness will block merging.
Step 5: Update Documentation
Add or update markdown files under docs/ if your change impacts public APIs, environment variables, or usage workflows. The docs/ directory contains the project's reference documentation, and maintaining it ensures users understand new capabilities.
Step 6: Include Benchmark Reports
For performance-related changes, provide comprehensive benchmark data including commit hash, exact command line, hardware specifications, storage specs, warm-up policy, run count, and median throughput. This data belongs in your PR description to justify optimization changes.
Working with Shared Headers and Model Files
The engine encourages code reuse through shared headers. When adding new capabilities like novel compression schemes or routing heuristics, implement them once in a shared header such as c/quant.h or c/expert_store.h, then reference them from each model's source file. This guarantees consistency across all supported families and minimizes regression risks.
GPU backend contributions must respect the "Rule for contributors" documented in GPU_BACKENDS.md. Files such as c/backend_cuda.cu, c/backend_metal.*, and c/backend_vulkan.* contain optional GPU implementations that must maintain compatibility with the core engine's streaming architecture.
Complete Contribution Workflow Example
Follow these commands to set up your environment and submit a contribution:
# Clone the repo and switch to the dev branch
git clone https://github.com/JustVugg/colibri.git
cd colibri
git checkout dev
# Build the default CPU engine (GLM-5.2 example)
make -C c glm # compiles c/colibri.c
./c/colibri --help # verify CLI functionality
# Run the full test suite locally
make check # portable CPU build + C + Python tests
# If you modified CUDA code, test on a GPU host
make -C c cuda-test CUDA_ARCH=native
# Generate a benchmark report for performance changes
COLI_MODEL=/path/to/glm52_i4 ./coli chat --model /path/to/glm52_i4 > bench.log
# Add hardware, storage, and throughput details to the log before PR
# Create and push your feature branch
git checkout -b feature/my-new-heuristic
# ... edit files (e.g., c/expert_store.h, c/colibri.c) ...
git commit -am "Add learning-cache heuristic for expert placement"
git push origin feature/my-new-heuristic
# Open a PR on GitHub targeting the `dev` branch
Summary
- Branch Strategy: Submit all pull requests to the
devbranch, never directly tomain. - Testing: Run
make checkfor CPU validation andmake -C c cuda-test CUDA_ARCH=nativefor GPU changes. - Code Quality: Maintain 0 compiler warnings and pass token-exact oracle verification using
SNAP=./glm_tiny TF=1 ./colibri 64 16 16. - Architecture: Implement shared features in headers like
c/st.horc/quant.hto propagate across all model families. - Documentation: Update
docs/files when changing public APIs or workflows. - Benchmarking: Include commit hash, hardware specs, storage details, and median throughput for performance PRs.
Frequently Asked Questions
Which branch should I target for pull requests?
Always target the dev branch. The maintainers batch and fast-forward tested PRs from dev to main after integration testing completes. Submitting directly to main will result in your PR being rejected or redirected.
How do I verify my changes meet the code quality standards?
Run make check to ensure zero compiler warnings and passing tests. For the strict token-exact verification requirement, execute SNAP=./glm_tiny TF=1 ./colibri 64 16 16 and confirm the output matches the oracle exactly. Any compiler warnings or output deviations will block merging.
What documentation is required for new features?
If your change affects public APIs, environment variables, or user workflows, you must add or update markdown files in the docs/ directory. Reference the existing structure in docs/README.md to maintain consistency with the project's technical documentation standards.
How do I test GPU backend changes locally?
Navigate to the c/ directory and run make cuda-test CUDA_ARCH=native on a CUDA-enabled machine. This command compiles and executes the CUDA test suite against your changes. For Metal or Vulkan backends, consult the specific build targets and validation steps outlined in GPU_BACKENDS.md and the backend-specific source files in c/.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →