How to Set Up Colibri Locally for Development: A Complete Guide

To set up Colibri locally for development, install build tools (GCC/Make), clone the JustVugg/colibri repository, run ./setup.sh to compile the C engine, execute make check to verify the build, and download a pre-converted model or run the conversion pipeline to begin inference.

Colibri is a lightweight, dependency-free inference engine that runs large Mixture-of-Experts (MoE) models by streaming shards from disk. This guide covers the exact steps to configure your environment, build the engine, and validate your installation using the actual source files from the repository.

Prerequisites

Before you begin, ensure your system meets these requirements:

Component Minimum Recommended
RAM ~16 GB 24 GB+
Free Disk ~380 GB (int4 container) Fast NVMe SSD
OS Linux, Windows 10/11, macOS Any
Tools C compiler, make, git, python3 —

No GPU is required; the engine runs on CPU by default. A GPU provides a speed boost but is not mandatory for development.

Install Build Tools

Linux (Ubuntu/Debian)

sudo apt update
sudo apt install -y build-essential git python3

Windows

Choose one of the following options:

  • Option A – Download a pre-built binary (no compiler needed).
  • Option B – Install MSYS2, then run:
pacman -S --needed mingw-w64-ucrt-x86_64-gcc make git python

macOS

xcode-select --install          # installs clang

brew install libomp git python  # OpenMP for multithreading

Clone and Compile the C Engine

Clone the repository and build the engine using the provided setup script:

git clone https://github.com/JustVugg/colibri.git
cd colibri/c
./setup.sh

In c/setup.sh, the script detects your compiler, builds the C engine with OpenMP support, and executes a self-test. When you see the following output, the engine is ready:


engine self-test: 32/32  (expected ~30-32/32)

Verify the Installation

Run the lightweight local checks to ensure the code builds cleanly and all unit tests pass:

make check

This command runs a portable CPU build, C unit tests, and Python standard library tests.

If you intend to work on CUDA-related code, also execute the CUDA test suite on a CUDA-capable host:

make -C c cuda-test CUDA_ARCH=native

According to the CONTRIBUTING.md source, these checks are required before submitting any pull request.

Download or Convert a Model

Download a Pre-Converted Model

The recommended GLM-5.2 int4 model (~372 GB) is hosted on Hugging Face:


https://huggingface.co/mastouri/GLM-5.2-colibri-int4-g64-with-int8-mtp

Unzip the container to a fast disk location, such as /nvme/glm52_i4 on Linux/macOS or D:\glm52_i4 on Windows.

Convert the Model Yourself

Alternatively, use the coli command to run the conversion pipeline on a raw FP8 model:

./coli convert --model /nvme/glm52_i4

The conversion process is resumable and produces the same group-scaled (gs64) container format as the pre-built version.

Run Your First Inference

Test your setup by launching an interactive chat session:


# Linux / macOS

COLI_MODEL=/nvme/glm52_i4 ./coli chat

# Windows (UCRT64 shell)

COLI_MODEL=/d/glm52_i4 ./coli chat

Additional utility commands include:

  • COLI_MODEL=/nvme/glm52_i4 ./coli doctor – Validates the environment configuration.
  • COLI_MODEL=/nvme/glm52_i4 ./coli plan – Displays RAM, disk, and GPU placement.
  • COLI_MODEL=/nvme/glm52_i4 ./coli chat --topp 0.85 – Adjusts sampling for faster token generation.

Development Workflow

Once set up, follow this workflow to contribute to the codebase:

  1. Edit source – Modify files in the c/ directory for engine changes or the colibri/ Python package for launcher and API modifications.
  2. Rebuild – After any change to C files, re-run ./setup.sh or make -C c.
  3. Run tests – Keep make check green; add new tests as needed.
  4. Submit a PR – Target the dev branch; maintainers fast-forward to main after review.

Troubleshooting Common Issues

  • Missing libgomp.so.1 on minimal cloud images – Install it via sudo apt install -y libgomp1.
  • Slow token generation – Disk speed dominates throughput; expect less than 1 token per second on slow disks.
  • ARM64 Linux builds – Pre-built binaries are only x86_64; ARM64 systems (AWS Graviton, Raspberry Pi) must build from source.

Summary

  • Install GCC, Make, Git, and Python3 for your operating system.
  • Clone JustVugg/colibri and run ./setup.sh in the c/ directory to compile the engine.
  • Verify the build with make check and CUDA-specific tests using make -C c cuda-test.
  • Download the GLM-5.2 int4 model from Hugging Face or convert your own using ./coli convert.
  • Launch inference with COLI_MODEL=/path ./coli chat and validate the environment with ./coli doctor.
  • Submit pull requests to the dev branch after ensuring all checks pass.

Frequently Asked Questions

Do I need a GPU to develop Colibri locally?

No. Colibri runs entirely on CPU by default. While a GPU can accelerate inference, it is not required for setting up the development environment or running the test suite. The ./setup.sh script and make check commands only require a C compiler and standard build tools.

What does the setup.sh script actually do?

The c/setup.sh script performs three critical tasks: it detects your system compiler, compiles the C engine with OpenMP multithreading support, and runs a 32-point self-test to verify matrix operations. When the script outputs engine self-test: 32/32, the build is successful and ready for development.

How do I convert a model for use with Colibri?

Use the coli CLI tool located in colibri/cli.py. Run ./coli convert --model /path/to/fp8 to process a raw FP8 checkpoint into the optimized int4 container format. The conversion is resumable if interrupted and produces a group-scaled (gs64) container identical to the pre-converted models available on Hugging Face.

Why is my inference speed extremely slow?

Disk I/O is the bottleneck for Colibri. The engine streams model shards from disk rather than loading them into RAM, so token throughput depends entirely on your storage speed. Slow mechanical hard drives or network-attached storage may yield less than 1 token per second, while fast NVMe SSDs provide significantly better performance.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →