# How to Set Up Colibri Locally for Development: A Complete Guide

> Set up Colibri locally for development with this guide. Clone the JustVugg/colibri repo, compile the C engine, and verify the build to start inference.

- Repository: [Vincenzo Fornaro/colibri](https://github.com/JustVugg/colibri)
- Tags: how-to-guide
- Published: 2026-09-12

---

**To set up Colibri locally for development, install build tools (GCC/Make), clone the JustVugg/colibri repository, run [`./setup.sh`](https://github.com/JustVugg/colibri/blob/main/./setup.sh) to compile the C engine, execute `make check` to verify the build, and download a pre-converted model or run the conversion pipeline to begin inference.**

Colibri is a lightweight, dependency-free inference engine that runs large Mixture-of-Experts (MoE) models by streaming shards from disk. This guide covers the exact steps to configure your environment, build the engine, and validate your installation using the actual source files from the repository.

## Prerequisites

Before you begin, ensure your system meets these requirements:

| Component | Minimum | Recommended |
|-----------|---------|-------------|
| **RAM** | ~16 GB | 24 GB+ |
| **Free Disk** | ~380 GB (int4 container) | Fast NVMe SSD |
| **OS** | Linux, Windows 10/11, macOS | Any |
| **Tools** | C compiler, `make`, `git`, `python3` | — |

No GPU is required; the engine runs on CPU by default. A GPU provides a speed boost but is not mandatory for development.

## Install Build Tools

### Linux (Ubuntu/Debian)

```bash
sudo apt update
sudo apt install -y build-essential git python3

```

### Windows

Choose one of the following options:

- **Option A** – Download a pre-built binary (no compiler needed).
- **Option B** – Install MSYS2, then run:

```bash
pacman -S --needed mingw-w64-ucrt-x86_64-gcc make git python

```

### macOS

```bash
xcode-select --install          # installs clang

brew install libomp git python  # OpenMP for multithreading

```

## Clone and Compile the C Engine

Clone the repository and build the engine using the provided setup script:

```bash
git clone https://github.com/JustVugg/colibri.git
cd colibri/c
./setup.sh

```

In [`c/setup.sh`](https://github.com/JustVugg/colibri/blob/main/c/setup.sh), the script detects your compiler, builds the C engine with OpenMP support, and executes a self-test. When you see the following output, the engine is ready:

```

engine self-test: 32/32  (expected ~30-32/32)

```

## Verify the Installation

Run the lightweight local checks to ensure the code builds cleanly and all unit tests pass:

```bash
make check

```

This command runs a portable CPU build, C unit tests, and Python standard library tests.

If you intend to work on CUDA-related code, also execute the CUDA test suite on a CUDA-capable host:

```bash
make -C c cuda-test CUDA_ARCH=native

```

According to the [`CONTRIBUTING.md`](https://github.com/JustVugg/colibri/blob/main/CONTRIBUTING.md) source, these checks are required before submitting any pull request.

## Download or Convert a Model

### Download a Pre-Converted Model

The recommended GLM-5.2 int4 model (~372 GB) is hosted on Hugging Face:

```

https://huggingface.co/mastouri/GLM-5.2-colibri-int4-g64-with-int8-mtp

```

Unzip the container to a fast disk location, such as `/nvme/glm52_i4` on Linux/macOS or `D:\glm52_i4` on Windows.

### Convert the Model Yourself

Alternatively, use the `coli` command to run the conversion pipeline on a raw FP8 model:

```bash
./coli convert --model /nvme/glm52_i4

```

The conversion process is resumable and produces the same group-scaled (gs64) container format as the pre-built version.

## Run Your First Inference

Test your setup by launching an interactive chat session:

```bash

# Linux / macOS

COLI_MODEL=/nvme/glm52_i4 ./coli chat

# Windows (UCRT64 shell)

COLI_MODEL=/d/glm52_i4 ./coli chat

```

Additional utility commands include:

- `COLI_MODEL=/nvme/glm52_i4 ./coli doctor` – Validates the environment configuration.
- `COLI_MODEL=/nvme/glm52_i4 ./coli plan` – Displays RAM, disk, and GPU placement.
- `COLI_MODEL=/nvme/glm52_i4 ./coli chat --topp 0.85` – Adjusts sampling for faster token generation.

## Development Workflow

Once set up, follow this workflow to contribute to the codebase:

1. **Edit source** – Modify files in the `c/` directory for engine changes or the `colibri/` Python package for launcher and API modifications.
2. **Rebuild** – After any change to C files, re-run [`./setup.sh`](https://github.com/JustVugg/colibri/blob/main/./setup.sh) or `make -C c`.
3. **Run tests** – Keep `make check` green; add new tests as needed.
4. **Submit a PR** – Target the `dev` branch; maintainers fast-forward to `main` after review.

## Troubleshooting Common Issues

- **Missing `libgomp.so.1`** on minimal cloud images – Install it via `sudo apt install -y libgomp1`.
- **Slow token generation** – Disk speed dominates throughput; expect less than 1 token per second on slow disks.
- **ARM64 Linux builds** – Pre-built binaries are only x86_64; ARM64 systems (AWS Graviton, Raspberry Pi) must build from source.

## Summary

- Install GCC, Make, Git, and Python3 for your operating system.
- Clone `JustVugg/colibri` and run [`./setup.sh`](https://github.com/JustVugg/colibri/blob/main/./setup.sh) in the `c/` directory to compile the engine.
- Verify the build with `make check` and CUDA-specific tests using `make -C c cuda-test`.
- Download the GLM-5.2 int4 model from Hugging Face or convert your own using `./coli convert`.
- Launch inference with `COLI_MODEL=/path ./coli chat` and validate the environment with `./coli doctor`.
- Submit pull requests to the `dev` branch after ensuring all checks pass.

## Frequently Asked Questions

### Do I need a GPU to develop Colibri locally?

No. Colibri runs entirely on CPU by default. While a GPU can accelerate inference, it is not required for setting up the development environment or running the test suite. The [`./setup.sh`](https://github.com/JustVugg/colibri/blob/main/./setup.sh) script and `make check` commands only require a C compiler and standard build tools.

### What does the [`setup.sh`](https://github.com/JustVugg/colibri/blob/main/setup.sh) script actually do?

The [`c/setup.sh`](https://github.com/JustVugg/colibri/blob/main/c/setup.sh) script performs three critical tasks: it detects your system compiler, compiles the C engine with OpenMP multithreading support, and runs a 32-point self-test to verify matrix operations. When the script outputs `engine self-test: 32/32`, the build is successful and ready for development.

### How do I convert a model for use with Colibri?

Use the `coli` CLI tool located in [`colibri/cli.py`](https://github.com/JustVugg/colibri/blob/main/colibri/cli.py). Run `./coli convert --model /path/to/fp8` to process a raw FP8 checkpoint into the optimized int4 container format. The conversion is resumable if interrupted and produces a group-scaled (gs64) container identical to the pre-converted models available on Hugging Face.

### Why is my inference speed extremely slow?

Disk I/O is the bottleneck for Colibri. The engine streams model shards from disk rather than loading them into RAM, so token throughput depends entirely on your storage speed. Slow mechanical hard drives or network-attached storage may yield less than 1 token per second, while fast NVMe SSDs provide significantly better performance.