How to Set Up a Development Environment for Colibri on macOS
Install clang and libomp via Homebrew, clone the repository, and compile with make -C c colibri METAL=1 to build the CPU and Metal GPU backend in minutes.
Colibri is a tiny, pure-C inference engine that runs frontier-size MoE models. Setting up a development environment for Colibri on macOS requires only a standard C toolchain plus OpenMP, with an optional Metal SDK for GPU acceleration on Apple silicon. The entire engine compiles to a single binary under 200 KB, making iteration cycles extremely fast.
Install Build Dependencies
The engine relies on OpenMP for parallel expert evaluation. Without it, the build falls back to single-threaded execution, which is strongly discouraged for non-trivial models.
Install clang and libomp
brew install clang libomp
The c/setup.sh script detects these dependencies and prints detected versions on macOS【c/setup.sh#L17-L23】. The libomp package supplies the OpenMP runtime required for multi-threaded routing.
Clone and Build the Engine
Clone the repository and navigate to the C source directory:
git clone https://github.com/JustVugg/colibri.git
cd colibri
The repository layout places all engine code in c/ and documentation in docs/, as described in the README【README.md#L104-L108】.
CPU-Only Build
For development on Intel Macs or when GPU acceleration is unnecessary:
make -C c colibri ARCH=native
This compiles c/colibri.c, the core token-by-token routing engine that implements JIT-style weight loading, KV cache management, and speculative decoding.
Metal-Enabled Build (Recommended for Apple Silicon)
To enable GPU-accelerated expert evaluation and fused attention:
make -C c colibri METAL=1
When METAL=1 is set, the Makefile also compiles c/backend_metal.c and c/backend_metal.h, which batch expert GEMMs into single command buffers【c/setup.sh#L38-L41】. The Metal backend requires no Xcode installation—the shaders compile at runtime【docs/metal.md#L29-L36】.
Verify the Installation
Run the Metal self-test to ensure GPU kernels match CPU reference implementations:
cd c
make metal-test
This executes a numerical consistency check with documented tolerances, verifying that backend_metal.* produces correct results before inference.
Download a Converted Model
Colibri requires a pre-converted safetensors container rather than raw HF weights. The reference GLM-5.2 int4 model requires approximately 372 GB of storage.
Download an official release to fast storage:
mkdir -p /Volumes/fast/glm52_i4
cd /Volumes/fast/glm52_i4
curl -L https://huggingface.co/mastouri/GLM-5.2-colibri-int4-g64-with-int8-mtp/resolve/main/glm52_i4.tar.gz | tar xz
This process is documented in the Quick-Start guide【docs/quickstart.md#L41-L47】. Alternatively, convert FP8 checkpoints manually using the Python launcher with coli convert【README.md#L86-L92】.
Run Your First Inference
Execute a chat session to validate your build:
# CPU-only
COLI_MODEL=/Volumes/fast/glm52_i4 ./coli chat --ram 24
# With Metal GPU acceleration
COLI_METAL=1 COLI_MODEL=/Volumes/fast/glm52_i4 ./coli chat --ram 96
The COLI_METAL environment variable activates the Metal backend at runtime, while --ram controls the host memory allocated for the learned expert cache【README.md#L41-L48】.
Tune macOS-Specific Environment Variables
Colibri auto-detects hardware capabilities, but you can override behavior via variables documented in docs/ENVIRONMENT.md:
COLI_METAL=1— Enables the Metal backend on Apple silicon.DIRECT=1— UsesF_NOCACHEviacompat.hto bypass the OS page cache.MLOCK=-1— Locks the expert cache into physical RAM, bypassing the memory compressor.COLI_METAL_GEMM_MIN=100000— Forces large GEMMs onto the CPU for exact-match prefill.
Run ./coli tune to benchmark your machine and persist optimal settings for subsequent runs【README.md#L55-L62】.
Development Workflow
- Edit the C source in
c/—core logic resides inc/colibri.cwhile Metal kernels live inc/backend_metal.c. - Rebuild with
make -C c colibri(addMETAL=1when testing GPU paths). - Test using
make -C c checkto catch regressions in routing or memory management. - Iterate—the compact binary size enables sub-second recompilation cycles.
The compat.h header ensures POSIX functions like posix_fadvise map correctly to macOS equivalents (e.g., F_NOCACHE), guaranteeing portable code without modification.
Summary
- Dependencies: Install
clangandlibompvia Homebrew; Metal SDK is optional. - Build: Use
make -C c colibri METAL=1to compile both CPU and GPU backends. - Models: Download pre-converted safetensors containers or use
coli convertfor custom checkpoints. - Runtime: Set
COLI_METAL=1andDIRECT=1for optimal Apple silicon performance. - Source: Key files include
c/colibri.c(core engine),c/backend_metal.c(GPU kernels), andc/setup.sh(dependency checks).
Frequently Asked Questions
Do I need Xcode to build Colibri on macOS?
No. The Metal backend compiles shaders at runtime, so you only need the command-line tools provided by clang and the Metal framework libraries included with macOS. The brew install clang command provides sufficient toolchain coverage【docs/metal.md#L29-L36】.
What is the difference between CPU and Metal builds?
The CPU build uses OpenMP to parallelize expert evaluation across cores, while the Metal build offloads matrix multiplications and attention blocks to the GPU via backend_metal.c. On Apple silicon, the Metal backend significantly reduces latency by overlapping I/O and compute in unified memory.
How much RAM is required for Colibri development?
The engine itself requires negligible RAM, but hosting the GLM-5.2 int4 model requires approximately 96 GB for GPU-accelerated inference or 24 GB for CPU-only mode. Use --ram to cap the learned expert cache size, and set MLOCK=-1 to prevent macOS memory compression from evicting weights.
Can I convert Hugging Face models on macOS?
Yes. Use the Python launcher (coli/cli.py) with the convert subcommand to transform FP8 checkpoints into Colibri's safetensors format. This process runs entirely on CPU and requires no GPU acceleration, making it feasible on any Mac with sufficient disk space【README.md#L86-L92】.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →