# How to Set Up a Development Environment for Colibri on macOS

> Set up your Colibri development environment on macOS quickly. Install dependencies, clone the JustVugg/colibri repo, and compile for CPU and Metal GPU in minutes.

- Repository: [Vincenzo Fornaro/colibri](https://github.com/JustVugg/colibri)
- Tags: getting-started
- Published: 2026-09-12

---

**Install clang and libomp via Homebrew, clone the repository, and compile with `make -C c colibri METAL=1` to build the CPU and Metal GPU backend in minutes.**

Colibri is a tiny, pure-C inference engine that runs frontier-size MoE models. Setting up a development environment for Colibri on macOS requires only a standard C toolchain plus OpenMP, with an optional Metal SDK for GPU acceleration on Apple silicon. The entire engine compiles to a single binary under 200 KB, making iteration cycles extremely fast.

## Install Build Dependencies

The engine relies on OpenMP for parallel expert evaluation. Without it, the build falls back to single-threaded execution, which is strongly discouraged for non-trivial models.

### Install clang and libomp

```bash
brew install clang libomp

```

The [`c/setup.sh`](https://github.com/JustVugg/colibri/blob/main/c/setup.sh) script detects these dependencies and prints detected versions on macOS【c/setup.sh#L17-L23】. The `libomp` package supplies the OpenMP runtime required for multi-threaded routing.

## Clone and Build the Engine

Clone the repository and navigate to the C source directory:

```bash
git clone https://github.com/JustVugg/colibri.git
cd colibri

```

The repository layout places all engine code in `c/` and documentation in `docs/`, as described in the README【README.md#L104-L108】.

### CPU-Only Build

For development on Intel Macs or when GPU acceleration is unnecessary:

```bash
make -C c colibri ARCH=native

```

This compiles [`c/colibri.c`](https://github.com/JustVugg/colibri/blob/main/c/colibri.c), the core token-by-token routing engine that implements JIT-style weight loading, KV cache management, and speculative decoding.

### Metal-Enabled Build (Recommended for Apple Silicon)

To enable GPU-accelerated expert evaluation and fused attention:

```bash
make -C c colibri METAL=1

```

When `METAL=1` is set, the Makefile also compiles [`c/backend_metal.c`](https://github.com/JustVugg/colibri/blob/main/c/backend_metal.c) and [`c/backend_metal.h`](https://github.com/JustVugg/colibri/blob/main/c/backend_metal.h), which batch expert GEMMs into single command buffers【c/setup.sh#L38-L41】. The Metal backend requires no Xcode installation—the shaders compile at runtime【docs/metal.md#L29-L36】.

## Verify the Installation

Run the Metal self-test to ensure GPU kernels match CPU reference implementations:

```bash
cd c
make metal-test

```

This executes a numerical consistency check with documented tolerances, verifying that `backend_metal.*` produces correct results before inference.

## Download a Converted Model

Colibri requires a **pre-converted safetensors container** rather than raw HF weights. The reference GLM-5.2 int4 model requires approximately 372 GB of storage.

Download an official release to fast storage:

```bash
mkdir -p /Volumes/fast/glm52_i4
cd /Volumes/fast/glm52_i4
curl -L https://huggingface.co/mastouri/GLM-5.2-colibri-int4-g64-with-int8-mtp/resolve/main/glm52_i4.tar.gz | tar xz

```

This process is documented in the Quick-Start guide【docs/quickstart.md#L41-L47】. Alternatively, convert FP8 checkpoints manually using the Python launcher with `coli convert`【README.md#L86-L92】.

## Run Your First Inference

Execute a chat session to validate your build:

```bash

# CPU-only

COLI_MODEL=/Volumes/fast/glm52_i4 ./coli chat --ram 24

# With Metal GPU acceleration

COLI_METAL=1 COLI_MODEL=/Volumes/fast/glm52_i4 ./coli chat --ram 96

```

The `COLI_METAL` environment variable activates the Metal backend at runtime, while `--ram` controls the host memory allocated for the learned expert cache【README.md#L41-L48】.

### Tune macOS-Specific Environment Variables

Colibri auto-detects hardware capabilities, but you can override behavior via variables documented in [`docs/ENVIRONMENT.md`](https://github.com/JustVugg/colibri/blob/main/docs/ENVIRONMENT.md):

- **`COLI_METAL=1`** — Enables the Metal backend on Apple silicon.
- **`DIRECT=1`** — Uses `F_NOCACHE` via [`compat.h`](https://github.com/JustVugg/colibri/blob/main/compat.h) to bypass the OS page cache.
- **`MLOCK=-1`** — Locks the expert cache into physical RAM, bypassing the memory compressor.
- **`COLI_METAL_GEMM_MIN=100000`** — Forces large GEMMs onto the CPU for exact-match prefill.

Run `./coli tune` to benchmark your machine and persist optimal settings for subsequent runs【README.md#L55-L62】.

## Development Workflow

1. **Edit** the C source in `c/`—core logic resides in [`c/colibri.c`](https://github.com/JustVugg/colibri/blob/main/c/colibri.c) while Metal kernels live in [`c/backend_metal.c`](https://github.com/JustVugg/colibri/blob/main/c/backend_metal.c).
2. **Rebuild** with `make -C c colibri` (add `METAL=1` when testing GPU paths).
3. **Test** using `make -C c check` to catch regressions in routing or memory management.
4. **Iterate**—the compact binary size enables sub-second recompilation cycles.

The [`compat.h`](https://github.com/JustVugg/colibri/blob/main/compat.h) header ensures POSIX functions like `posix_fadvise` map correctly to macOS equivalents (e.g., `F_NOCACHE`), guaranteeing portable code without modification.

## Summary

- **Dependencies**: Install `clang` and `libomp` via Homebrew; Metal SDK is optional.
- **Build**: Use `make -C c colibri METAL=1` to compile both CPU and GPU backends.
- **Models**: Download pre-converted safetensors containers or use `coli convert` for custom checkpoints.
- **Runtime**: Set `COLI_METAL=1` and `DIRECT=1` for optimal Apple silicon performance.
- **Source**: Key files include [`c/colibri.c`](https://github.com/JustVugg/colibri/blob/main/c/colibri.c) (core engine), [`c/backend_metal.c`](https://github.com/JustVugg/colibri/blob/main/c/backend_metal.c) (GPU kernels), and [`c/setup.sh`](https://github.com/JustVugg/colibri/blob/main/c/setup.sh) (dependency checks).

## Frequently Asked Questions

### Do I need Xcode to build Colibri on macOS?

No. The Metal backend compiles shaders at runtime, so you only need the command-line tools provided by `clang` and the Metal framework libraries included with macOS. The `brew install clang` command provides sufficient toolchain coverage【docs/metal.md#L29-L36】.

### What is the difference between CPU and Metal builds?

The CPU build uses OpenMP to parallelize expert evaluation across cores, while the Metal build offloads matrix multiplications and attention blocks to the GPU via [`backend_metal.c`](https://github.com/JustVugg/colibri/blob/main/backend_metal.c). On Apple silicon, the Metal backend significantly reduces latency by overlapping I/O and compute in unified memory.

### How much RAM is required for Colibri development?

The engine itself requires negligible RAM, but hosting the GLM-5.2 int4 model requires approximately 96 GB for GPU-accelerated inference or 24 GB for CPU-only mode. Use `--ram` to cap the learned expert cache size, and set `MLOCK=-1` to prevent macOS memory compression from evicting weights.

### Can I convert Hugging Face models on macOS?

Yes. Use the Python launcher ([`coli/cli.py`](https://github.com/JustVugg/colibri/blob/main/coli/cli.py)) with the `convert` subcommand to transform FP8 checkpoints into Colibri's safetensors format. This process runs entirely on CPU and requires no GPU acceleration, making it feasible on any Mac with sufficient disk space【README.md#L86-L92】.