# How Database Management Works in Colibri: File-Based Storage Architecture

> Discover how Colibri manages its database using file-based storage architecture with JSON config files and binary shards for zero-dependency deployment and massive model streaming.

- Repository: [Vincenzo Fornaro/colibri](https://github.com/JustVugg/colibri)
- Tags: architecture
- Published: 2026-09-12

---

**Colibri abandons traditional relational database management entirely, instead using direct file-based storage with JSON configuration files, binary expert shards, and in-memory global pointers to enable zero-dependency deployment and streaming of massive models.**

The Colibri inference engine (`JustVugg/colibri`) rejects conventional database architectures in favor of a minimalist, file-centric design. Rather than relying on SQLite, PostgreSQL, or similar systems to persist model data, it treats weights, routing metadata, and configuration as ordinary files on disk. This approach eliminates database drivers and connection overhead while enabling the engine to stream parameters for models too large to fit in system RAM.

## Why Colibri Eliminates Traditional Database Management

Relational databases impose dependencies and runtime overhead that conflict with Colibri's deployment goals. By managing data through direct filesystem operations, the engine avoids schema migrations, query planning, and connection pooling. The architecture ensures that Colibri runs on minimal systems without requiring a database engine installation, reducing the attack surface and binary size.

## Configuration Storage and the Cfg Struct

### Parsing Model Configuration from JSON

The command-line entry point ([`colibri/cli.py`](https://github.com/JustVugg/colibri/blob/main/colibri/cli.py)) handles configuration by parsing JSON or binary files provided at startup. These files define critical parameters such as `ebits` and `dbits`—the bit-widths for expert and dense layers throughout the model. The CLI populates an in-memory structure rather than writing to a database.

### The Cfg Struct Definition

In [`c/colibri.c`](https://github.com/JustVugg/colibri/blob/main/c/colibri.c) at lines 3215-3220, the core engine defines a `Cfg` struct that holds these configuration values in memory during execution. The fields exist as stack or heap allocations after parsing; they are never persisted to a database backend or transaction log. This struct serves as the single source of truth for model behavior throughout the inference session.

## Expert Weight Storage and Streaming

### Per-Expert File Shards

Expert parameters are not stored in database tables or blob columns. As documented in the comment at the top of [`c/colibri.c`](https://github.com/JustVugg/colibri/blob/main/c/colibri.c) (lines 5-6), "expert shards … are stored *per‑expert* on disk." Each expert's weight matrix resides in its own binary file, allowing the engine to load only the specific parameters required for the current token.

### Direct Memory Mapping

When inference requires a particular expert, the engine opens the corresponding file directly using `mmap` or standard read operations. The following pattern from [`c/colibri.c`](https://github.com/JustVugg/colibri/blob/main/c/colibri.c) demonstrates how expert data bypasses database layers entirely:

```c
/* Example excerpt from c/colibri.c – loading a per‑expert shard */
static void load_expert_shard(int expert_id, const char *path) {
    int fd = open(path, O_RDONLY);
    if (fd < 0) { perror("open expert shard"); exit(1); }
    /* mmap the file directly – no DB calls */
    void *data = mmap(NULL, sharded_size, PROT_READ, MAP_PRIVATE, fd, 0);
    close(fd);
    /* store pointer in the in‑memory expert table */
    expert_table[expert_id] = data;
}

```

This file-based database management strategy enables Colibri to stream hundreds of gigabytes of model weights from disk without loading the entire parameter set into RAM.

## Routing Metadata Management

### The Expert Atlas JSON Files

Routing metadata and benchmarking prompts live in plain JSON files under `c/tools/expert_atlas/`. For example, [`c/tools/expert_atlas/probes.json`](https://github.com/JustVugg/colibri/blob/main/c/tools/expert_atlas/probes.json) contains an array of natural-language prompts that the engine uses to test routing decisions. These files function as a read-only metadata database.

The Python utilities load this data without any database drivers:

```python

# Example: loading an expert‑atlas JSON file (used by the CLI for benchmarking)

import json, pathlib

PROBES_PATH = pathlib.Path(__file__).parent / "expert_atlas" / "probes.json"
with PROBES_PATH.open() as f:
    probes = json.load(f)          # plain list of prompt strings

print(probes[:3])

```

This implementation in [`c/tools/datapoint.py`](https://github.com/JustVugg/colibri/blob/main/c/tools/datapoint.py) uses only the standard library `json` module, maintaining the zero-dependency philosophy.

## Runtime In-Memory State Management

During execution, Colibri manages transient state through global pointers defined in [`c/colibri.c`](https://github.com/JustVugg/colibri/blob/main/c/colibri.c) around line 48: `g_pre_idx`, `g_pre_w`, and `g_pre_keff`. These variables hold routing indices and weights derived from the on-disk expert data. Memory allocation occurs via standard heap or `mmap` calls—not through database driver APIs—ensuring minimal latency and maximum control over memory layout.

## Summary

- **Colibri uses file-based storage** rather than relational databases for all persistence needs, eliminating DBMS dependencies.
- **Expert shards** are stored as individual binary files on disk and loaded via `mmap` on demand, as implemented in [`c/colibri.c`](https://github.com/JustVugg/colibri/blob/main/c/colibri.c).
- **Configuration data** is parsed from JSON into the `Cfg` struct defined at lines 3215-3220 of [`c/colibri.c`](https://github.com/JustVugg/colibri/blob/main/c/colibri.c), residing only in memory.
- **Routing metadata** lives in plain JSON files under `c/tools/expert_atlas/` (e.g., [`probes.json`](https://github.com/JustVugg/colibri/blob/main/probes.json)), loaded by Python utilities without drivers.
- **Runtime state** is managed through global pointers like `g_pre_idx` without database abstraction layers.
- This architecture enables **streaming inference** on models hundreds of gigabytes in size while maintaining zero database dependencies.

## Frequently Asked Questions

### Does Colibri support SQLite or PostgreSQL for model storage?

No. According to the `JustVugg/colibri` source code, the engine deliberately avoids all relational database management systems. Model weights are stored as binary expert shards on disk, while configuration and routing data use JSON files. The engine accesses these directly through filesystem and `mmap` operations, not SQL queries.

### How does Colibri handle large models that exceed available RAM?

Colibri implements a streaming architecture where expert shards remain on disk until needed. As shown in [`c/colibri.c`](https://github.com/JustVugg/colibri/blob/main/c/colibri.c), the engine uses `mmap` to map specific expert files into memory on a per-token basis. This file-based database management approach allows models hundreds of gigabytes in size to run on systems with limited RAM, loading only the active expert parameters.

### Where is the routing configuration stored in Colibri?

Routing metadata is stored in plain JSON files located in `c/tools/expert_atlas/`. For instance, [`probes.json`](https://github.com/JustVugg/colibri/blob/main/probes.json) contains an array of natural-language prompts used for benchmarking routing decisions. The Python module [`c/tools/datapoint.py`](https://github.com/JustVugg/colibri/blob/main/c/tools/datapoint.py) loads these files using the standard `json` module, parsing them into native data structures without any database intermediary.

### How is configuration data loaded at runtime?

The CLI entry point ([`colibri/cli.py`](https://github.com/JustVugg/colibri/blob/main/colibri/cli.py)) parses JSON or binary configuration files and populates the `Cfg` struct defined in [`c/colibri.c`](https://github.com/JustVugg/colibri/blob/main/c/colibri.c) (lines 3215-3220). Critical fields like `ebits` and `dbits` are stored in this memory structure immediately after parsing. They are never inserted into database tables, ensuring immediate access without query overhead or connection latency.