How Database Management Works in Colibri: File-Based Storage Architecture

Colibri abandons traditional relational database management entirely, instead using direct file-based storage with JSON configuration files, binary expert shards, and in-memory global pointers to enable zero-dependency deployment and streaming of massive models.

The Colibri inference engine (JustVugg/colibri) rejects conventional database architectures in favor of a minimalist, file-centric design. Rather than relying on SQLite, PostgreSQL, or similar systems to persist model data, it treats weights, routing metadata, and configuration as ordinary files on disk. This approach eliminates database drivers and connection overhead while enabling the engine to stream parameters for models too large to fit in system RAM.

Why Colibri Eliminates Traditional Database Management

Relational databases impose dependencies and runtime overhead that conflict with Colibri's deployment goals. By managing data through direct filesystem operations, the engine avoids schema migrations, query planning, and connection pooling. The architecture ensures that Colibri runs on minimal systems without requiring a database engine installation, reducing the attack surface and binary size.

Configuration Storage and the Cfg Struct

Parsing Model Configuration from JSON

The command-line entry point (colibri/cli.py) handles configuration by parsing JSON or binary files provided at startup. These files define critical parameters such as ebits and dbits—the bit-widths for expert and dense layers throughout the model. The CLI populates an in-memory structure rather than writing to a database.

The Cfg Struct Definition

In c/colibri.c at lines 3215-3220, the core engine defines a Cfg struct that holds these configuration values in memory during execution. The fields exist as stack or heap allocations after parsing; they are never persisted to a database backend or transaction log. This struct serves as the single source of truth for model behavior throughout the inference session.

Expert Weight Storage and Streaming

Per-Expert File Shards

Expert parameters are not stored in database tables or blob columns. As documented in the comment at the top of c/colibri.c (lines 5-6), "expert shards … are stored per‑expert on disk." Each expert's weight matrix resides in its own binary file, allowing the engine to load only the specific parameters required for the current token.

Direct Memory Mapping

When inference requires a particular expert, the engine opens the corresponding file directly using mmap or standard read operations. The following pattern from c/colibri.c demonstrates how expert data bypasses database layers entirely:

/* Example excerpt from c/colibri.c – loading a per‑expert shard */
static void load_expert_shard(int expert_id, const char *path) {
    int fd = open(path, O_RDONLY);
    if (fd < 0) { perror("open expert shard"); exit(1); }
    /* mmap the file directly – no DB calls */
    void *data = mmap(NULL, sharded_size, PROT_READ, MAP_PRIVATE, fd, 0);
    close(fd);
    /* store pointer in the in‑memory expert table */
    expert_table[expert_id] = data;
}

This file-based database management strategy enables Colibri to stream hundreds of gigabytes of model weights from disk without loading the entire parameter set into RAM.

Routing Metadata Management

The Expert Atlas JSON Files

Routing metadata and benchmarking prompts live in plain JSON files under c/tools/expert_atlas/. For example, c/tools/expert_atlas/probes.json contains an array of natural-language prompts that the engine uses to test routing decisions. These files function as a read-only metadata database.

The Python utilities load this data without any database drivers:


# Example: loading an expert‑atlas JSON file (used by the CLI for benchmarking)

import json, pathlib

PROBES_PATH = pathlib.Path(__file__).parent / "expert_atlas" / "probes.json"
with PROBES_PATH.open() as f:
    probes = json.load(f)          # plain list of prompt strings

print(probes[:3])

This implementation in c/tools/datapoint.py uses only the standard library json module, maintaining the zero-dependency philosophy.

Runtime In-Memory State Management

During execution, Colibri manages transient state through global pointers defined in c/colibri.c around line 48: g_pre_idx, g_pre_w, and g_pre_keff. These variables hold routing indices and weights derived from the on-disk expert data. Memory allocation occurs via standard heap or mmap calls—not through database driver APIs—ensuring minimal latency and maximum control over memory layout.

Summary

  • Colibri uses file-based storage rather than relational databases for all persistence needs, eliminating DBMS dependencies.
  • Expert shards are stored as individual binary files on disk and loaded via mmap on demand, as implemented in c/colibri.c.
  • Configuration data is parsed from JSON into the Cfg struct defined at lines 3215-3220 of c/colibri.c, residing only in memory.
  • Routing metadata lives in plain JSON files under c/tools/expert_atlas/ (e.g., probes.json), loaded by Python utilities without drivers.
  • Runtime state is managed through global pointers like g_pre_idx without database abstraction layers.
  • This architecture enables streaming inference on models hundreds of gigabytes in size while maintaining zero database dependencies.

Frequently Asked Questions

Does Colibri support SQLite or PostgreSQL for model storage?

No. According to the JustVugg/colibri source code, the engine deliberately avoids all relational database management systems. Model weights are stored as binary expert shards on disk, while configuration and routing data use JSON files. The engine accesses these directly through filesystem and mmap operations, not SQL queries.

How does Colibri handle large models that exceed available RAM?

Colibri implements a streaming architecture where expert shards remain on disk until needed. As shown in c/colibri.c, the engine uses mmap to map specific expert files into memory on a per-token basis. This file-based database management approach allows models hundreds of gigabytes in size to run on systems with limited RAM, loading only the active expert parameters.

Where is the routing configuration stored in Colibri?

Routing metadata is stored in plain JSON files located in c/tools/expert_atlas/. For instance, probes.json contains an array of natural-language prompts used for benchmarking routing decisions. The Python module c/tools/datapoint.py loads these files using the standard json module, parsing them into native data structures without any database intermediary.

How is configuration data loaded at runtime?

The CLI entry point (colibri/cli.py) parses JSON or binary configuration files and populates the Cfg struct defined in c/colibri.c (lines 3215-3220). Critical fields like ebits and dbits are stored in this memory structure immediately after parsing. They are never inserted into database tables, ensuring immediate access without query overhead or connection latency.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →