codebase-memory-mcp Memory Management Strategy: RAM-First Pipeline and LZ4 Compression

codebase-memory-mcp stores and processes code-intelligence data using a RAM-first pipeline that performs all analysis passes in memory before persisting results to an LZ4-compressed blob store, achieving fast incremental indexing and compact persistent storage.

The codebase-memory-mcp repository implements a high-performance indexing system for code intelligence. According to the source code in internal/cbm/lz4_store.c and the pipeline implementation, the project employs a RAM-first pipeline coupled with LZ4 high-compression (HC) to minimize disk I/O during analysis while maintaining efficient persistent storage.

The RAM-First Pipeline Architecture

In-Memory Processing Model

The pipeline loads the entire intermediate representation of a project into RAM during the indexing phase. As implemented in pipeline/pipeline.c, the system performs all analysis passes—including syntax parsing, semantic analysis, and similarity detection—entirely in memory. Only after the in-memory work completes does the pipeline persist the final results to the on-disk store.

This design eliminates disk I/O bottlenecks during the analysis phase, allowing the system to repeatedly re-run passes (such as similarity matching and route canonicalization) without reading from or writing to persistent storage.

Worker Pool Concurrency

The concurrency model relies on a worker pool defined in pipeline/worker_pool.h. This architecture creates a set of threads that operate on shared in-memory structures, synchronizing only when persisting the final compressed blocks to the SQLite-based store. By keeping the working set memory-resident, the worker pool avoids expensive disk operations during parallel processing.

LZ4 Compression Implementation

Core Compression API in lz4_store.c

The on-disk storage layer uses an LZ4-compressed blob store implemented in internal/cbm/lz4_store.c. The C implementation exposes three critical functions:

  • cbm_lz4_compress_hc(src, srcLen, dst, dstCap) – High-compression (HC) LZ4 mode used for persisting intermediate data
  • cbm_lz4_decompress(src, srcLen, dst, originalLen) – Safe decompression that prevents buffer overruns by validating against the original size
  • cbm_lz4_bound(inputSize) – Returns the worst-case size needed for the compressed buffer, guaranteeing that malloc(cbm_lz4_bound(n)) is always sufficient

Buffer Management Strategy

Before compression, the pipeline allocates temporary buffers sized using cbm_lz4_bound(). This guarantees that the allocated memory provides sufficient space for the compressed output even in the worst-case scenario. After compression, the block is written to the SQLite-based store. A similar pattern exists in internal/cbm/zstd_store.c for projects using ZSTD compression instead of LZ4.

When the store is later read, the data is decompressed in-place into a RAM buffer supplied by the pipeline, keeping the rest of the system fully memory-resident until the final commit.

Persistence and Storage Flow

The persistence layer uses SQLite as the underlying database, with individual blobs compressed using the LZ4 HC algorithm. During the write phase, the pipeline compresses the in-memory representation and stores the resulting blobs. During reads, the system decompresses these blobs back into RAM buffers for further processing.

This architecture provides a clear separation between the high-speed in-memory processing phase and the compact storage phase, ensuring that disk I/O only occurs at the boundaries of the pipeline execution.

Practical Code Examples

Compression and Decompression

/* Allocate a compression buffer sized for the input */
int bound = cbm_lz4_bound(input_len);
char *cbuf = malloc(bound);

/* Compress the data using high-compression mode */
int c_len = cbm_lz4_compress_hc(input, input_len, cbuf, bound);
assert(c_len > 0);

/* Persist the compressed blob (e.g., into SQLite) */
store_write_blob(store, cbuf, c_len);
free(cbuf);

/* Later – read & decompress */
char *decomp = malloc(original_len);
int d_len = cbm_lz4_decompress(blob, blob_len, decomp, original_len);
assert(d_len == original_len);

Pipeline Usage

/* Initialize the RAM-first pipeline */
cbm_pipeline_t *p = cbm_pipeline_new(project_dir, db_path, CBM_MODE_FULL);

/* Run all analysis passes; everything stays in RAM until final store step */
int rc = cbm_pipeline_run(p);
assert(rc == 0);

/* The pipeline internally compresses its store blobs with LZ4 */
cbm_pipeline_free(p);

Performance Benefits

The RAM-first architecture with LZ4 compression provides two major advantages:

  • Fast Incremental Indexing: By keeping the working set in RAM, the pipeline can re-run analysis passes without incurring disk I/O, significantly reducing iteration time for large codebases.
  • Compact Persistent Storage: LZ4 HC achieves greater than 2× compression on highly repetitive source files, as verified in tests/test_lz4.c, dramatically reducing SQLite blob size while maintaining swift decompression speeds for fast loading.

Summary

  • codebase-memory-mcp uses a RAM-first pipeline that processes all code intelligence data in memory before persistence, as implemented in pipeline/pipeline.c
  • The LZ4 compression implementation in internal/cbm/lz4_store.c provides high-compression (HC) mode via cbm_lz4_compress_hc() and safe decompression via cbm_lz4_decompress()
  • Buffer allocation uses cbm_lz4_bound() to guarantee sufficient space for worst-case compression scenarios before writing to the SQLite-based store
  • The architecture enables fast incremental indexing by eliminating disk I/O during analysis passes and compact storage through efficient compression
  • Concurrency is managed through a worker pool (pipeline/worker_pool.h) that operates on shared memory structures with synchronization only at persistence boundaries

Frequently Asked Questions

How does codebase-memory-mcp ensure data integrity during compression?

The implementation uses cbm_lz4_decompress() with explicit original length validation, ensuring that decompressed data matches the expected size and preventing buffer overruns. Unit tests in tests/test_lz4.c verify round-trip correctness and compression ratios, confirming that data integrity is maintained through the compression and decompression cycle.

Why does the pipeline use LZ4 HC instead of standard LZ4?

The cbm_lz4_compress_hc() function provides higher compression ratios at the cost of slightly slower compression speed, which is optimal for the write-once-read-many pattern of code indexing. This trade-off reduces storage requirements while maintaining the decompression speed necessary for fast loading when the pipeline restarts.

Can the pipeline handle projects larger than available RAM?

The current implementation in pipeline/pipeline.c assumes the intermediate representation fits entirely in memory during the analysis phase. For projects exceeding RAM capacity, the system would require external memory management or batch processing strategies not present in the current codebase.

Where can I find integration tests for the full pipeline?

The complete RAM-first pipeline lifecycle is tested in tests/test_pipeline.c, which exercises the initialization (cbm_pipeline_new), execution (cbm_pipeline_run), and teardown phases, verifying that the LZ4 compression and in-memory processing work correctly end-to-end.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →