Graph Storage Dependencies in Codebase-Memory-MCP: Internal, Third-Party, and Standard Library Components

The graph storage functionality in DeusData/codebase-memory-mCP relies on internal foundation libraries for data structures, the store abstraction layer for persistence, third-party libraries yyjson and SQLite3 for JSON parsing and database operations, and standard C library headers for atomic operations and memory management.

The graph storage subsystem serves as the in-memory indexing engine for the DeusData/codebase-memory-mcp project, maintaining nodes and edges that represent code entities and their relationships. Implemented as a pure C component within src/graph_buffer, this subsystem minimizes external dependencies while leveraging high-performance internal libraries for hash-based indexing and dynamic array management. Understanding these graph storage dependencies is critical for developers contributing to the MCP server or integrating the buffer into custom pipelines.

Internal Foundation Dependencies

The graph buffer builds upon a suite of internal foundation headers that provide core data structures and cross-platform utilities.

Hash Tables and Dynamic Arrays

At the core of the indexing mechanism are two primary data structure dependencies:

  • src/foundation/hash_table.h — Provides hash tables that enable O(1) lookups of nodes by qualified name and edges by composite key. These hash tables serve as the primary and secondary indexes for the graph.

  • src/foundation/dyn_array.h — Implements dynamic arrays that store ordered lists of node and edge pointers. These arrays grow automatically as the graph expands during indexing operations.

System Utilities and Memory Management

The subsystem includes several foundational utility headers that handle cross-cutting concerns:

External Storage and Parsing Libraries

Beyond internal foundations, the graph storage integrates third-party libraries for data serialization and persistence.

SQLite Persistence Layer

Persistent storage operates through two related dependencies:

  • src/store/store.h — Implements the store API that abstracts database connection handling. The graph buffer uses this layer when flushing to persistent storage via cbm_gbuf_flush_to_store().

  • vendored/sqlite3/sqlite3.c (and sqlite3.h) — The bundled SQLite3 engine provides the final persistence mechanism. The dump routine cbm_gbuf_dump_to_sqlite() writes the complete in-memory graph structure to a SQLite database file.

JSON Processing with yyjson

During graph dump operations, the system requires fast JSON parsing to extract properties from edge metadata:

  • vendored/yyjson/yyjson.h — This header-only library provides fast JSON parsing required for extracting url_path and local_name values from edge property JSON strings during the dump phase.

Standard Library Requirements

The implementation relies on standard C library headers for fundamental language features:

  • <stdbool.h> and <stdint.h> — Boolean types and fixed-width integers.
  • <stdatomic.h> — Atomic counters for thread-safe, shared ID generation across nodes and edges.
  • <stdio.h>, <stdlib.h>, <string.h> — Standard I/O, memory management, and string handling.
  • <time.h> — Timestamp generation for graph operations.

Key Implementation Files

The graph storage functionality spans these primary source files:

File Purpose
src/graph_buffer/graph_buffer.h Public API and data-type definitions, including the cbm_gbuf_t structure and function declarations.
src/graph_buffer/graph_buffer.c Full implementation containing node/edge storage logic, indexing, merging algorithms, and dump routines.
src/store/store.h Store abstraction interface used by cbm_gbuf_flush_to_store().
src/foundation/hash_table.h Hash table implementation supporting the primary node index.
src/foundation/dyn_array.h Dynamic array utilities for managing node and edge pointer collections.

Working with the Graph Buffer API

The following examples demonstrate how to interact with the graph buffer using the dependencies outlined above.

Initialize a New Graph Buffer

#include "graph_buffer.h"

/* Create a buffer for a project called "myproj". */
cbm_gbuf_t *gb = cbm_gbuf_new("myproj", "/path/to/project");

Upsert a Node

int64_t node_id = cbm_gbuf_upsert_node(
    gb,
    "Function",                 /* label */
    "my_func",                  /* name */
    "myproj:src/file.c:my_func",/* qualified_name */
    "src/file.c",               /* file_path */
    10,                         /* start_line */
    20,                         /* end_line */
    "{\"is_public\":true}"      /* properties_json */
);

Insert an Edge

int64_t edge_id = cbm_gbuf_insert_edge(
    gb,
    src_node_id,                /* source_id */
    dst_node_id,                /* target_id */
    "CALLS",                    /* edge type */
    "{\"url_path\":\"/api/foo\"}" /* properties_json */
);

Dump to SQLite Database

int rc = cbm_gbuf_dump_to_sqlite(gb, "/tmp/myproj.db");
if (rc != 0) {
    /* Handle error */
}

Flush to Store Interface

cbm_store_t *store = cbm_store_open_path("/tmp/myproj.db");
int rc = cbm_gbuf_flush_to_store(gb, store);
cbm_store_close(store);

Summary

  • Graph storage dependencies in Codebase-Memory-MCP are organized into three layers: internal foundation libraries, third-party vendored code, and standard C library headers.
  • Internal foundations (hash_table.h, dyn_array.h) provide O(1) node lookups and dynamic memory expansion for graph structures.
  • Persistence dependencies include the store abstraction layer (store.h) and bundled SQLite3 for database operations.
  • JSON processing relies on the high-performance yyjson library for extracting properties during graph dumps.
  • Atomic operations from <stdatomic.h> ensure thread-safe ID generation when multiple threads access the graph buffer.

Frequently Asked Questions

What standard library headers does the graph storage functionality require?

The implementation requires <stdbool.h>, <stdint.h>, <stdatomic.h>, <stdio.h>, <stdlib.h>, <string.h>, and <time.h>. These provide essential types for boolean values, fixed-width integers, atomic counters for ID generation, and standard I/O operations.

Why does the graph buffer use yyjson instead of standard library functions?

The graph buffer uses yyjson because it provides optimized, fast JSON parsing capabilities necessary for extracting specific fields like url_path and local_name from edge property strings during the dump to SQLite. Standard C does not include JSON parsing, and yyjson offers superior performance compared to generic parsing implementations.

How does the graph storage handle concurrent ID generation?

According to the source code in src/graph_buffer/graph_buffer.c, the system utilizes atomic counters from <stdatomic.h> to manage shared ID generation for nodes and edges. This ensures thread-safe operations when multiple indexing threads request new identifiers simultaneously.

Can the graph storage operate without SQLite?

While the graph buffer maintains its entire structure in memory using internal hash tables and dynamic arrays, SQLite is required for persistence. The cbm_gbuf_dump_to_sqlite() and cbm_gbuf_flush_to_store() functions depend on the bundled SQLite3 library to write graph data to disk. Without these components, the buffer can still function for temporary in-memory operations but cannot persist data across process restarts.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →