How Codebase-Memory-MCP Handles Schema Evolution for Graph Data
Codebase-Memory-MCP handles schema evolution through a compatibility-probe-and-rebuild strategy that detects missing columns at runtime and triggers a full store rebuild rather than applying incremental migrations.
The graph data in Codebase-Memory-MCP persists in an SQLite database managed by the store implementation in src/store/store.c. Unlike traditional migration-based systems, this MCP server handles schema evolution by validating compatibility at startup and rebuilding the derived cache when structural changes are detected. This ensures the compiled schema always matches the database file without requiring manual migration scripts.
The Compatibility-Probe-and-Rebuild Strategy
Rather than maintaining version tables or migration scripts, the project uses a compatibility-probe-and-rebuild approach. When the software initializes a store, it probes for the presence of expected schema elements. If the probe fails, the system treats the database as incompatible and delegates rebuilding to higher-level logic.
Idempotent Schema Initialization
The init_schema function in src/store/store.c creates tables using idempotent CREATE TABLE IF NOT EXISTS statements, making the initial schema safe to run on fresh databases. The DDL defines core tables including projects, nodes, and edges with their respective columns and constraints.
const char *ddl =
"CREATE TABLE IF NOT EXISTS projects ("
" name TEXT PRIMARY KEY,"
" indexed_at TEXT NOT NULL,"
" root_path TEXT NOT NULL"
");"
"CREATE TABLE IF NOT EXISTS nodes ("
" id INTEGER PRIMARY KEY AUTOINCREMENT,"
" project TEXT NOT NULL REFERENCES projects(name) ON DELETE CASCADE,"
" label TEXT NOT NULL,"
" name TEXT NOT NULL,"
" qualified_name TEXT NOT NULL,"
" file_path TEXT DEFAULT '',"
" start_line INTEGER DEFAULT 0,"
" end_line INTEGER DEFAULT 0,"
" properties TEXT DEFAULT '{}',"
" UNIQUE(project, qualified_name)"
");"
/* …additional tables… */
;
int rc = exec_sql(s, ddl);
Source: src/store/store.c lines 219-247
Detecting Schema Changes
After creating the base schema, the system probes for new columns that may have been added in later versions. For example, when the edges table required a local_name_gen generated column for IMPORTS edges, the code added a compatibility check:
sqlite3_stmt *probe = NULL;
if (sqlite3_prepare_v2(s->db,
"SELECT local_name_gen FROM edges LIMIT 0;", CBM_NOT_FOUND,
&probe, NULL) != SQLITE_OK) {
cbm_log_warn("store.schema", "result", "incompatible",
"missing", "edges.local_name_gen");
return CBM_STORE_ERR;
}
sqlite3_finalize(probe);
Source: src/store/store.c lines 788-794
Fail-Fast Behavior and Automatic Rebuilding
When the compatibility probe fails, init_schema immediately returns CBM_STORE_ERR. Callers interpret this error as "the DB cannot be opened," prompting them to delete the existing store and recreate it from scratch. This fail-fast approach is safe because the store contains only derived information— the original source files remain unchanged, allowing the system to re-index the codebase and populate the new schema automatically.
The comment in init_schema explicitly notes that this mechanism results in "full index deletes + rebuilds it," ensuring the database always matches the expected structure defined in the current version of src/store/store.c.
Read-Only Mode and Backward Compatibility
When opening a store in read-only mode, the system skips init_schema entirely, allowing older databases to be queried for existing data. However, any write operations requiring new columns will fail, causing the caller to fall back to a re-index operation. As noted in the source code comments: "Read-only query opens skip init_schema and keep working."
Source: src/store/store.c lines 179-183
Schema Introspection via get_graph_schema
The MCP server exposes the current schema through the get_graph_schema tool, implemented in src/mcp/mcp.c. This tool calls cbm_store_get_schema to inspect the live database and return JSON describing node labels and edge types, allowing clients to discover the schema without hardcoding version details.
// From the MCP client
char *resp = cbm_mcp_handle_tool(srv, "get_graph_schema",
"{\"project\":\"my-project\"}");
printf("%s\n", resp); // JSON listing of node labels and edge types
Source: src/mcp/mcp.c lines 1446-1485
Summary
- Compatibility probing detects schema mismatches by attempting to prepare statements against new columns (e.g.,
local_name_gen). - Fail-fast error handling returns
CBM_STORE_ERRwhen probes fail, preventing data corruption from schema mismatches. - Automatic rebuilding deletes and recreates the store when incompatibility is detected, which is safe because the store is a cache of derived data.
- Read-only bypass allows querying legacy databases without triggering schema initialization.
- Runtime introspection via
get_graph_schemaenables clients to discover the current schema dynamically.
Frequently Asked Questions
What happens when the database schema is incompatible?
When init_schema detects a missing column during its compatibility probe, it returns CBM_STORE_ERR. The caller then deletes the existing database file and rebuilds it from scratch by re-indexing the source files. This ensures the database always matches the schema expected by the current software version.
Does Codebase-Memory-MCP support incremental migrations?
No. The project explicitly avoids incremental migration scripts. Instead, it uses a compatibility-probe-and-rebuild strategy where the entire store is rebuilt when schema changes are detected. This approach is viable because the store functions as a cache—it can always be regenerated from the original source code.
How can I check the current graph schema without triggering a rebuild?
Use the get_graph_schema MCP tool, which calls cbm_store_get_schema in src/mcp/mcp.c. This function inspects the live database and returns the current node labels and edge types as JSON without modifying the schema or triggering a rebuild, even on older databases.
Is data lost during a schema rebuild?
No permanent data is lost because the SQLite store contains only derived information (indexed code structure). The original source files remain intact, and the rebuild process simply re-parses the codebase to populate the new schema. However, any unsaved analysis state that exists only in the old database file will be lost when the file is deleted.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →