How CBM Represents Code Architecture for Analysis: Inside the Knowledge Graph
CBM represents code architecture as a structured knowledge graph built from C-level data structures that capture syntactic and semantic elements extracted via Tree-sitter and hybrid LSP resolution, storing results in arena-allocated arrays before persisting to a SQLite-backed graph.
The Codebase Memory (CBM) MCP server from DeusData/codebase-memory-mcp transforms raw source trees into a rich, queryable semantic model. Unlike traditional static analysis tools that rely solely on AST traversal, CBM employs a two-layer extraction pipeline that bridges syntactic parsing with lightweight type resolution to build a comprehensive architecture graph.
The Core C Data Structures
At the heart of CBM's representation lies a set of carefully designed C structs defined in internal/cbm/cbm.h. These structures serve as the intermediate representation between raw source code and the final graph database.
Language Identification
Every analyzed file is tagged with a language identifier from the CBMLanguage enum. Defined in cbm.h, this enumeration supports over 150 programming languages:
typedef enum {
CBM_LANG_GO = 0,
CBM_LANG_PYTHON,
CBM_LANG_JAVASCRIPT,
…,
CBM_LANG_COUNT
} CBMLanguage;
Node Type Definitions
The primary graph nodes materialize through several key structs (lines 78-115 of cbm.h):
-
CBMDefinition: Models functions, classes, methods, variables, and modules with fields includingname,qualified_name,signature,return_type,is_exported, andstructural_profile. -
CBMCall: Captures invocation expressions withcallee_name,enclosing_func_qn, argument arrays, and control-flow context likeloop_depthandbranch_depth. -
CBMImport,CBMUsage,CBMTypeRef, and domain-specific types likeCBMInfraBinding(for URLs and config keys) andCBMChannel(for pub/sub channels) complete the semantic picture.
Arena-AAllocated Arrays
During extraction, each node type accumulates in growable arrays backed by a custom arena allocator:
typedef struct { CBMDefinition *items; int count; int cap; } CBMDefArray;
typedef struct { CBMCall *items; int count; int cap; } CBMCallArray;
These arrays aggregate into the CBMFileResult struct (lines 224-274 of cbm.h), the atomic unit of extraction output:
typedef struct {
CBMArena arena; // owns all strings
CBMDefArray defs;
CBMCallArray calls;
CBMImportArray imports;
const char *module_qn; // qualified module name
const char *namespace_name; // package/namespace
bool has_error; // parse-error flag
TSTree *cached_tree; // retained Tree-sitter tree
CBMLanguage cached_lang; // language of the file
const char *source; // retained source bytes
int source_len;
} CBMFileResult;
The cbm_extract_file function defined in internal/cbm/cbm.c populates this structure, returning a pointer to a CBMFileResult containing all discovered entities for a single source file.
Two-Layer Extraction Architecture
CBM's representation emerges from a dual-pass process described in the README (lines 438-465).
Pass 1: Syntactic Tree-Sitter Analysis
The first pass performs a fast syntactic walk using Tree-sitter grammars. This populates the basic CBMFileResult arrays with raw definitions, imports, and unresolved call expressions. The parser identifies structural elements without requiring compilation units or build system integration.
Pass 2: Hybrid LSP Resolution
The second pass implements a lightweight, C-based type-resolution layer that enriches the graph with semantic information. This Hybrid LSP pass resolves method dispatch targets, computes fully-qualified names for callees, and adds semantic edge types like HTTP_CALLS and ASYNC_CALLS to the CBMCall structures.
From Structs to Graph: The Transformation Pipeline
After per-file extraction, CBM merges the individual CBMFileResult instances into a global SQLite-backed graph database. The transformation follows a predictable mapping:
-
defsarrays become graph nodes labeled:Function,:Class, or:Variable, inheriting properties likequalified_nameandsignature. -
callsarrays generateCALLSedges (syntactic) andRESOLVED_CALLSedges (semantic), linkingenclosing_func_qnto resolved callee nodes. -
importsarrays createIMPORTSedges between module nodes. -
CBMChannelandCBMInfraBindinginstances materialize asEMITS,LISTENS_ON, andHTTP_CALLSedges connecting runtime entities.
Querying Architecture with get_architecture
The get_architecture tool traverses this graph to produce high-level visualizations and analysis. It aggregates language statistics, package hierarchies, HTTP routes, and entry points while running Louvain community detection to identify architectural clusters.
CLI Usage
Index a repository and query its architecture:
# Index the repository
codebase-memory-mcp cli index_repository '{"repo_path":"/path/to/repo"}'
# Retrieve architecture overview
codebase-memory-mcp cli get_architecture '{"project":"my-repo"}'
C API Integration
Extract file-level data programmatically using the public API defined in cbm.h:
#include "cbm.h"
int main(void) {
cbm_init();
const char *src = "package main\nfunc Add(a, b int) int { return a + b }";
int src_len = strlen(src);
CBMFileResult *res = cbm_extract_file(
src, src_len, CBM_LANG_GO,
"myproj", "example.go", 0, NULL, NULL);
if (res->defs.count > 0) {
CBMDefinition *def = &res->defs.items[0];
printf("Function %s at %s:%u-%u\n",
def->name, def->file_path,
def->start_line, def->end_line);
}
cbm_free_result(res);
cbm_shutdown();
return 0;
}
Summary
- CBM represents code architecture through a structured knowledge graph built from C-level data structures including
CBMDefinition,CBMCall, andCBMFileResult. - The two-layer extraction pipeline combines Tree-sitter parsing for syntax with Hybrid LSP resolution for semantics.
- Arena-allocated arrays in
CBMFileResultserve as the intermediate representation before merging into a SQLite-backed graph. - File paths like
internal/cbm/cbm.h(lines 78-115 and 224-274) define the core data model used throughout the system. - The
get_architectureCLI tool and C API (cbm_extract_file) provide programmable access to the graph for analysis and visualization.
Frequently Asked Questions
How does CBM handle multiple programming languages in the same repository?
CBM uses the CBMLanguage enum in cbm.h to tag every file with one of 150+ supported languages. The extraction pipeline automatically selects the appropriate Tree-sitter grammar and language-specific extraction logic based on this identifier, allowing the graph to represent polyglot architectures uniformly.
What is the difference between CALLS and RESOLVED_CALLS edges in the CBM graph?
CALLS edges represent syntactic invocation relationships discovered during the initial Tree-sitter pass, linking a function to the literal callee name found in source code. RESOLVED_CALLS edges are added during the Hybrid LSP pass, connecting the caller to fully-qualified target names with resolved method dispatch and type information.
Can I extract architectural data from CBM without using the CLI?
Yes. The cbm_extract_file function in internal/cbm/cbm.c provides a direct C API for programmatic extraction. You initialize the library with cbm_init, pass source code buffers to cbm_extract_file, and receive a CBMFileResult containing populated arrays of definitions, calls, and imports that you can process without invoking the SQLite graph layer.
Where does CBM store the extracted architecture graph?
After processing individual files into CBMFileResult structures, CBM merges these into a SQLite-backed graph database. This persistent store enables the get_architecture queries and supports the 3-D visualization UI served via the embedded HTTP server on port 9749.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →