How CBM Represents Code Architecture for Analysis: Inside the Knowledge Graph

CBM represents code architecture as a structured knowledge graph built from C-level data structures that capture syntactic and semantic elements extracted via Tree-sitter and hybrid LSP resolution, storing results in arena-allocated arrays before persisting to a SQLite-backed graph.

The Codebase Memory (CBM) MCP server from DeusData/codebase-memory-mcp transforms raw source trees into a rich, queryable semantic model. Unlike traditional static analysis tools that rely solely on AST traversal, CBM employs a two-layer extraction pipeline that bridges syntactic parsing with lightweight type resolution to build a comprehensive architecture graph.

The Core C Data Structures

At the heart of CBM's representation lies a set of carefully designed C structs defined in internal/cbm/cbm.h. These structures serve as the intermediate representation between raw source code and the final graph database.

Language Identification

Every analyzed file is tagged with a language identifier from the CBMLanguage enum. Defined in cbm.h, this enumeration supports over 150 programming languages:

typedef enum {
    CBM_LANG_GO = 0,
    CBM_LANG_PYTHON,
    CBM_LANG_JAVASCRIPT,
    …,
    CBM_LANG_COUNT
} CBMLanguage;

Node Type Definitions

The primary graph nodes materialize through several key structs (lines 78-115 of cbm.h):

  • CBMDefinition: Models functions, classes, methods, variables, and modules with fields including name, qualified_name, signature, return_type, is_exported, and structural_profile.

  • CBMCall: Captures invocation expressions with callee_name, enclosing_func_qn, argument arrays, and control-flow context like loop_depth and branch_depth.

  • CBMImport, CBMUsage, CBMTypeRef, and domain-specific types like CBMInfraBinding (for URLs and config keys) and CBMChannel (for pub/sub channels) complete the semantic picture.

Arena-AAllocated Arrays

During extraction, each node type accumulates in growable arrays backed by a custom arena allocator:

typedef struct { CBMDefinition *items; int count; int cap; } CBMDefArray;
typedef struct { CBMCall *items;        int count; int cap; } CBMCallArray;

These arrays aggregate into the CBMFileResult struct (lines 224-274 of cbm.h), the atomic unit of extraction output:

typedef struct {
    CBMArena arena;               // owns all strings
    CBMDefArray    defs;
    CBMCallArray   calls;
    CBMImportArray imports;
    const char *module_qn;        // qualified module name
    const char *namespace_name;   // package/namespace
    bool has_error;               // parse-error flag
    TSTree *cached_tree;          // retained Tree-sitter tree
    CBMLanguage cached_lang;      // language of the file
    const char *source;           // retained source bytes
    int source_len;
} CBMFileResult;

The cbm_extract_file function defined in internal/cbm/cbm.c populates this structure, returning a pointer to a CBMFileResult containing all discovered entities for a single source file.

Two-Layer Extraction Architecture

CBM's representation emerges from a dual-pass process described in the README (lines 438-465).

Pass 1: Syntactic Tree-Sitter Analysis

The first pass performs a fast syntactic walk using Tree-sitter grammars. This populates the basic CBMFileResult arrays with raw definitions, imports, and unresolved call expressions. The parser identifies structural elements without requiring compilation units or build system integration.

Pass 2: Hybrid LSP Resolution

The second pass implements a lightweight, C-based type-resolution layer that enriches the graph with semantic information. This Hybrid LSP pass resolves method dispatch targets, computes fully-qualified names for callees, and adds semantic edge types like HTTP_CALLS and ASYNC_CALLS to the CBMCall structures.

From Structs to Graph: The Transformation Pipeline

After per-file extraction, CBM merges the individual CBMFileResult instances into a global SQLite-backed graph database. The transformation follows a predictable mapping:

  • defs arrays become graph nodes labeled :Function, :Class, or :Variable, inheriting properties like qualified_name and signature.

  • calls arrays generate CALLS edges (syntactic) and RESOLVED_CALLS edges (semantic), linking enclosing_func_qn to resolved callee nodes.

  • imports arrays create IMPORTS edges between module nodes.

  • CBMChannel and CBMInfraBinding instances materialize as EMITS, LISTENS_ON, and HTTP_CALLS edges connecting runtime entities.

Querying Architecture with get_architecture

The get_architecture tool traverses this graph to produce high-level visualizations and analysis. It aggregates language statistics, package hierarchies, HTTP routes, and entry points while running Louvain community detection to identify architectural clusters.

CLI Usage

Index a repository and query its architecture:


# Index the repository

codebase-memory-mcp cli index_repository '{"repo_path":"/path/to/repo"}'

# Retrieve architecture overview

codebase-memory-mcp cli get_architecture '{"project":"my-repo"}'

C API Integration

Extract file-level data programmatically using the public API defined in cbm.h:

#include "cbm.h"

int main(void) {
    cbm_init();
    
    const char *src = "package main\nfunc Add(a, b int) int { return a + b }";
    int src_len = strlen(src);
    
    CBMFileResult *res = cbm_extract_file(
        src, src_len, CBM_LANG_GO,
        "myproj", "example.go", 0, NULL, NULL);
    
    if (res->defs.count > 0) {
        CBMDefinition *def = &res->defs.items[0];
        printf("Function %s at %s:%u-%u\n",
               def->name, def->file_path, 
               def->start_line, def->end_line);
    }
    
    cbm_free_result(res);
    cbm_shutdown();
    return 0;
}

Summary

  • CBM represents code architecture through a structured knowledge graph built from C-level data structures including CBMDefinition, CBMCall, and CBMFileResult.
  • The two-layer extraction pipeline combines Tree-sitter parsing for syntax with Hybrid LSP resolution for semantics.
  • Arena-allocated arrays in CBMFileResult serve as the intermediate representation before merging into a SQLite-backed graph.
  • File paths like internal/cbm/cbm.h (lines 78-115 and 224-274) define the core data model used throughout the system.
  • The get_architecture CLI tool and C API (cbm_extract_file) provide programmable access to the graph for analysis and visualization.

Frequently Asked Questions

How does CBM handle multiple programming languages in the same repository?

CBM uses the CBMLanguage enum in cbm.h to tag every file with one of 150+ supported languages. The extraction pipeline automatically selects the appropriate Tree-sitter grammar and language-specific extraction logic based on this identifier, allowing the graph to represent polyglot architectures uniformly.

What is the difference between CALLS and RESOLVED_CALLS edges in the CBM graph?

CALLS edges represent syntactic invocation relationships discovered during the initial Tree-sitter pass, linking a function to the literal callee name found in source code. RESOLVED_CALLS edges are added during the Hybrid LSP pass, connecting the caller to fully-qualified target names with resolved method dispatch and type information.

Can I extract architectural data from CBM without using the CLI?

Yes. The cbm_extract_file function in internal/cbm/cbm.c provides a direct C API for programmatic extraction. You initialize the library with cbm_init, pass source code buffers to cbm_extract_file, and receive a CBMFileResult containing populated arrays of definitions, calls, and imports that you can process without invoking the SQLite graph layer.

Where does CBM store the extracted architecture graph?

After processing individual files into CBMFileResult structures, CBM merges these into a SQLite-backed graph database. This persistent store enables the get_architecture queries and supports the 3-D visualization UI served via the embedded HTTP server on port 9749.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →