# How CBM Represents Code Architecture for Analysis: Inside the Knowledge Graph

> Discover how CBM represents code architecture for analysis by building a knowledge graph from C-level data structures, capturing syntactic and semantic elements for in-depth insights.

- Repository: [Martin Vogel/codebase-memory-mcp](https://github.com/DeusData/codebase-memory-mcp)
- Tags: internals
- Published: 2026-07-09

---

**CBM represents code architecture as a structured knowledge graph built from C-level data structures that capture syntactic and semantic elements extracted via Tree-sitter and hybrid LSP resolution, storing results in arena-allocated arrays before persisting to a SQLite-backed graph.**

The Codebase Memory (CBM) MCP server from DeusData/codebase-memory-mcp transforms raw source trees into a rich, queryable semantic model. Unlike traditional static analysis tools that rely solely on AST traversal, CBM employs a two-layer extraction pipeline that bridges syntactic parsing with lightweight type resolution to build a comprehensive architecture graph.

## The Core C Data Structures

At the heart of CBM's representation lies a set of carefully designed C structs defined in [`internal/cbm/cbm.h`](https://github.com/DeusData/codebase-memory-mcp/blob/main/internal/cbm/cbm.h). These structures serve as the intermediate representation between raw source code and the final graph database.

### Language Identification

Every analyzed file is tagged with a language identifier from the `CBMLanguage` enum. Defined in [`cbm.h`](https://github.com/DeusData/codebase-memory-mcp/blob/main/cbm.h), this enumeration supports over 150 programming languages:

```c
typedef enum {
    CBM_LANG_GO = 0,
    CBM_LANG_PYTHON,
    CBM_LANG_JAVASCRIPT,
    …,
    CBM_LANG_COUNT
} CBMLanguage;

```

### Node Type Definitions

The primary graph nodes materialize through several key structs (lines `78-115` of [`cbm.h`](https://github.com/DeusData/codebase-memory-mcp/blob/main/cbm.h)):

- **`CBMDefinition`**: Models functions, classes, methods, variables, and modules with fields including `name`, `qualified_name`, `signature`, `return_type`, `is_exported`, and `structural_profile`.

- **`CBMCall`**: Captures invocation expressions with `callee_name`, `enclosing_func_qn`, argument arrays, and control-flow context like `loop_depth` and `branch_depth`.

- **`CBMImport`**, **`CBMUsage`**, **`CBMTypeRef`**, and domain-specific types like **`CBMInfraBinding`** (for URLs and config keys) and **`CBMChannel`** (for pub/sub channels) complete the semantic picture.

### Arena-AAllocated Arrays

During extraction, each node type accumulates in growable arrays backed by a custom arena allocator:

```c
typedef struct { CBMDefinition *items; int count; int cap; } CBMDefArray;
typedef struct { CBMCall *items;        int count; int cap; } CBMCallArray;

```

These arrays aggregate into the **`CBMFileResult`** struct (lines `224-274` of [`cbm.h`](https://github.com/DeusData/codebase-memory-mcp/blob/main/cbm.h)), the atomic unit of extraction output:

```c
typedef struct {
    CBMArena arena;               // owns all strings
    CBMDefArray    defs;
    CBMCallArray   calls;
    CBMImportArray imports;
    const char *module_qn;        // qualified module name
    const char *namespace_name;   // package/namespace
    bool has_error;               // parse-error flag
    TSTree *cached_tree;          // retained Tree-sitter tree
    CBMLanguage cached_lang;      // language of the file
    const char *source;           // retained source bytes
    int source_len;
} CBMFileResult;

```

The `cbm_extract_file` function defined in [`internal/cbm/cbm.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/internal/cbm/cbm.c) populates this structure, returning a pointer to a `CBMFileResult` containing all discovered entities for a single source file.

## Two-Layer Extraction Architecture

CBM's representation emerges from a dual-pass process described in the README (lines `438-465`).

### Pass 1: Syntactic Tree-Sitter Analysis

The first pass performs a fast syntactic walk using Tree-sitter grammars. This populates the basic `CBMFileResult` arrays with raw definitions, imports, and unresolved call expressions. The parser identifies structural elements without requiring compilation units or build system integration.

### Pass 2: Hybrid LSP Resolution

The second pass implements a lightweight, C-based type-resolution layer that enriches the graph with semantic information. This **Hybrid LSP** pass resolves method dispatch targets, computes fully-qualified names for callees, and adds semantic edge types like `HTTP_CALLS` and `ASYNC_CALLS` to the `CBMCall` structures.

## From Structs to Graph: The Transformation Pipeline

After per-file extraction, CBM merges the individual `CBMFileResult` instances into a global SQLite-backed graph database. The transformation follows a predictable mapping:

- **`defs` arrays** become graph nodes labeled `:Function`, `:Class`, or `:Variable`, inheriting properties like `qualified_name` and `signature`.

- **`calls` arrays** generate `CALLS` edges (syntactic) and `RESOLVED_CALLS` edges (semantic), linking `enclosing_func_qn` to resolved callee nodes.

- **`imports` arrays** create `IMPORTS` edges between module nodes.

- **`CBMChannel` and `CBMInfraBinding`** instances materialize as `EMITS`, `LISTENS_ON`, and `HTTP_CALLS` edges connecting runtime entities.

## Querying Architecture with get_architecture

The `get_architecture` tool traverses this graph to produce high-level visualizations and analysis. It aggregates language statistics, package hierarchies, HTTP routes, and entry points while running Louvain community detection to identify architectural clusters.

### CLI Usage

Index a repository and query its architecture:

```bash

# Index the repository

codebase-memory-mcp cli index_repository '{"repo_path":"/path/to/repo"}'

# Retrieve architecture overview

codebase-memory-mcp cli get_architecture '{"project":"my-repo"}'

```

### C API Integration

Extract file-level data programmatically using the public API defined in [`cbm.h`](https://github.com/DeusData/codebase-memory-mcp/blob/main/cbm.h):

```c
#include "cbm.h"

int main(void) {
    cbm_init();
    
    const char *src = "package main\nfunc Add(a, b int) int { return a + b }";
    int src_len = strlen(src);
    
    CBMFileResult *res = cbm_extract_file(
        src, src_len, CBM_LANG_GO,
        "myproj", "example.go", 0, NULL, NULL);
    
    if (res->defs.count > 0) {
        CBMDefinition *def = &res->defs.items[0];
        printf("Function %s at %s:%u-%u\n",
               def->name, def->file_path, 
               def->start_line, def->end_line);
    }
    
    cbm_free_result(res);
    cbm_shutdown();
    return 0;
}

```

## Summary

- **CBM represents code architecture** through a structured knowledge graph built from C-level data structures including `CBMDefinition`, `CBMCall`, and `CBMFileResult`.
- The **two-layer extraction pipeline** combines Tree-sitter parsing for syntax with Hybrid LSP resolution for semantics.
- **Arena-allocated arrays** in `CBMFileResult` serve as the intermediate representation before merging into a SQLite-backed graph.
- **File paths** like [`internal/cbm/cbm.h`](https://github.com/DeusData/codebase-memory-mcp/blob/main/internal/cbm/cbm.h) (lines 78-115 and 224-274) define the core data model used throughout the system.
- The **`get_architecture`** CLI tool and C API (`cbm_extract_file`) provide programmable access to the graph for analysis and visualization.

## Frequently Asked Questions

### How does CBM handle multiple programming languages in the same repository?

CBM uses the `CBMLanguage` enum in [`cbm.h`](https://github.com/DeusData/codebase-memory-mcp/blob/main/cbm.h) to tag every file with one of 150+ supported languages. The extraction pipeline automatically selects the appropriate Tree-sitter grammar and language-specific extraction logic based on this identifier, allowing the graph to represent polyglot architectures uniformly.

### What is the difference between CALLS and RESOLVED_CALLS edges in the CBM graph?

`CALLS` edges represent syntactic invocation relationships discovered during the initial Tree-sitter pass, linking a function to the literal callee name found in source code. `RESOLVED_CALLS` edges are added during the Hybrid LSP pass, connecting the caller to fully-qualified target names with resolved method dispatch and type information.

### Can I extract architectural data from CBM without using the CLI?

Yes. The `cbm_extract_file` function in [`internal/cbm/cbm.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/internal/cbm/cbm.c) provides a direct C API for programmatic extraction. You initialize the library with `cbm_init`, pass source code buffers to `cbm_extract_file`, and receive a `CBMFileResult` containing populated arrays of definitions, calls, and imports that you can process without invoking the SQLite graph layer.

### Where does CBM store the extracted architecture graph?

After processing individual files into `CBMFileResult` structures, CBM merges these into a **SQLite-backed graph database**. This persistent store enables the `get_architecture` queries and supports the 3-D visualization UI served via the embedded HTTP server on port 9749.