What Indexing Is Supported for Graph Nodes and Edges in codebase-memory-mcp

codebase-memory-mcp indexes repositories into a persistent knowledge graph using 13+ node labels and 15+ edge types, storing the data in SQLite and exposing it through Cypher-like queries.

The codebase-memory-mcp project builds a persistent knowledge graph of your repository, enabling semantic code search through structured indexing of graph nodes and edges. During the indexing phase, the engine extracts a fixed taxonomy of structural elements and their relationships, persisting them to a local SQLite database for fast querying.

Supported Node Labels for Codebase Indexing

The index contains a fixed taxonomy of node labels covering every structural element of a codebase, plus infrastructure-as-code artifacts. According to the README (lines 3375‑3379), the supported node labels include:

Node label Meaning
Project Top-level repository
Package Language package (npm, Go module, Maven artifact, etc.)
Folder Directory in the source tree
File Individual source file
Module Language-specific module (Python module, TS/JS module, etc.)
Class Class / struct / record definition
Function Free-standing function
Method Method belonging to a class
Interface Interface / trait definition
Enum Enumeration type
Type Alias, generic, or other type definition
Route HTTP / gRPC / GraphQL endpoint
Resource K8s resource, Docker image, etc.

These labels are defined in the README's Node Labels section and implemented in the pipeline logic under src/pipeline/.

Supported Edge Types and Relationships

Edges capture the relationships that tie nodes together, enabling traversal from callers to callees, modules to imports, and routes to handlers. The edge taxonomy appears in the README at lines 3390‑3394 and includes:

Edge type Relationship
CONTAINS_PACKAGE Project → Package
CONTAINS_FOLDER Folder → Sub‑folder
CONTAINS_FILE Folder → File
DEFINES Symbol → its definition (e.g., Function → its AST)
DEFINES_METHOD Class → Method
IMPORTS Symbol → imported module/package
CALLS Function/Method → called symbol
HTTP_CALLS / ASYNC_CALLS HTTP / async call across services
IMPLEMENTS Class → Interface
HANDLES Route → handler function
USAGE, CONFIGURES, WRITES Data‑flow relationships
MEMBER_OF Method → Class
TESTS Test function → target under test
USES_TYPE Symbol → type it manipulates
FILE_CHANGES_WITH File ↔ change‑set (used by the watcher)

The storage layer in src/store/ persists these edges in SQLite tables optimized for fast graph traversal.

How the Graph Indexing Pipeline Works

The indexing process runs through four distinct phases to build the queryable graph, as described in the README section "Indexing pipeline" (lines 777‑783).

Tree-sitter AST Extraction

The engine performs a Tree-sitter pass that extracts syntactic ASTs for all 158 vendored languages. This creates the initial DEFINES and DEFINES_METHOD edges linking symbols to their source locations.

Hybrid LSP Resolution

A Hybrid LSP pass refines the graph with type-aware resolution, identifying imports, generics, and inheritance. This phase strengthens CALLS edges and establishes IMPLEMENTS and USES_TYPE relationships that static analysis alone cannot determine.

SQLite Persistence and Compression

The graph is assembled in RAM, compressed with LZ4, then dumped to a SQLite database located at ~/.cache/codebase-memory-mcp/…. The src/store/ directory handles the node/edge tables, while a background watcher updates indices incrementally on git changes via FILE_CHANGES_WITH edges.

Querying Indexed Graph Nodes and Edges

Once indexed, the graph supports three query interfaces operating on the node and edge model:

Structural search with search_graph filters by label, name regex, file path, or degree:

codebase-memory-mcp cli search_graph '{"label":"Function","name_pattern":"^handle.*"}'

Cypher-like queries with query_graph runs read-only openCypher against the SQLite-backed graph:

codebase-memory-mcp cli query_graph '{"query":"MATCH (f:Function)-[:CALLS]->(g) WHERE f.name=\"processOrder\" RETURN g.name"}'

Schema introspection with get_graph_schema reports node counts, edge counts, and property definitions without reading the full dataset.

The query engine lives in src/cypher/ and operates directly on the tables defined in src/store/.

Summary

  • codebase-memory-mcp indexes repositories into a persistent knowledge graph with 13+ node types and 15+ relationship types.
  • Node labels cover structural elements from Project down to Resource, while edge types capture containment, calls, imports, and data flow.
  • The pipeline uses Tree-sitter and Hybrid LSP passes, storing results in LZ4-compressed SQLite at ~/.cache/codebase-memory-mcp/.
  • Query the indexed graph via search_graph, query_graph (openCypher), or get_graph_schema APIs.

Frequently Asked Questions

What node labels does codebase-memory-mcp index?

The engine indexes 13 structural node labels including Project, Package, Class, Function, Method, Interface, Enum, Type, Route, and Resource. These labels cover language constructs and infrastructure artifacts, defined in the README at lines 3375‑3379.

How are relationships between code elements stored?

Relationships are stored as typed edges in a SQLite database. Key edge types include CALLS for function invocations, IMPORTS for module dependencies, IMPLEMENTS for interface adherence, and HANDLES for route-to-function mappings. The schema supports 15+ edge types defined at lines 3390‑3394 of the README.

Can I query the graph using standard Cypher syntax?

Yes, the query_graph CLI command accepts read-only openCypher syntax. You can match nodes by label, traverse edges using -[rel:TYPE]-> notation, and filter with WHERE clauses. The query executor is implemented in src/cypher/.

Where is the indexed graph data stored locally?

The indexed graph persists as a SQLite database in ~/.cache/codebase-memory-mcp/, compressed with LZ4. The storage implementation in src/store/ manages node tables, edge tables, and incremental updates triggered by the file watcher.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →