What Indexing Is Supported for Graph Nodes and Edges in codebase-memory-mcp
codebase-memory-mcp indexes repositories into a persistent knowledge graph using 13+ node labels and 15+ edge types, storing the data in SQLite and exposing it through Cypher-like queries.
The codebase-memory-mcp project builds a persistent knowledge graph of your repository, enabling semantic code search through structured indexing of graph nodes and edges. During the indexing phase, the engine extracts a fixed taxonomy of structural elements and their relationships, persisting them to a local SQLite database for fast querying.
Supported Node Labels for Codebase Indexing
The index contains a fixed taxonomy of node labels covering every structural element of a codebase, plus infrastructure-as-code artifacts. According to the README (lines 3375‑3379), the supported node labels include:
| Node label | Meaning |
|---|---|
| Project | Top-level repository |
| Package | Language package (npm, Go module, Maven artifact, etc.) |
| Folder | Directory in the source tree |
| File | Individual source file |
| Module | Language-specific module (Python module, TS/JS module, etc.) |
| Class | Class / struct / record definition |
| Function | Free-standing function |
| Method | Method belonging to a class |
| Interface | Interface / trait definition |
| Enum | Enumeration type |
| Type | Alias, generic, or other type definition |
| Route | HTTP / gRPC / GraphQL endpoint |
| Resource | K8s resource, Docker image, etc. |
These labels are defined in the README's Node Labels section and implemented in the pipeline logic under src/pipeline/.
Supported Edge Types and Relationships
Edges capture the relationships that tie nodes together, enabling traversal from callers to callees, modules to imports, and routes to handlers. The edge taxonomy appears in the README at lines 3390‑3394 and includes:
| Edge type | Relationship |
|---|---|
| CONTAINS_PACKAGE | Project → Package |
| CONTAINS_FOLDER | Folder → Sub‑folder |
| CONTAINS_FILE | Folder → File |
| DEFINES | Symbol → its definition (e.g., Function → its AST) |
| DEFINES_METHOD | Class → Method |
| IMPORTS | Symbol → imported module/package |
| CALLS | Function/Method → called symbol |
| HTTP_CALLS / ASYNC_CALLS | HTTP / async call across services |
| IMPLEMENTS | Class → Interface |
| HANDLES | Route → handler function |
| USAGE, CONFIGURES, WRITES | Data‑flow relationships |
| MEMBER_OF | Method → Class |
| TESTS | Test function → target under test |
| USES_TYPE | Symbol → type it manipulates |
| FILE_CHANGES_WITH | File ↔ change‑set (used by the watcher) |
The storage layer in src/store/ persists these edges in SQLite tables optimized for fast graph traversal.
How the Graph Indexing Pipeline Works
The indexing process runs through four distinct phases to build the queryable graph, as described in the README section "Indexing pipeline" (lines 777‑783).
Tree-sitter AST Extraction
The engine performs a Tree-sitter pass that extracts syntactic ASTs for all 158 vendored languages. This creates the initial DEFINES and DEFINES_METHOD edges linking symbols to their source locations.
Hybrid LSP Resolution
A Hybrid LSP pass refines the graph with type-aware resolution, identifying imports, generics, and inheritance. This phase strengthens CALLS edges and establishes IMPLEMENTS and USES_TYPE relationships that static analysis alone cannot determine.
SQLite Persistence and Compression
The graph is assembled in RAM, compressed with LZ4, then dumped to a SQLite database located at ~/.cache/codebase-memory-mcp/…. The src/store/ directory handles the node/edge tables, while a background watcher updates indices incrementally on git changes via FILE_CHANGES_WITH edges.
Querying Indexed Graph Nodes and Edges
Once indexed, the graph supports three query interfaces operating on the node and edge model:
Structural search with search_graph filters by label, name regex, file path, or degree:
codebase-memory-mcp cli search_graph '{"label":"Function","name_pattern":"^handle.*"}'
Cypher-like queries with query_graph runs read-only openCypher against the SQLite-backed graph:
codebase-memory-mcp cli query_graph '{"query":"MATCH (f:Function)-[:CALLS]->(g) WHERE f.name=\"processOrder\" RETURN g.name"}'
Schema introspection with get_graph_schema reports node counts, edge counts, and property definitions without reading the full dataset.
The query engine lives in src/cypher/ and operates directly on the tables defined in src/store/.
Summary
- codebase-memory-mcp indexes repositories into a persistent knowledge graph with 13+ node types and 15+ relationship types.
- Node labels cover structural elements from
Projectdown toResource, while edge types capture containment, calls, imports, and data flow. - The pipeline uses Tree-sitter and Hybrid LSP passes, storing results in LZ4-compressed SQLite at
~/.cache/codebase-memory-mcp/. - Query the indexed graph via
search_graph,query_graph(openCypher), orget_graph_schemaAPIs.
Frequently Asked Questions
What node labels does codebase-memory-mcp index?
The engine indexes 13 structural node labels including Project, Package, Class, Function, Method, Interface, Enum, Type, Route, and Resource. These labels cover language constructs and infrastructure artifacts, defined in the README at lines 3375‑3379.
How are relationships between code elements stored?
Relationships are stored as typed edges in a SQLite database. Key edge types include CALLS for function invocations, IMPORTS for module dependencies, IMPLEMENTS for interface adherence, and HANDLES for route-to-function mappings. The schema supports 15+ edge types defined at lines 3390‑3394 of the README.
Can I query the graph using standard Cypher syntax?
Yes, the query_graph CLI command accepts read-only openCypher syntax. You can match nodes by label, traverse edges using -[rel:TYPE]-> notation, and filter with WHERE clauses. The query executor is implemented in src/cypher/.
Where is the indexed graph data stored locally?
The indexed graph persists as a SQLite database in ~/.cache/codebase-memory-mcp/, compressed with LZ4. The storage implementation in src/store/ manages node tables, edge tables, and incremental updates triggered by the file watcher.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →