Codebase-Memory-MCP Graph Schema: Complete Guide to Node Labels and Edge Types

The Codebase-Memory-MCP knowledge graph uses seventeen distinct node labels—including Project, Package, File, Class, and Function—and fifteen edge types such as CALLS, DATA_FLOWS, and HTTP_CALLS to model software architecture, dependencies, and cross-service communication.

The DeusData/codebase-memory-mcp repository implements a Model Context Protocol (MCP) server that constructs a knowledge graph from indexed source code. Understanding the specific node labels and edge types in this graph schema is essential for writing accurate Cypher queries and interpreting the relationships within your codebase.

Node Labels in Codebase-Memory-MCP

Each entity in the graph receives a label that identifies its semantic type. The authoritative list of supported labels appears in src/ui/layout3d.c (lines 131-155), where the renderer maps string labels to visual attributes using string comparisons like if (strcmp(label, "Project") == 0).

Project Structure Nodes

These labels organize the hierarchical containment of your codebase:

  • Project: The top-level root node created when the indexer initializes a new repository.
  • Package: A language-level collection of modules or namespaces (e.g., Python packages or Go import paths).
  • Module: A language-specific module unit, such as a Go package or Python module file.
  • Folder: Represents a directory in the source tree.
  • File: A specific source file that carries additional metadata including kind and detail for partial parsing scenarios.

Code Entity Nodes

These labels represent concrete programming constructs extracted by the parser:

  • Class: Object-oriented class definitions found in languages like Java, Python, or C++.
  • Struct: C-style or Go-style struct definitions.
  • Interface: Interface definitions from Go, Java, TypeScript, or similar languages.
  • Function: Free-standing functions that are not bound to a particular class or struct.
  • Method: Functions bound to a class or struct (instance or static methods).
  • Variable: Variable definitions emitted by many language grammars (implicit in some contexts).
  • Constant: Immutable constant definitions (implicit in some grammars).
  • Enum: Enumeration type definitions (implicitly supported by specific language grammars).

Service and Concurrency Nodes

Specialized labels for distributed systems and concurrent programming:

  • Route: HTTP route or endpoint definitions used when building service-level graphs.
  • Channel: Concurrency primitives such as Go channels or async queues.
  • GRPC_CALLS, GRAPHQL_CALLS, TRPC_CALLS: Specialized labels used in some graph configurations to represent RPC entry points (note these edge-type names sometimes function as labels in specialized contexts).

Edge Types in the Graph Schema

Edge types define the relationships between nodes. The trace_path tool definition in src/mcp/mcp.c (lines 26-31) enumerates the available edge types, which the system stores in the SQLite edges table with a distinct type column.

Call and Data Flow Relationships

  • CALLS: Represents a call-site relationship where one function or method invokes another.
  • DATA_FLOWS: Models the propagation of values through function arguments, return statements, or variable assignments.
  • USES: A generic dependency edge indicating that a code entity (such as a variable) is utilized within another scope (e.g., a variable used inside a function).

Cross-Service Communication

These edges model interactions between microservices or external systems:

  • HTTP_CALLS: An HTTP request initiated from one service to another.
  • ASYNC_CALLS: Asynchronous invocations via message queues or event buses.
  • CROSS_HTTP_CALLS: HTTP calls that cross repository boundaries.
  • CROSS_ASYNC_CALLS: Cross-repository asynchronous communication.
  • CROSS_CHANNEL: Cross-repository channel usage for streaming data.
  • CROSS_GRPC_CALLS, CROSS_GRAPHQL_CALLS, CROSS_TRPC_CALLS: Cross-repository RPC calls using their respective protocols.

Structural and Dependency Edges

  • CONTAINS_FOLDER: Hierarchical edge from a project to a folder.
  • CONTAINS_FILE: Hierarchical edge from a folder to a specific file.
  • IMPORTS: Language-specific import relationships (e.g., Go import statements or Python import clauses).
  • DEPENDS_ON: General dependency edges used by the "dependencies" aspect of the graph.
  • HAS_RISK: A specialized edge that carries risk classification metadata (critical, high, medium, low).

How to Query the Graph Schema

The get_graph_schema tool, implemented in src/mcp/mcp.c, returns the complete schema as a JSON object containing two arrays: node_labels and edge_types. Internally, the implementation queries the SQLite store in src/store/store.c using SELECT label … GROUP BY label for nodes and SELECT type … GROUP BY type for edges.

Retrieving the Schema via CLI

Use the MCP CLI to fetch the current schema for a specific project:

mcp get_graph_schema --project my-repo

The command outputs structured JSON:

{
  "node_labels": ["Project","Package","Module","Folder","File","Class","Struct",
                  "Interface","Function","Method","Variable","Constant","Route"],
  "edge_types":  ["CALLS","DATA_FLOWS","HTTP_CALLS","ASYNC_CALLS",
                  "CROSS_HTTP_CALLS","CROSS_ASYNC_CALLS","CONTAINS_FOLDER",
                  "CONTAINS_FILE","IMPORTS","DEPENDS_ON"]
}

Programmatic Access with Python

Query the HTTP API to retrieve schema information programmatically:

import requests
import json

resp = requests.post(
    "http://localhost:8080/api/get_graph_schema",
    json={"project": "my-repo"}
)
schema = resp.json()

print("Available node labels:", schema["node_labels"])
print("Available edge types:", schema["edge_types"])

Example Cypher Query

Once you understand the node labels and edge types, you can write precise Cypher queries. This example finds function call relationships:

MATCH (src:Function)-[c:CALLS]->(dst:Function)
RETURN src.qualified_name AS caller, dst.qualified_name AS callee
LIMIT 20;

Summary

  • Codebase-Memory-MCP defines seventeen node labels ranging from structural entities (Project, Folder, File) to code constructs (Class, Function, Method) and service components (Route, Channel).
  • The system supports fifteen edge types including CALLS and DATA_FLOWS for code analysis, HTTP_CALLS and CROSS_* variants for service mesh visualization, and IMPORTS for dependency tracking.
  • Label definitions reside in src/ui/layout3d.c (lines 131-155), while edge type enumerations appear in src/mcp/mcp.c (lines 26-31).
  • The get_graph_schema tool queries src/store/store.c to return distinct labels and types from the SQLite backing store.
  • Understanding these schema elements enables accurate graph traversal and custom tool development against the MCP server.

Frequently Asked Questions

How does Codebase-Memory-MCP determine which node label to assign?

The indexer analyzes source code using language-specific grammars and AST parsers. When the parser encounters a construct like a function definition or class declaration, it emits a node with the corresponding label (e.g., Function or Class). The label strings are then interned in the graph buffer via src/graph_buffer/graph_buffer.c to ensure consistency across the dataset.

Can I extend the graph schema with custom node labels or edge types?

The core schema is defined in the source code at src/ui/layout3d.c for labels and src/mcp/mcp.c for edge types. Adding custom labels requires modifying the C source, recompiling the indexer, and updating the SQLite schema initialization. The HAS_RISK edge type demonstrates how the codebase supports edges with metadata, suggesting a pattern for custom relationship attributes.

What is the difference between CALLS and DATA_FLOWS edge types?

The CALLS edge represents control flow—indicating that one function invokes another—while DATA_FLOWS tracks the movement of data values through arguments, returns, or variable assignments. Use CALLS to analyze execution paths and DATA_FLOWS to trace how specific values propagate through your application, which is critical for security auditing and impact analysis.

Where are the node labels and edge types physically stored?

The graph persists in a SQLite database managed by src/store/store.c. The nodes table contains a label column, and the edges table contains a type column. When you invoke get_graph_schema, the implementation executes SELECT DISTINCT label FROM nodes and SELECT DISTINCT type FROM edges queries against this store to return the current schema state.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →