How to Parse a Code Repository into a Knowledge Graph Using Tree-sitter
The code-graph-rag repository uses a lazy-loading grammar store and a custom DefinitionProcessor to parse source code with Tree-sitter, extract typed entities, and ingest them into a knowledge graph via modular mix-ins.
Parsing a codebase into a navigable knowledge graph requires more than running a parser over every file. The vitali87/code-graph-rag project demonstrates a production-ready architecture that combines lazy grammar loading, qualified name resolution, and deferred fact extraction to build cross-language code graphs at scale. This article walks through the complete pipeline implemented in the source, from parser initialization to final relationship linking.
Architecture Overview: From Source Code to Knowledge Graph
The ingestion pipeline follows seven distinct stages, each handled by specialized components in the codebase_rag package. Understanding this flow is essential when adapting the approach to your own tools.
1. Lazy Grammar Loading with load_parsers()
The system avoids the memory cost of loading all 14 supported Tree-sitter grammars upfront. Instead, load_parsers() in codebase_rag/parser_loader.py implements a _LazyGrammarStore that caches Parser and LanguageQueries objects per language.
When parsers[lang] or queries[lang] is first accessed, _process_language() builds a tree_sitter.Language object from the compiled grammar and registers the parser. This pattern ensures that parsing Python triggers only the Python grammar, not JavaScript, Go, or C++.
from codebase_rag.parser_loader import load_parsers
from tree_sitter import Parser
# Triggers lazy load for Python only
parsers, queries = load_parsers()
python_parser: Parser = parsers["python"] # First access triggers _process_language()
The COMBINED_FUNC_CLASS_IMPORT_QUERIES constant (defined at line 17) pre-builds the Tree-sitter query strings for each language, covering functions, classes, calls, imports, and more.
2. File Walking and Qualified Name Generation
The DefinitionProcessor in codebase_rag/parsers/definition_processor.py orchestrates the repository walk. For each file, process_file() computes a qualified name (QN) for the module using:
base_module_qn()— transformssrc/util/helpers.pyintomy_project.src.util.helpers_disambiguate_module_qn()— handles name collisions when multiple files share the same relative path
These QNs become stable identifiers for graph nodes, enabling cross-file reference resolution later.
from codebase_rag.utils.path_utils import base_module_qn
from pathlib import Path
module_qn = base_module_qn(Path("src/util/helpers.py"), "my_project")
print(module_qn) # → my_project.src.util.helpers
3. Tree-sitter Parsing with Preprocessor Recovery
For each file, the processor retrieves the appropriate parser and calls parse_with_preproc_recovery() (line 21). This function wraps parser.parse(source_bytes) with special handling for C/C++ files where preprocessor directives may fragment the syntax tree.
tree = parse_with_preproc_recovery(parser, source_bytes, language)
The resulting Tree object contains the complete syntax tree for query execution.
4. Executing Tree-sitter Queries
The processor runs language-specific queries using tree_sitter.QueryCursor. Captured nodes are sorted and cached in _func_class_captures_cache to avoid re-querying during mix-in processing.
func_query = queries["python"].functions
cursor = tree.walk()
# Query execution returns named captures for function definitions, calls, etc.
5. Fact Extraction via Mix-ins
DefinitionProcessor inherits behavior from three specialized mix-ins:
FunctionIngestMixin— extracts functions, methods, and their signaturesClassIngestMixin— extracts classes, inheritance relationships, and membersJsTsIngestMixin— handles JavaScript/TypeScript-specific constructs like arrow functions and interfaces
Each mix-in creates graph entities with stable QNs. Deferred items — such as C++ macro-generated nodes or anonymous JavaScript functions — are buffered internally and flushed after the full file set is processed. This two-phase approach ensures that forward references can be resolved once all files are parsed.
6. Graph Ingestion via IngestorProtocol
Extracted facts are handed to an IngestorProtocol implementation through two batch methods:
ingestor.ensure_node_batch()— createsMODULE,FUNCTION,CLASSnodesingestor.ensure_relationship_batch()— createsCONTAINS_MODULE,CALLS,EXTENDSedges
The concrete ingestor translates these into Cypher statements for Neo4j or any backend you implement.
7. Post-Processing and Cross-File Resolution
After DefinitionProcessor.walk_repo() completes, emit_type_edges() (line 67) finalizes the graph:
- Resolves deferred imports
- Creates type edges for inferred relationships
- Links cross-file references (e.g., a function call to a method defined in another module)
Complete Code Example: Ingesting a Repository
The ingest_repo function in codebase_rag/workspaces/cli.py wires all components together for end-to-end usage:
from codebase_rag.workspaces.cli import ingest_repo
from pathlib import Path
repo_path = Path("/path/to/your/project")
# Creates DefinitionProcessor, loads parsers lazily, walks repository
graph = ingest_repo(repo_path, project_name="my_project")
print("Graph contains", graph.node_count(), "nodes")
print("Relationships:", graph.relationship_count())
For finer control, instantiate components directly:
from codebase_rag.parser_loader import load_parsers
from codebase_rag.parsers.definition_processor import DefinitionProcessor
from codebase_rag.ingestors.neo4j_ingestor import Neo4jIngestor
parsers, queries = load_parsers()
ingestor = Neo4jIngestor(uri="bolt://localhost:7687", user="neo4j", password="password")
processor = DefinitionProcessor(
parsers=parsers,
queries=queries,
ingestor=ingestor,
project_name="my_analysis"
)
# Walk and ingest
for file_path in Path("src").rglob("*.py"):
processor.process_file(file_path)
processor.finalize() # Emit type edges and resolve deferred items
Key Files and Their Responsibilities
| File | Primary Role |
|---|---|
codebase_rag/parsers/definition_processor.py |
Core orchestrator — walks files, runs Tree-sitter, extracts facts, manages deferred buffers |
codebase_rag/parser_loader.py |
Lazy grammar loading, query construction, _LazyGrammarStore implementation |
codebase_rag/parsers/function_ingest.py |
FunctionIngestMixin — function and method extraction |
codebase_rag/parsers/class_ingest.py |
ClassIngestMixin — class and inheritance extraction |
codebase_rag/constants.py |
Language enums, node labels (MODULE, FUNCTION, CLASS), relationship types |
codebase_rag/utils/path_utils.py |
Path-to-QN conversion and module disambiguation |
codebase_rag/workspaces/cli.py |
Public API entry point (ingest_repo) |
Summary
- Lazy loading via
load_parsers()minimizes memory overhead by importing Tree-sitter grammars on first use, not at startup. - Qualified names (QNs) computed from file paths provide stable identifiers for cross-file reference resolution.
- Mix-in architecture separates extraction logic by entity type (functions, classes, JS/TS specifics) for maintainability.
- Deferred processing buffers ambiguous entities until the full repository is parsed, enabling accurate forward reference linking.
IngestorProtocolabstraction decouples fact extraction from storage backends, supporting Neo4j or custom graph databases.
Frequently Asked Questions
What Tree-sitter grammars does code-graph-rag support?
The repository supports 14 languages through compiled Tree-sitter grammars. The exact set is defined in codebase_rag/constants.py and loaded lazily via TreeSitterModule specifications in parser_loader.py. You can extend support by adding new grammar packages and registering them in the LanguageQueries construction.
How does the system handle parsing errors in malformed source files?
The parse_with_preproc_recovery() function in definition_processor.py implements recovery logic, particularly for C/C++ where preprocessor directives may create non-contiguous syntax trees. For unrecoverable errors, the file is skipped and logged, allowing the pipeline to continue with remaining sources.
Can I use a different graph database instead of Neo4j?
Yes. The IngestorProtocol abstract base class defines the interface (ensure_node_batch, ensure_relationship_batch, etc.). Implement this protocol for your target database — Apache TinkerPop, Amazon Neptune, or a custom in-memory graph — and pass your ingestor to DefinitionProcessor.
Why are some entities deferred instead of ingested immediately?
Deferred ingestion handles cases where entity identity depends on context not yet available. Examples include C++ symbols generated by macros (where the macro definition may appear later) and anonymous JavaScript functions that need call-site analysis for naming. The finalize() method resolves these after all files are processed.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →