# How to Parse a Code Repository into a Knowledge Graph Using Tree-sitter

> Learn to parse code repositories into knowledge graphs with Tree-sitter. This guide explores extracting typed entities and ingesting them using a custom DefinitionProcessor and modular mix-ins.

- Repository: [Vitali Avagyan/code-graph-rag](https://github.com/vitali87/code-graph-rag)
- Tags: how-to-guide
- Published: 2026-09-06

---

**The *code-graph-rag* repository uses a lazy-loading grammar store and a custom `DefinitionProcessor` to parse source code with Tree-sitter, extract typed entities, and ingest them into a knowledge graph via modular mix-ins.**

Parsing a codebase into a navigable knowledge graph requires more than running a parser over every file. The [vitali87/code-graph-rag](https://github.com/vitali87/code-graph-rag) project demonstrates a production-ready architecture that combines **lazy grammar loading**, **qualified name resolution**, and **deferred fact extraction** to build cross-language code graphs at scale. This article walks through the complete pipeline implemented in the source, from parser initialization to final relationship linking.

## Architecture Overview: From Source Code to Knowledge Graph

The ingestion pipeline follows seven distinct stages, each handled by specialized components in the `codebase_rag` package. Understanding this flow is essential when adapting the approach to your own tools.

### 1. Lazy Grammar Loading with `load_parsers()`

The system avoids the memory cost of loading all 14 supported Tree-sitter grammars upfront. Instead, `load_parsers()` in [`codebase_rag/parser_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/parser_loader.py) implements a **`_LazyGrammarStore`** that caches `Parser` and `LanguageQueries` objects per language.

When `parsers[lang]` or `queries[lang]` is first accessed, `_process_language()` builds a `tree_sitter.Language` object from the compiled grammar and registers the parser. This pattern ensures that parsing Python triggers only the Python grammar, not JavaScript, Go, or C++.

```python
from codebase_rag.parser_loader import load_parsers
from tree_sitter import Parser

# Triggers lazy load for Python only

parsers, queries = load_parsers()
python_parser: Parser = parsers["python"]  # First access triggers _process_language()

```

The `COMBINED_FUNC_CLASS_IMPORT_QUERIES` constant (defined at line 17) pre-builds the Tree-sitter query strings for each language, covering `functions`, `classes`, `calls`, `imports`, and more.

### 2. File Walking and Qualified Name Generation

The **`DefinitionProcessor`** in [`codebase_rag/parsers/definition_processor.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/parsers/definition_processor.py) orchestrates the repository walk. For each file, `process_file()` computes a **qualified name (QN)** for the module using:

- `base_module_qn()` — transforms [`src/util/helpers.py`](https://github.com/vitali87/code-graph-rag/blob/main/src/util/helpers.py) into `my_project.src.util.helpers`
- `_disambiguate_module_qn()` — handles name collisions when multiple files share the same relative path

These QNs become stable identifiers for graph nodes, enabling cross-file reference resolution later.

```python
from codebase_rag.utils.path_utils import base_module_qn
from pathlib import Path

module_qn = base_module_qn(Path("src/util/helpers.py"), "my_project")
print(module_qn)  # → my_project.src.util.helpers

```

### 3. Tree-sitter Parsing with Preprocessor Recovery

For each file, the processor retrieves the appropriate parser and calls `parse_with_preproc_recovery()` (line 21). This function wraps `parser.parse(source_bytes)` with special handling for C/C++ files where preprocessor directives may fragment the syntax tree.

```python
tree = parse_with_preproc_recovery(parser, source_bytes, language)

```

The resulting `Tree` object contains the complete syntax tree for query execution.

### 4. Executing Tree-sitter Queries

The processor runs language-specific queries using `tree_sitter.QueryCursor`. Captured nodes are sorted and cached in `_func_class_captures_cache` to avoid re-querying during mix-in processing.

```python
func_query = queries["python"].functions
cursor = tree.walk()

# Query execution returns named captures for function definitions, calls, etc.

```

### 5. Fact Extraction via Mix-ins

`DefinitionProcessor` inherits behavior from three specialized mix-ins:

- **`FunctionIngestMixin`** — extracts functions, methods, and their signatures
- **`ClassIngestMixin`** — extracts classes, inheritance relationships, and members
- **`JsTsIngestMixin`** — handles JavaScript/TypeScript-specific constructs like arrow functions and interfaces

Each mix-in creates graph entities with stable QNs. **Deferred items** — such as C++ macro-generated nodes or anonymous JavaScript functions — are buffered internally and flushed after the full file set is processed. This two-phase approach ensures that forward references can be resolved once all files are parsed.

### 6. Graph Ingestion via `IngestorProtocol`

Extracted facts are handed to an `IngestorProtocol` implementation through two batch methods:

- `ingestor.ensure_node_batch()` — creates `MODULE`, `FUNCTION`, `CLASS` nodes
- `ingestor.ensure_relationship_batch()` — creates `CONTAINS_MODULE`, `CALLS`, `EXTENDS` edges

The concrete ingestor translates these into Cypher statements for Neo4j or any backend you implement.

### 7. Post-Processing and Cross-File Resolution

After `DefinitionProcessor.walk_repo()` completes, `emit_type_edges()` (line 67) finalizes the graph:

- Resolves deferred imports
- Creates type edges for inferred relationships
- Links cross-file references (e.g., a function call to a method defined in another module)

## Complete Code Example: Ingesting a Repository

The `ingest_repo` function in [`codebase_rag/workspaces/cli.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/workspaces/cli.py) wires all components together for end-to-end usage:

```python
from codebase_rag.workspaces.cli import ingest_repo
from pathlib import Path

repo_path = Path("/path/to/your/project")

# Creates DefinitionProcessor, loads parsers lazily, walks repository

graph = ingest_repo(repo_path, project_name="my_project")

print("Graph contains", graph.node_count(), "nodes")
print("Relationships:", graph.relationship_count())

```

For finer control, instantiate components directly:

```python
from codebase_rag.parser_loader import load_parsers
from codebase_rag.parsers.definition_processor import DefinitionProcessor
from codebase_rag.ingestors.neo4j_ingestor import Neo4jIngestor

parsers, queries = load_parsers()
ingestor = Neo4jIngestor(uri="bolt://localhost:7687", user="neo4j", password="password")

processor = DefinitionProcessor(
    parsers=parsers,
    queries=queries,
    ingestor=ingestor,
    project_name="my_analysis"
)

# Walk and ingest

for file_path in Path("src").rglob("*.py"):
    processor.process_file(file_path)

processor.finalize()  # Emit type edges and resolve deferred items

```

## Key Files and Their Responsibilities

| File | Primary Role |
|------|-------------|
| [`codebase_rag/parsers/definition_processor.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/parsers/definition_processor.py) | Core orchestrator — walks files, runs Tree-sitter, extracts facts, manages deferred buffers |
| [`codebase_rag/parser_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/parser_loader.py) | Lazy grammar loading, query construction, `_LazyGrammarStore` implementation |
| [`codebase_rag/parsers/function_ingest.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/parsers/function_ingest.py) | `FunctionIngestMixin` — function and method extraction |
| [`codebase_rag/parsers/class_ingest.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/parsers/class_ingest.py) | `ClassIngestMixin` — class and inheritance extraction |
| [`codebase_rag/constants.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/constants.py) | Language enums, node labels (`MODULE`, `FUNCTION`, `CLASS`), relationship types |
| [`codebase_rag/utils/path_utils.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/utils/path_utils.py) | Path-to-QN conversion and module disambiguation |
| [`codebase_rag/workspaces/cli.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/workspaces/cli.py) | Public API entry point (`ingest_repo`) |

## Summary

- **Lazy loading** via `load_parsers()` minimizes memory overhead by importing Tree-sitter grammars on first use, not at startup.
- **Qualified names (QNs)** computed from file paths provide stable identifiers for cross-file reference resolution.
- **Mix-in architecture** separates extraction logic by entity type (functions, classes, JS/TS specifics) for maintainability.
- **Deferred processing** buffers ambiguous entities until the full repository is parsed, enabling accurate forward reference linking.
- **`IngestorProtocol`** abstraction decouples fact extraction from storage backends, supporting Neo4j or custom graph databases.

## Frequently Asked Questions

### What Tree-sitter grammars does code-graph-rag support?

The repository supports 14 languages through compiled Tree-sitter grammars. The exact set is defined in [`codebase_rag/constants.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/constants.py) and loaded lazily via `TreeSitterModule` specifications in [`parser_loader.py`](https://github.com/vitali87/code-graph-rag/blob/main/parser_loader.py). You can extend support by adding new grammar packages and registering them in the `LanguageQueries` construction.

### How does the system handle parsing errors in malformed source files?

The `parse_with_preproc_recovery()` function in [`definition_processor.py`](https://github.com/vitali87/code-graph-rag/blob/main/definition_processor.py) implements recovery logic, particularly for C/C++ where preprocessor directives may create non-contiguous syntax trees. For unrecoverable errors, the file is skipped and logged, allowing the pipeline to continue with remaining sources.

### Can I use a different graph database instead of Neo4j?

Yes. The `IngestorProtocol` abstract base class defines the interface (`ensure_node_batch`, `ensure_relationship_batch`, etc.). Implement this protocol for your target database — Apache TinkerPop, Amazon Neptune, or a custom in-memory graph — and pass your ingestor to `DefinitionProcessor`.

### Why are some entities deferred instead of ingested immediately?

Deferred ingestion handles cases where entity identity depends on context not yet available. Examples include C++ symbols generated by macros (where the macro definition may appear later) and anonymous JavaScript functions that need call-site analysis for naming. The `finalize()` method resolves these after all files are processed.