How code-review-graph Parses Source Code to Build Its Knowledge Graph: A Deep Dive into the Parser Architecture
code-review-graph parses source code by detecting the programming language, delegating to language-specific AST resolvers to extract symbols and relationships, then running cross-file post-processors to construct a unified NetworkX knowledge graph of nodes and edges.
The code-review-graph repository implements a sophisticated multi-language parser that transforms raw source files into a queryable knowledge graph. At its core, the CodeParser class in code_review_graph/parser.py orchestrates a pipeline of language detection, AST resolution, and graph-wide post-processing to produce a structured representation of code semantics.
The Core Entry Point: parse_bytes
Every parsing operation flows through the parse_bytes method of the CodeParser class.
def parse_bytes(self, path: Path, source: bytes) -> tuple[list[NodeInfo], list[EdgeInfo]]:
Source: parser.py L2598
This method executes a four-stage pipeline:
- Detect the language using file extension, shebang, or content heuristics
- Select a resolver that understands the concrete syntax of that language
- Collect nodes and edges from the language-specific AST walk
- Run graph-wide post-processors that add missing cross-file links
The result is a pair of lists: NodeInfo objects representing symbols (functions, classes, variables) and EdgeInfo objects representing relationships (calls, imports, inheritance).
Stage 1: Language Detection
Before any AST parsing occurs, code-review-graph must determine which language a file contains.
def detect_language(self, path: Path, source: Optional[bytes] = None) -> Optional[str]:
Source: parser.py L2483-L2520
The detection strategy checks multiple signals in priority order:
- File extension —
.pyfor Python,.javafor Java,.js/.tsfor JavaScript/TypeScript,.rsfor Rust - Shebang line — scripts without extensions are analyzed via
_detect_language_from_shebang - Special-case heuristics — HCL files, Spring configuration files, and other framework-specific patterns
Files that cannot be classified are skipped during graph construction, preventing parse errors from contaminating the knowledge base.
Stage 2: Resolver Dispatch
Once a language is identified, CodeParser instantiates or reuses a resolver that encapsulates language-specific parsing logic:
if language == "python":
resolver = self._python_resolver
elif language == "java":
resolver = self._java_resolver
# … additional cases: cpp, ts, rust, php, spring, etc.
Source: parser.py L2410-L2440 (excerpt)
Each resolver lives in its own module and implements a standard resolve interface:
| Resolver Module | Language | Parsing Technology |
|---|---|---|
python_resolver.py |
Python | ast module + jedi |
jedi_resolver.py |
Python (enhanced) | jedi for type inference |
spring_resolver.py |
Java Spring | Custom AST + framework conventions |
rust_resolver.py |
Rust | tree-sitter or custom parser |
tsconfig_resolver.py |
TypeScript | tsconfig.json + path mapping |
Resolvers emit NodeInfo and EdgeInfo objects during their AST traversal, capturing both structural and semantic relationships.
Stage 3: Python Resolver Deep Dive
The Python resolver (python_resolver.py) demonstrates the typical resolver implementation. It walks Python's AST to create nodes for:
- Modules, Classes, Functions, Variables
- Decorators, Async functions, Lambdas
And edges representing:
- Function calls
- Class inheritance
- Attribute accesses
- Import relationships
Source: [python_resolver.py relevant sections] — see the resolver's resolve method for concrete node/edge creation
The resolver uses Python's built-in ast module for syntax analysis, optionally enhanced by jedi for richer type information and cross-module symbol resolution.
Stage 4: Cross-File Linking with Post-Processors
After language-specific parsing, CodeParser runs a series of graph-wide resolvers that operate on the complete node/edge collection to resolve symbols across file boundaries:
| Post-Processor | Function | Source Location |
|---|---|---|
| Scoped resolver | Validates scope-visible symbols and adds edges for lexical containment (functions within classes, classes within modules) | parser.py L5123-L5459 |
| TSConfig resolver | Resolves TypeScript path aliases defined in tsconfig.json to physical file paths |
parser.py L13589-L13682 |
| Spring resolver | Handles Spring DI bean wiring, constructor injection, and @Autowired method resolution for Java projects |
parser.py L10720-L10826 |
| Temporal resolver | Links temporally-scoped symbols such as asyncio tasks or Spring-scheduled jobs |
parser.py L9791-L9805 |
These resolvers examine the global graph state, resolve import statements against discovered symbols, apply framework-specific conventions, and emit additional edges that complete the knowledge graph.
Building the Final Knowledge Graph
The output of parse_bytes — and its convenience wrappers parse_file and parse_repo — feeds into code_review_graph/graph.py. This module constructs a NetworkX-compatible directed graph that powers:
- Semantic search via
semantic_search_nodes_tool - Graph queries via
query_graph_tool - Impact assessment for code changes
- Skill extraction and developer expertise mapping
Practical Usage Examples
Parse a Single Python File
from pathlib import Path
from code_review_graph.parser import CodeParser
parser = CodeParser()
nodes, edges = parser.parse_file(Path("my_project/main.py"))
print("Nodes:", len(nodes))
print("Edges:", len(edges))
# Nodes contain identifiers like "/my_project/main.py::MyClass.method"
# Edges contain calls such as ("...::MyClass.method", "...::helper")
Parse an Entire Repository
from code_review_graph.parser import CodeParser
from code_review_graph.graph import GraphBuilder
repo_root = Path("/path/to/repo")
parser = CodeParser(repo_root=repo_root)
graph_builder = GraphBuilder()
graph_builder.ingest_repo(parser) # walks the repo, parses each file
graph = graph_builder.graph # a NetworkX DiGraph
# Example query: find all callers of a function
callers = [e.source for e in graph.edges(data=True)
if e.target == "/repo/src/util.py::process"]
Direct Resolver Access (Advanced)
from code_review_graph.python_resolver import PythonResolver
from pathlib import Path
resolver = PythonResolver()
source = Path("example.py").read_bytes()
nodes, edges = resolver.resolve(Path("example.py"), source)
Key Source Files
| File | Role |
|---|---|
code_review_graph/parser.py |
Central orchestrator: language detection, resolver dispatch, post-processing |
code_review_graph/python_resolver.py |
AST walk for Python; creates nodes/edges for functions, classes, imports |
code_review_graph/jedi_resolver.py |
Enhances Python parsing with Jedi for richer type information |
code_review_graph/spring_resolver.py |
Resolves Spring-specific DI and bean wiring in Java |
code_review_graph/tsconfig_resolver.py |
Handles TypeScript path-alias resolution |
code_review_graph/graph.py |
Constructs NetworkX graph from parsed structures |
code_review_graph/graph_diff.py |
Computes incremental diffs between graph snapshots |
code_review_graph/visualization.py |
Generates HTML/SVG graph visualizations |
Summary
parse_bytesis the central entry point that orchestrates the complete parsing pipeline incode_review_graph/parser.py- Language detection combines file extension, shebang analysis, and heuristic checks to select the appropriate resolver
- Language-specific resolvers extract
NodeInfoandEdgeInfofrom AST structures using tools likeast,jedi, andtree-sitter - Post-processors resolve cross-file references, framework conventions, and path aliases to complete the knowledge graph
- NetworkX integration enables powerful graph queries, visualizations, and downstream code analysis workflows
Frequently Asked Questions
What languages does code-review-graph support?
code-review-graph supports Python, Java (including Spring Framework), JavaScript, TypeScript, Rust, C++, PHP, and HCL. Language support is modular — each language has its own resolver module in the code_review_graph/ directory. New languages can be added by implementing the resolver interface.
How does the Spring resolver handle dependency injection?
The Spring resolver in spring_resolver.py and post-processing logic in parser.py (lines 10720-10826) analyzes @Component, @Service, @Repository, and @Autowired annotations. It creates edges representing bean wiring relationships and constructor/method injection points, linking interface definitions to concrete implementations in the knowledge graph.
Can I use code-review-graph without the full graph builder?
Yes. The CodeParser class can be used standalone to produce NodeInfo and EdgeInfo lists for individual files or directories. You can also instantiate language-specific resolvers directly (like PythonResolver) for custom parsing pipelines that don't require cross-file resolution or the full NetworkX graph structure.
How does incremental parsing work?
Incremental updates are handled by code_review_graph/graph_diff.py, which computes differences between two graph snapshots. When files change, only affected nodes and edges are re-parsed and merged, rather than rebuilding the entire knowledge graph from scratch.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →