How code-review-graph Parses Source Code to Build Its Knowledge Graph: A Deep Dive into the Parser Architecture

code-review-graph parses source code by detecting the programming language, delegating to language-specific AST resolvers to extract symbols and relationships, then running cross-file post-processors to construct a unified NetworkX knowledge graph of nodes and edges.

The code-review-graph repository implements a sophisticated multi-language parser that transforms raw source files into a queryable knowledge graph. At its core, the CodeParser class in code_review_graph/parser.py orchestrates a pipeline of language detection, AST resolution, and graph-wide post-processing to produce a structured representation of code semantics.

The Core Entry Point: parse_bytes

Every parsing operation flows through the parse_bytes method of the CodeParser class.

def parse_bytes(self, path: Path, source: bytes) -> tuple[list[NodeInfo], list[EdgeInfo]]:

Source: parser.py L2598

This method executes a four-stage pipeline:

  1. Detect the language using file extension, shebang, or content heuristics
  2. Select a resolver that understands the concrete syntax of that language
  3. Collect nodes and edges from the language-specific AST walk
  4. Run graph-wide post-processors that add missing cross-file links

The result is a pair of lists: NodeInfo objects representing symbols (functions, classes, variables) and EdgeInfo objects representing relationships (calls, imports, inheritance).

Stage 1: Language Detection

Before any AST parsing occurs, code-review-graph must determine which language a file contains.

def detect_language(self, path: Path, source: Optional[bytes] = None) -> Optional[str]:

Source: parser.py L2483-L2520

The detection strategy checks multiple signals in priority order:

  • File extension — .py for Python, .java for Java, .js / .ts for JavaScript/TypeScript, .rs for Rust
  • Shebang line — scripts without extensions are analyzed via _detect_language_from_shebang
  • Special-case heuristics — HCL files, Spring configuration files, and other framework-specific patterns

Files that cannot be classified are skipped during graph construction, preventing parse errors from contaminating the knowledge base.

Stage 2: Resolver Dispatch

Once a language is identified, CodeParser instantiates or reuses a resolver that encapsulates language-specific parsing logic:

if language == "python":
    resolver = self._python_resolver
elif language == "java":
    resolver = self._java_resolver

# … additional cases: cpp, ts, rust, php, spring, etc.

Source: parser.py L2410-L2440 (excerpt)

Each resolver lives in its own module and implements a standard resolve interface:

Resolver Module Language Parsing Technology
python_resolver.py Python ast module + jedi
jedi_resolver.py Python (enhanced) jedi for type inference
spring_resolver.py Java Spring Custom AST + framework conventions
rust_resolver.py Rust tree-sitter or custom parser
tsconfig_resolver.py TypeScript tsconfig.json + path mapping

Resolvers emit NodeInfo and EdgeInfo objects during their AST traversal, capturing both structural and semantic relationships.

Stage 3: Python Resolver Deep Dive

The Python resolver (python_resolver.py) demonstrates the typical resolver implementation. It walks Python's AST to create nodes for:

  • Modules, Classes, Functions, Variables
  • Decorators, Async functions, Lambdas

And edges representing:

  • Function calls
  • Class inheritance
  • Attribute accesses
  • Import relationships

Source: [python_resolver.py relevant sections] — see the resolver's resolve method for concrete node/edge creation

The resolver uses Python's built-in ast module for syntax analysis, optionally enhanced by jedi for richer type information and cross-module symbol resolution.

Stage 4: Cross-File Linking with Post-Processors

After language-specific parsing, CodeParser runs a series of graph-wide resolvers that operate on the complete node/edge collection to resolve symbols across file boundaries:

Post-Processor Function Source Location
Scoped resolver Validates scope-visible symbols and adds edges for lexical containment (functions within classes, classes within modules) parser.py L5123-L5459
TSConfig resolver Resolves TypeScript path aliases defined in tsconfig.json to physical file paths parser.py L13589-L13682
Spring resolver Handles Spring DI bean wiring, constructor injection, and @Autowired method resolution for Java projects parser.py L10720-L10826
Temporal resolver Links temporally-scoped symbols such as asyncio tasks or Spring-scheduled jobs parser.py L9791-L9805

These resolvers examine the global graph state, resolve import statements against discovered symbols, apply framework-specific conventions, and emit additional edges that complete the knowledge graph.

Building the Final Knowledge Graph

The output of parse_bytes — and its convenience wrappers parse_file and parse_repo — feeds into code_review_graph/graph.py. This module constructs a NetworkX-compatible directed graph that powers:

  • Semantic search via semantic_search_nodes_tool
  • Graph queries via query_graph_tool
  • Impact assessment for code changes
  • Skill extraction and developer expertise mapping

Practical Usage Examples

Parse a Single Python File

from pathlib import Path
from code_review_graph.parser import CodeParser

parser = CodeParser()
nodes, edges = parser.parse_file(Path("my_project/main.py"))

print("Nodes:", len(nodes))
print("Edges:", len(edges))

# Nodes contain identifiers like "/my_project/main.py::MyClass.method"

# Edges contain calls such as ("...::MyClass.method", "...::helper")

Parse an Entire Repository

from code_review_graph.parser import CodeParser
from code_review_graph.graph import GraphBuilder

repo_root = Path("/path/to/repo")
parser = CodeParser(repo_root=repo_root)

graph_builder = GraphBuilder()
graph_builder.ingest_repo(parser)   # walks the repo, parses each file

graph = graph_builder.graph        # a NetworkX DiGraph

# Example query: find all callers of a function

callers = [e.source for e in graph.edges(data=True) 
           if e.target == "/repo/src/util.py::process"]

Direct Resolver Access (Advanced)

from code_review_graph.python_resolver import PythonResolver
from pathlib import Path

resolver = PythonResolver()
source = Path("example.py").read_bytes()
nodes, edges = resolver.resolve(Path("example.py"), source)

Key Source Files

File Role
code_review_graph/parser.py Central orchestrator: language detection, resolver dispatch, post-processing
code_review_graph/python_resolver.py AST walk for Python; creates nodes/edges for functions, classes, imports
code_review_graph/jedi_resolver.py Enhances Python parsing with Jedi for richer type information
code_review_graph/spring_resolver.py Resolves Spring-specific DI and bean wiring in Java
code_review_graph/tsconfig_resolver.py Handles TypeScript path-alias resolution
code_review_graph/graph.py Constructs NetworkX graph from parsed structures
code_review_graph/graph_diff.py Computes incremental diffs between graph snapshots
code_review_graph/visualization.py Generates HTML/SVG graph visualizations

Summary

  • parse_bytes is the central entry point that orchestrates the complete parsing pipeline in code_review_graph/parser.py
  • Language detection combines file extension, shebang analysis, and heuristic checks to select the appropriate resolver
  • Language-specific resolvers extract NodeInfo and EdgeInfo from AST structures using tools like ast, jedi, and tree-sitter
  • Post-processors resolve cross-file references, framework conventions, and path aliases to complete the knowledge graph
  • NetworkX integration enables powerful graph queries, visualizations, and downstream code analysis workflows

Frequently Asked Questions

What languages does code-review-graph support?

code-review-graph supports Python, Java (including Spring Framework), JavaScript, TypeScript, Rust, C++, PHP, and HCL. Language support is modular — each language has its own resolver module in the code_review_graph/ directory. New languages can be added by implementing the resolver interface.

How does the Spring resolver handle dependency injection?

The Spring resolver in spring_resolver.py and post-processing logic in parser.py (lines 10720-10826) analyzes @Component, @Service, @Repository, and @Autowired annotations. It creates edges representing bean wiring relationships and constructor/method injection points, linking interface definitions to concrete implementations in the knowledge graph.

Can I use code-review-graph without the full graph builder?

Yes. The CodeParser class can be used standalone to produce NodeInfo and EdgeInfo lists for individual files or directories. You can also instantiate language-specific resolvers directly (like PythonResolver) for custom parsing pipelines that don't require cross-file resolution or the full NetworkX graph structure.

How does incremental parsing work?

Incremental updates are handled by code_review_graph/graph_diff.py, which computes differences between two graph snapshots. When files change, only affected nodes and edges are re-parsed and merged, rather than rebuilding the entire knowledge graph from scratch.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →