How code-review-graph Builds a Knowledge Graph from Source Code: Parser Architecture Explained

The code-review-graph framework transforms raw source files into a queryable knowledge graph by detecting programming languages, delegating to specialized AST resolvers, and executing post-processing linkers to resolve cross-file symbol relationships.

The code-review-graph repository by tirth8205 provides a robust pipeline for constructing knowledge graphs from heterogeneous codebases. By parsing source code into nodes (symbols) and edges (relationships), the system enables semantic search, impact analysis, and dependency visualization. At the heart of this pipeline sits the CodeParser class in code_review_graph/parser.py, which orchestrates the entire extraction process from raw bytes to graph structures.

Entry Point: The parse_bytes Method

The primary interface for file ingestion is parse_bytes at line 2598 of parser.py:

def parse_bytes(self, path: Path, source: bytes) -> tuple[list[NodeInfo], list[EdgeInfo]]:

This method executes a four-stage pipeline:

  1. Language Detection: Identifies the programming language via file extension, shebang analysis, or content heuristics.
  2. Resolver Selection: Instantiates a language-specific resolver capable of walking that language's AST.
  3. Symbol Extraction: Collects NodeInfo and EdgeInfo objects representing entities and their relationships.
  4. Post-Processing: Runs graph-wide resolvers to link symbols across files and frameworks.

The method returns two lists containing fully-qualified nodes and their connecting edges, which subsequently feed into the graph builder.

Language Detection and Resolver Dispatch

Before AST traversal begins, detect_language (lines 2483-2520) determines how to parse the file:

def detect_language(self, path: Path, source: Optional[bytes] = None) -> Optional[str]:

The detection logic checks file extensions (.py, .java, .ts), shebang lines for extensionless scripts, and special-case heuristics for frameworks like Spring or HCL configurations. If detection fails, the file is excluded from graph construction.

Following detection, CodeParser dispatches to a specialized resolver around lines 2410-2440:

if language == "python":
    resolver = self._python_resolver
elif language == "java":
    resolver = self._java_resolver

# … additional resolvers for cpp, ts, rust, php, spring, etc.

Each resolver resides in its own module—python_resolver.py, jedi_resolver.py, spring_resolver.py, rust_resolver.py—and implements a resolve method that traverses language-specific ASTs using tools like Python's built-in ast module, jedi, or tree-sitter.

AST Traversal and Symbol Extraction

Language-specific resolvers extract semantic entities by walking abstract syntax trees. The Python resolver (python_resolver.py) demonstrates this pattern by creating:

  • Nodes: Modules, classes, functions, variables, decorators, async functions, and lambdas.
  • Edges: Function calls, class inheritance chains, attribute accesses, and import relationships.

Resolvers emit NodeInfo objects containing fully-qualified identifiers, file locations, and metadata (docstrings, type hints), alongside EdgeInfo objects capturing relationship types—calls, imports, inheritance, and references.

Cross-File Linking with Post-Processing Resolvers

After language-specific parsing, CodeParser executes a series of graph-wide resolvers to connect symbols across file boundaries:

  • Scoped Resolver (lines 5123-5459): Validates visibility constraints for functions, classes, and modules, adding edges for symbols accessible only within specific scopes.
  • TSConfig Resolver (lines 13589-13682): Resolves TypeScript path aliases defined in tsconfig.json, mapping abstract import paths to concrete file locations.
  • Spring Resolver (lines 10720-10826): Handles Spring Framework dependency injection, wiring beans and method injections for Java projects.
  • Temporal Resolver (lines 9791-9805): Links temporally-scoped symbols such as scheduled tasks in Spring or Python's asyncio.

These resolvers examine the global graph state, resolve import mappings and framework conventions, then emit additional edges to complete the knowledge graph.

Constructing the NetworkX Graph

The final output of parse_bytes—along with convenience methods parse_file and parse_repo—consists of structured node and edge lists. These feed into code_review_graph/graph.py, which constructs a NetworkX-compatible directed graph supporting complex queries.

The graph enables semantic search tools, impact assessment, and visualization workflows by storing symbols as nodes and their relationships (calls, imports, inheritance) as typed edges.

Practical Code Examples

Parsing a Single Python File

from pathlib import Path
from code_review_graph.parser import CodeParser

parser = CodeParser()
nodes, edges = parser.parse_file(Path("my_project/main.py"))

print("Nodes:", len(nodes))
print("Edges:", len(edges))

# Nodes contain identifiers like "/my_project/main.py::MyClass.method"

# Edges contain calls such as ("...::MyClass.method", "...::helper")

Processing an Entire Repository

from code_review_graph.parser import CodeParser
from code_review_graph.graph import GraphBuilder

repo_root = Path("/path/to/repo")
parser = CodeParser(repo_root=repo_root)

graph_builder = GraphBuilder()
graph_builder.ingest_repo(parser)   # walks the repo, parses each file

graph = graph_builder.graph        # a NetworkX DiGraph

# Example query: find all callers of a function

callers = [e.source for e in graph.edges(data=True) if e.target == "/repo/src/util.py::process"]

Direct Resolver Access (Advanced)

from code_review_graph.python_resolver import PythonResolver
from pathlib import Path

resolver = PythonResolver()
source = Path("example.py").read_bytes()
nodes, edges = resolver.resolve(Path("example.py"), source)

Key Source Files

The knowledge graph pipeline spans several critical modules:

Summary

  • code-review-graph builds knowledge graphs through a multi-stage pipeline involving language detection, AST traversal, and cross-file linking.
  • The CodeParser class in parser.py serves as the central orchestrator, with parse_bytes as the primary entry point for converting source files to graph data.
  • Language-specific resolvers handle AST walking for Python, Java, TypeScript, Rust, and other languages, extracting NodeInfo and EdgeInfo objects.
  • Post-processing resolvers (Scoped, TSConfig, Spring, Temporal) resolve cross-file dependencies and framework-specific relationships after initial parsing.
  • The final output feeds into NetworkX via graph.py, creating a queryable directed graph for semantic search and dependency analysis.

Frequently Asked Questions

How does code-review-graph handle multiple programming languages in the same repository?

The framework uses the detect_language method in parser.py (lines 2483-2520) to identify each file's language via extension, shebang, or content analysis. It then dispatches to specialized resolvers—such as python_resolver.py for Python or spring_resolver.py for Java Spring—each implementing language-specific AST traversal logic while emitting standardized NodeInfo and EdgeInfo objects.

What types of relationships does the knowledge graph capture?

The graph captures semantic relationships including function calls, class inheritance, import dependencies, attribute accesses, and framework-specific wiring (such as Spring dependency injection). Post-processing resolvers add cross-file edges for scoped visibility, TypeScript path aliases, and temporal task scheduling.

Can I use code-review-graph to analyze a single file without processing the entire repository?

Yes. The CodeParser class provides the parse_file convenience method, which wraps parse_bytes to handle individual files. For advanced use cases, you can instantiate language-specific resolvers directly—such as PythonResolver—and call their resolve method with a file path and byte content.

Where does the graph data structure reside after parsing?

The raw node and edge lists returned by parse_bytes are consumed by code_review_graph/graph.py, which constructs a NetworkX DiGraph (directed graph). This graph object supports complex queries for callers, callees, inheritance chains, and semantic search through tools like semantic_search_nodes_tool and query_graph_tool.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →