How Diff Impact Analysis in Egonex-AI Identifies Affected Parts of the Codebase
Egonex-AI's diff impact analysis identifies affected parts of the codebase by traversing a static knowledge graph where nodes represent code entities and edges represent relationships, expanding from changed files to include child components and immediate neighbors to calculate the full scope of architectural impact.
The Understand-Anything platform constructs a knowledge graph of your project to power its analysis engine. Within this graph, nodes correspond to code entities, edges define relationships between them, and layers group nodes into architectural tiers. When a changeset arrives, the diff impact analysis walks this structure to pinpoint every component that may be affected by the modifications, all without requiring runtime execution.
The Knowledge Graph Architecture
Before analyzing diffs, Understand-Anything builds a comprehensive graph representation of the codebase. This graph encodes code entities as nodes, dependencies and containment as edges, and architectural layers as logical groupings of nodes. This static structure enables the system to understand semantic relationships that traditional text-based diffing cannot capture, such as which services depend on a modified utility function or which UI components reside within a changed directory.
Step-by-Step Diff Analysis in diff-analyzer.ts
The core logic resides in understand-anything-plugin/src/diff-analyzer.ts, specifically within the buildDiffContext function. The analyzer processes a list of changed file paths through a seven-stage pipeline to determine impact.
Mapping Changed Files to Graph Nodes
First, the analyzer maps each file path in the changeset to corresponding nodes in the knowledge graph (lines 31-42). It compares the input path against the node.filePath property of each graph node. Any file paths that fail to match existing nodes are recorded as unmapped files (line 40), flagging potential new additions or orphaned resources.
Traversing Contains Relationships
Next, the analyzer expands the set of changed nodes by following containment edges. For every edge of type "contains" where the source node appears in the changed set, the analyzer pulls in the target node (lines 44-49). This step captures files that are physically or logically inside a changed directory, ensuring that parent directory modifications correctly flag all child components.
Identifying 1-Hop Neighbors
The analyzer then scans the entire graph to find immediate dependencies. It iterates through every edge (lines 57-70), checking if either endpoint belongs to changedNodeIds. When a match occurs, the edge is added to impactedEdges, and the opposite endpoint—if not already marked as changed—becomes an affected node. This step identifies downstream consumers and upstream dependencies that interact directly with modified code.
Aggregating Impacted Layers
Finally, the analyzer determines which architectural layers are compromised. Any layer whose nodeIds intersect with the union of changed nodes and affected nodes gets flagged as impacted (lines 74-77). This aggregation helps developers understand whether changes remain isolated within a single architectural tier or cross critical boundaries.
Building the DiffContext Output
After completing the graph traversal, the analyzer constructs a comprehensive DiffContext object (lines 79-87). This structure contains:
changedFiles– the original input pathschangedNodes– nodes directly modified by the diffaffectedNodes– nodes identified through 1-hop neighbor traversalimpactedEdges– relationships connecting changed and affected componentsaffectedLayers– architectural layers containing touched nodesunmappedFiles– paths not currently represented in the graph
Generating Human-Readable Reports
The formatDiffAnalysis function (lines 90-198) transforms the raw DiffContext into a markdown report. This output lists changed components with their complexity metrics, enumerates downstream affected components, identifies compromised architectural layers, and provides a risk assessment based on the scope of the impact. The report enables code review bots and explanation agents to focus their attention precisely where the changes matter most.
Practical Implementation Example
To execute diff impact analysis in your own workflow, import the analyzer functions and provide a knowledge graph instance along with your changed file paths:
import { buildDiffContext, formatDiffAnalysis } from "./diff-analyzer";
import type { KnowledgeGraph } from "@understand-anything/core";
/* 1️⃣ Load the knowledge graph (e.g. from a JSON export) */
const graph: KnowledgeGraph = await fetch("/knowledge-graph.json").then(r => r.json());
/* 2️⃣ List of files changed in the latest commit */
const changedFiles = [
"src/utils/helpers.ts",
"src/components/Button.tsx",
];
/* 3️⃣ Build the diff context */
const ctx = buildDiffContext(graph, changedFiles);
/* 4️⃣ Produce a markdown report */
const report = formatDiffAnalysis(ctx);
console.log(report);
Result (abridged)
# Diff Analysis: MyProject
## Changed Components
- **helpers** (function) — Utility helpers
- File: `src/utils/helpers.ts`
- Complexity: simple
- **Button** (component) — UI button component
...
## Affected Components
- **Icon** (component) — Shared icon library
- **ThemeProvider** (service) — Provides theming
...
## Affected Layers
- **UI Layer**: Presentation components
- **Core Logic Layer**: Business rules
Integration with the Analysis Pipeline
The context-builder.ts file integrates this diff analysis into the broader LLM-driven workflow. By feeding the DiffContext into downstream agents, the system ensures that explanation generators, onboarding tools, and review bots receive precisely the subset of the codebase relevant to the current changeset, rather than processing the entire repository.
Summary
- Egonex-AI uses a static knowledge graph to determine code impact without runtime execution.
- The
buildDiffContextfunction indiff-analyzer.tsimplements a seven-stage traversal algorithm. - Analysis includes direct file mapping, containment expansion, 1-hop neighbor detection, and layer aggregation.
- The resulting
DiffContextcaptures changed nodes, affected dependencies, impacted edges, and unmapped files. formatDiffAnalysisrenders the technical results into developer-friendly markdown reports.
Frequently Asked Questions
What is the primary source file for diff impact analysis in Egonex-AI?
The core implementation lives in understand-anything-plugin/src/diff-analyzer.ts within the Understand-Anything repository. This file defines both the buildDiffContext function for graph traversal and the formatDiffAnalysis function for report generation.
How does the analyzer handle files not yet represented in the knowledge graph?
Unmatched file paths are recorded in the unmappedFiles array (line 40 of diff-analyzer.ts). These entries highlight new files or resources that have not been indexed into the graph, allowing developers to update the knowledge base or investigate orphaned code.
What relationship types does the diff analyzer traverse to find affected components?
The analyzer specifically looks for edges of type "contains" to identify parent-child relationships (lines 44-49) and general edges to find 1-hop neighbors (lines 57-70). It does not use heuristics or runtime analysis—only the topology of the static knowledge graph.
How does this differ from standard git diff output?
While git diff shows textual changes between file versions, diff impact analysis understands semantic relationships. It identifies which architectural layers, dependent services, and sibling components are affected by a change, providing architectural context that line-by-line comparisons cannot capture.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →