# How Egonex-AI's Diff Impact Analysis Tracks Ripple Effects in a Codebase

> Discover how Egonex-AI's diff impact analysis tracks codebase ripple effects by mapping changes to a knowledge graph and analyzing dependencies. Understand the full impact without external services.

- Repository: [Egonex/Understand-Anything](https://github.com/Egonex-AI/Understand-Anything)
- Tags: deep-dive
- Published: 2026-06-09

---

**Egonex-AI's diff impact analysis maps changed file paths to a knowledge graph and propagates impact through "contains" relationships and direct dependencies to identify affected nodes, edges, and architectural layers without external service calls.**

The Understand-Anything platform transforms raw Git diffs into actionable intelligence by analyzing how changes propagate through your codebase's structural graph. By leveraging the knowledge graph produced by its core engine, this **diff impact analysis** reveals not just what files changed, but which downstream components, relationships, and architectural layers face potential **ripple effects**. The entire pipeline is self-contained in the `@understand-anything/plugin` package and operates purely on in-memory graph data.

## The Three-Step Algorithm in [`diff-analyzer.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/diff-analyzer.ts)

The heart of the ripple tracking logic resides in [`src/diff-analyzer.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/src/diff-analyzer.ts), specifically within the `buildDiffContext` function (lines 31–77). This algorithm processes the knowledge graph through three sequential phases to transform a simple list of changed paths into a comprehensive impact assessment.

### Step 1: Mapping Changed Files to Graph Nodes

The analyzer begins by correlating file paths from your diff (typically sourced via `git diff --name-only`) with nodes in the knowledge graph. During this phase, the function loops over every node in the graph and matches `node.filePath` against each changed file path.

Nodes that match become **changed nodes**, while paths without corresponding graph nodes are collected as **unmapped files** for incremental graph updates. This mapping occurs in `buildDiffContext` at lines 31–42.

### Step 2: Propagating Through "contains" Edges

Once initial nodes are identified, the algorithm expands the change set by traversing **"contains"** relationships. If a directory node contains files or subdirectories, all immediate children of a changed node are also marked as changed.

This containment expansion ensures that modifying a parent directory correctly flags all nested components as directly touched. The logic scans all edges for `type === "contains"` and adds the child node (`edge.target`) to the changed set, as implemented in lines 44–49.

### Step 3: Walking the Graph to Calculate Ripple Effects

The final phase identifies downstream impact by walking the graph topology:

- **Affected nodes**: One-hop neighbors of changed nodes (excluding already-changed nodes) are flagged as potentially impacted downstream components.
- **Impacted edges**: Every edge that touches either a changed node or an affected node is recorded, highlighting relationships that cross the change boundary.
- **Affected layers**: Any architectural layer whose `nodeIds` intersect with the union of changed and affected IDs is marked as impacted.

This neighbor discovery and layer aggregation logic appears in lines 53–77 of `buildDiffContext`.

## The DiffContext Data Structure

The `buildDiffContext` function returns a structured `DiffContext` object (defined in [`src/diff-analyzer.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/src/diff-analyzer.ts)) that captures the complete impact surface:

```typescript
{
  projectName,
  changedFiles,    // Original input paths
  changedNodes,    // Directly modified graph nodes
  affectedNodes,   // Downstream neighbors requiring attention
  impactedEdges,   // Relationships crossing change boundaries
  affectedLayers,  // Architectural layers touched by the change
  unmappedFiles    // Files not yet in the knowledge graph
}

```

This object serves as the single source of truth for all subsequent reporting and risk assessment calculations.

## Generating Human-Readable Reports with `formatDiffAnalysis`

After computing the `DiffContext`, the `formatDiffAnalysis` helper (lines 58–94 in the same file) renders a Markdown report that surfaces the ripple effects for developers. The report includes:

- Changed components with name, type, summary, and complexity metrics
- Downstream components that may need regression testing
- Architectural layers impacted by the change set
- Relationships (edges) that cross the changed/affected boundary
- Unmapped files requiring graph updates
- A risk assessment based on complexity, cross-layer impact, blast radius, and unmapped file count

The risk scoring logic evaluates high-complexity changes, cross-layer impacts, wide blast radii, and unmapped files to generate an overall change risk profile.

## Practical Implementation Example

The following TypeScript example demonstrates how to load a persisted knowledge graph and execute the diff impact analysis:

```typescript
import { readKnowledgeGraph } from '@understand-anything/core';
import { buildDiffContext, formatDiffAnalysis } from '@understand-anything/plugin';

// 1️⃣ Load the persisted knowledge graph
const graph = await readKnowledgeGraph('/path/to/.understand-anything/knowledge-graph.json');

// 2️⃣ Gather changed files from your CI pipeline
const changedFiles = [
  'src/services/userService.ts',
  'src/utils/auth.ts',
];

// 3️⃣ Build the diff context – this is where ripple tracking happens
const diffCtx = buildDiffContext(graph, changedFiles);

// 4️⃣ Generate the Markdown report
const report = formatDiffAnalysis(diffCtx);
console.log(report);

```

Running this snippet produces a Markdown document identical in structure to the UI's "Diff Overlay" view, with the underlying algorithm mirroring the production implementation exactly.

## Architecture and Key Source Files

The diff impact analysis spans multiple packages within the Understand-Anything repository:

- **[`src/diff-analyzer.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/src/diff-analyzer.ts)**: Core logic implementing `buildDiffContext` and `formatDiffAnalysis` for ripple effect calculation and report generation.
- **[`src/__tests__/diff-analyzer.test.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/src/__tests__/diff-analyzer.test.ts)**: Test suite validating empty-diff handling, node mapping accuracy, and ripple calculations.
- **[`packages/core/src/types.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/types.ts)**: TypeScript definitions for `KnowledgeGraph`, `GraphNode`, `GraphEdge`, and `Layer` used throughout the analyzer.
- **[`packages/core/src/graph-builder.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/core/src/graph-builder.ts)**: Generates the knowledge graph from source files, supplying the structural data consumed by the diff analyzer.
- **[`packages/dashboard/src/store.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/packages/dashboard/src/store.ts)**: UI state management including the `diffMode` flag that toggles the overlay visualization of diff analysis results.

## Summary

- **Diff impact analysis** in Understand-Anything converts raw Git file lists into structured impact assessments using the knowledge graph.
- The `buildDiffContext` function in [`src/diff-analyzer.ts`](https://github.com/Egonex-AI/Understand-Anything/blob/main/src/diff-analyzer.ts) executes a three-phase algorithm: node mapping, containment expansion via "contains" edges, and neighbor discovery for ripple detection.
- **Affected nodes** represent one-hop neighbors of changed files, while **impacted edges** highlight relationships crossing the change boundary.
- The analysis operates entirely in-memory without external service dependencies, making it suitable for CI/CD pipelines.
- The `formatDiffAnalysis` function generates Markdown reports with built-in risk assessment based on complexity, blast radius, and cross-layer impact.

## Frequently Asked Questions

### How does the algorithm handle files not yet mapped in the knowledge graph?

Files that do not match any `node.filePath` in the knowledge graph are collected into the `unmappedFiles` array within the `DiffContext`. These entries are surfaced in the final report to flag components requiring incremental graph updates, ensuring the analysis remains transparent about coverage gaps.

### What defines an "affected node" versus a "changed node" in the analysis?

**Changed nodes** are graph entities directly matching the input file paths or their immediate children via "contains" relationships. **Affected nodes** are one-hop neighbors connected to changed nodes through any edge type, representing downstream components that may experience ripple effects but were not directly modified in the diff.

### Does the diff impact analysis require external services or APIs?

No. The analysis is entirely self-contained and operates purely on the in-memory knowledge graph. No external services are called during the execution of `buildDiffContext` or `formatDiffAnalysis`, making the pipeline deterministic and suitable for offline CI environments.

### How is the risk assessment score calculated in the Markdown report?

The risk assessment in `formatDiffAnalysis` (lines 58–94) evaluates four factors: high complexity scores of changed components, cross-layer architectural impact, wide blast radius (large number of affected nodes), and the presence of unmapped files. When multiple factors are present, the report flags the change as higher risk.