How Egonex AI Handles Changed Files in Large Codebases with Incremental Analysis
Egonex AI's "Understand Anything" plugin fingerprints every source file to detect structural changes, then classifies updates into four distinct tiers—ranging from skip to full rebuild—to ensure only the minimal necessary graph components are recomputed.
The Egonex AI "Understand Anything" plugin solves the scalability challenge of maintaining accurate knowledge graphs in rapidly evolving repositories. Instead of reprocessing thousands of files after every modification, the system employs structural fingerprinting and intelligent change classification to update only affected graph segments. This incremental analysis pipeline keeps performance proportional to change magnitude, making it feasible to track massive codebases in real time.
File Fingerprinting and Change Detection
The foundation of incremental analysis lies in lightweight file descriptors that capture structural essence without storing full source text.
Structural Fingerprints and SHA-256 Hashing
In packages/core/src/fingerprint.ts, the system generates two complementary identifiers for every tracked file. The structural fingerprint extracts function signatures, class members, and import/export statements to represent the file's API surface. This is paired with a SHA-256 content hash that detects any textual modification. Together, these fingerprints distinguish between cosmetic changes (whitespace, comments) and structural modifications that alter the knowledge graph's topology.
The analyzeChanges Function
When the IDE or git hook reports modified files, the analyzeChanges function compares current fingerprints against the stored state in .understand-anything/fingerprint-store.json. This comparison produces a ChangeAnalysis object that categorizes each file as new, deleted, structurally changed, or cosmetic-only. The function accepts the project root, changed file paths, the stored fingerprint registry, and a PluginRegistry instance to handle language-specific parsing.
Four-Tier Update Classification
After detecting changes, the system must decide how much of the knowledge graph to invalidate. In packages/core/src/change-classifier.ts, the classifyUpdate function implements a deterministic decision matrix that scales from zero-cost skips to complete rebuilds.
Decision Matrix and Thresholds
The classifier evaluates the ChangeAnalysis against the total file count and existing paths to select one of four actions:
- SKIP: Triggers when no structural changes exist—only cosmetic modifications or no changes at all. The existing knowledge graph remains valid.
- PARTIAL_UPDATE: Activates when fewer than 10 files have structural changes confined to specific directories. Only the listed files require re-analysis.
- ARCHITECTURE_UPDATE: Occurs when new or deleted top-level directories appear, or when more than 10 files change structurally. This triggers regeneration of module-dependency layouts via
packages/core/src/analyzer/graph-builder.ts. - FULL_UPDATE: Reserved for massive refactors exceeding 30 structural changes or affecting more than 50% of the project's files. This rebuilds the entire graph from scratch.
The classifyUpdate Function
This function returns a structured decision object containing the selected action, the specific filesToReanalyze, and boolean flags indicating whether rerunArchitecture analysis or tour generation steps must execute. The thresholds are hardcoded based on performance benchmarks to balance accuracy against computational cost.
Incremental Pipeline Execution
The main analysis workflow chains analyzeChanges → classifyUpdate → specialized execution paths based on the classification result.
Partial Updates and Patching
For PARTIAL_UPDATE decisions, the pipeline sends only the affected files through language-specific extractors provided by packages/core/src/plugins/registry.ts. The graph builder then patches the knowledge graph in-place, updating nodes and edges without touching unrelated structures. This path executes in milliseconds for single-file changes.
Architecture-Wide Updates
When rerunArchitecture is true, the system recomputes the high-level module dependency graph and directory structure. This handles scenarios like adding new service directories or moving files between major components, ensuring the architectural visualization remains accurate.
Full Rebuild Scenarios
FULL_UPDATE bypasses all incremental optimizations, clearing the existing graph and reprocessing every file through the complete extraction pipeline. The system reserves this expensive operation for branch switches, massive refactoring operations, or initial project imports.
Implementation Example
The following TypeScript demonstrates the complete incremental analysis workflow:
import { analyzeChanges } from "./packages/core/src/fingerprint.js";
import { classifyUpdate } from "./packages/core/src/change-classifier.js";
import { PluginRegistry } from "./packages/core/src/plugins/registry.js";
// 1️⃣ Load the previously stored fingerprint store
const stored = await import("./.understand-anything/fingerprint-store.json");
// 2️⃣ Build a plugin registry for language-specific parsing
const registry = new PluginRegistry();
// 3️⃣ Detect changed files from git diff or IDE events
const changedFiles = ["src/foo.ts", "src/bar.go"];
// 4️⃣ Produce detailed change analysis
const analysis = analyzeChanges(
"/path/to/project",
changedFiles,
stored,
registry,
);
// 5️⃣ Decide recomputation scope
const decision = classifyUpdate(
analysis,
Object.keys(stored.files).length,
Object.keys(stored.files),
);
console.log(decision);
/*
{
action: "PARTIAL_UPDATE",
filesToReanalyze: ["src/foo.ts"],
rerunArchitecture: false,
rerunTour: false,
reason: "1 file have structural changes: 1 new"
}
*/
Partial update processes only src/foo.ts through the extractors. Architecture update would set rerunArchitecture: true, triggering the graph-builder module. Full update reprocesses the entire project structure.
Summary
- File fingerprinting in
packages/core/src/fingerprint.tscombines structural signatures with SHA-256 hashes to distinguish meaningful changes from cosmetic edits. - Change classification via
classifyUpdatesorts modifications into SKIP, PARTIAL_UPDATE, ARCHITECTURE_UPDATE, or FULL_UPDATE tiers based on configurable thresholds. - Incremental execution patches the knowledge graph in-place for small changes, while triggering architecture-wide or full rebuilds only when necessary.
- Performance scaling ensures that modifying a single file in a 10,000-file repository completes in milliseconds rather than minutes.
Frequently Asked Questions
What is the difference between a structural change and a cosmetic change?
A structural change modifies function signatures, class definitions, import/export statements, or other elements that affect the knowledge graph's nodes and edges. Cosmetic changes alter whitespace, comments, or formatting without changing the API surface. The fingerprinting system detects structural changes by comparing extracted signatures, while SHA-256 hashes catch cosmetic edits, allowing the classifier to skip graph updates for purely cosmetic modifications.
How does the system handle newly created or deleted directories?
New or deleted top-level directories trigger the ARCHITECTURE_UPDATE action, as implemented in packages/core/src/change-classifier.ts. This recomputes the module dependency graph and directory structure through packages/core/src/analyzer/graph-builder.ts, ensuring the architectural visualization reflects the new project layout. Subdirectory changes within existing modules typically trigger PARTIAL_UPDATE unless they exceed the 10-file structural change threshold.
Can the thresholds for update classification be configured?
The current implementation in packages/core/src/change-classifier.ts uses hardcoded thresholds—10 files for partial updates, 30 files for full updates, and 50% of total files for mass rebuilds—based on performance benchmarks in the packages/core/src/__tests__/change-classifier.test.ts suite. These values optimize the balance between analysis accuracy and computational cost for typical enterprise codebases.
How does fingerprinting maintain performance across thousands of files?
Structural fingerprints are lightweight descriptors containing only function signatures, class members, and import/export data, not full source text. Stored in .understand-anything/fingerprint-store.json, these descriptors enable analyzeChanges to compare file states using hash comparisons rather than text diffing, resulting in O(1) per-file operations. This design ensures the initial change detection phase completes in milliseconds regardless of repository size.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →