Fingerprint-Based Change Detection for Incremental Updates in Understand-Anything

The fingerprint-based change detection system in Understand-Anything uses SHA-256 content hashes and structural signatures to classify file changes as NONE, COSMETIC, or STRUCTURAL, enabling incremental knowledge graph updates that skip unchanged files and minimize recomputation.

The Egonex-AI/Understand-Anything repository implements a sophisticated incremental update mechanism that avoids rebuilding the entire knowledge graph on every code change. By leveraging fingerprint-based change detection, the system compares cryptographic hashes and parsed structural signatures to determine exactly which files require reprocessing. This approach dramatically reduces CPU and I/O overhead for large codebases while maintaining accuracy through conservative fallbacks for unsupported languages.

Core Architecture of the Fingerprint System

At the heart of the incremental update pipeline lies the FileFingerprint interface, defined in understand-anything-plugin/packages/core/src/fingerprint.ts (lines 9-38). This data structure captures both content identity and semantic structure for every source file in the project.

export interface FileFingerprint {
  filePath: string;
  contentHash: string;
  functions: FunctionFingerprint[];
  classes:   ClassFingerprint[];
  imports:   ImportFingerprint[];
  exports:   string[];
  totalLines: number;
  hasStructuralAnalysis: boolean;
}

The content hash stores a SHA-256 digest of the entire file, while the structural arrays (functions, classes, imports, exports) contain signatures extracted by tree-sitter parsers. The hasStructuralAnalysis flag indicates whether full parsing succeeded or if the system must rely on hash-only comparison.

Building the Fingerprint Store

The buildFingerprintStore function (see fingerprint.ts lines 53-84) initializes the fingerprint database by walking every file in the project. It accepts a PluginRegistry that supplies language-specific tree-sitter parsers via registry.analyzeFile.

export function buildFingerprintStore(
  projectDir: string,
  filePaths: string[],
  registry: PluginRegistry,
  gitCommitHash: string,
): FingerprintStore {
  // Implementation walks files, extracts fingerprints, 
  // and creates hash-only fallbacks for unsupported languages
}

Files without tree-sitter support receive a hash-only fingerprint and are treated conservatively as potential structural changes. The resulting FingerprintStore is serialized to fingerprints.json as a versioned map of filePath → FileFingerprint, anchored to a specific Git commit hash.

Detecting Changes with Structural Analysis

When the system runs incrementally, only files reported by the VCS as modified are re-fingerprinted. The compareFingerprints function (located in fingerprint.ts lines 31-46) evaluates the delta between the stored fingerprint and the newly computed one.

export function compareFingerprints(
  oldFp: FileFingerprint, 
  newFp: FileFingerprint
): FileChangeResult {
  // Compares content hashes, then structural signatures
  // Returns classification with human-readable details
}

The comparison logic follows a strict hierarchy:

  1. Hash comparison – If SHA-256 digests match, the result is NONE (no work required).
  2. Structural signature comparison – If the hash differs but all function signatures, class members, and import/export lists remain identical, the change is COSMETIC (internal logic changed without affecting the knowledge graph structure).
  3. Structural mismatch – If any signature differs or structural analysis is unavailable (hasStructuralAnalysis: false), the change is STRUCTURAL (requires graph recomputation).

The function also generates a details array describing specific deltas (e.g., "new function: foo", "imports changed") for downstream reporting.

Aggregating Results for Incremental Updates

The analyzeChanges function (see fingerprint.ts lines 94-108) orchestrates the incremental pipeline by aggregating per-file results into a ChangeAnalysis object.

export function analyzeChanges(
  projectDir: string,
  changedFiles: string[],
  existingStore: FingerprintStore,
  registry: PluginRegistry,
): ChangeAnalysis {
  // Re-fingerprints changed files, compares with stored versions,
  // and categorizes results for the graph builder
}

This categorization separates files into new, deleted, unchanged, cosmetically changed, and structurally changed buckets. The incremental graph builder uses this ChangeAnalysis to trigger node and edge updates only for STRUCTURAL changes, while COSMETIC modifications preserve the existing knowledge graph intact.

Practical Implementation Example

The following TypeScript example demonstrates the complete workflow from initial store generation to incremental analysis:

import { 
  buildFingerprintStore,
  analyzeChanges,
  readFileSync,
  writeFileSync,
} from '@understand-anything/core';

// 1️⃣ Generate a fingerprint store for the whole project
const allFiles = await globby(['**/*.{ts,js,tsx,jsx,py,go,rs}']);
const registry = await import('@understand-anything/core/src/plugins/registry.js');
const store = buildFingerprintStore('/my/project', allFiles, registry, 'HEAD');

// Persist for later runs
writeFileSync('fingerprints.json', JSON.stringify(store));

// 2️⃣ Later – after git reports changed files
const changed = ['src/utils.ts', 'README.md']; // from `git diff --name-only`
const previousStore = JSON.parse(readFileSync('fingerprints.json', 'utf-8'));

const analysis = analyzeChanges('/my/project', changed, previousStore, registry);

// Use the analysis to trigger incremental graph updates
for (const file of analysis.structurallyChangedFiles) {
  console.log(`⚙️ Re‑process structural change in ${file}`);
}
for (const file of analysis.cosmeticOnlyFiles) {
  console.log(`✨ Cosmetic change only – graph unchanged for ${file}`);
}

Key supporting files include types.ts for shared type definitions, registry.ts for language-specific parser management, and staleness.ts for higher-level staleness detection logic. The test suite in change-classifier.test.ts validates the correct classification of change types across edge cases.

Summary

  • Fingerprint-based change detection in Understand-Anything combines SHA-256 content hashes with tree-sitter structural signatures to minimize unnecessary recomputation.
  • The FileFingerprint interface stores both content identity and semantic structure, enabling precise change classification.
  • Hash-only fallbacks ensure correctness for languages lacking tree-sitter parsers by conservatively marking changes as STRUCTURAL.
  • The compareFingerprints function implements a three-tier classification (NONE, COSMETIC, STRUCTURAL) that determines whether the knowledge graph requires updates.
  • Only files flagged as STRUCTURAL trigger graph recomputation, while COSMETIC changes preserve existing nodes and edges.

Frequently Asked Questions

How does Understand-Anything handle files without tree-sitter support?

Files for which no tree-sitter parser exists in the PluginRegistry receive a hash-only fingerprint with hasStructuralAnalysis: false. According to the source code in fingerprint.ts, these files are treated conservatively: any content change is classified as STRUCTURAL rather than COSMETIC, ensuring the knowledge graph remains accurate even when deep parsing is unavailable.

What distinguishes COSMETIC from STRUCTURAL changes in the fingerprint system?

COSMETIC changes occur when the SHA-256 hash differs but all structural signatures (function signatures, class definitions, import/export lists) remain identical, indicating only internal implementation logic changed. STRUCTURAL changes occur when any signature differs, imports/exports are modified, or when structural analysis is unavailable. Only STRUCTURAL changes require knowledge graph updates, while COSMETIC changes skip recomputation entirely.

Where is the fingerprint data persisted between analysis runs?

The buildFingerprintStore function serializes the FingerprintStore object to a JSON file named fingerprints.json in the project root. This file contains a versioned map of file paths to FileFingerprint objects, including the Git commit hash at the time of creation. Subsequent runs load this store via readFileSync and pass it to analyzeChanges for comparison against newly computed fingerprints.

How does fingerprint-based detection improve performance on large codebases?

By comparing SHA-256 hashes first, the system eliminates parsing overhead for unchanged files entirely. The structural signature comparison further filters out changes that do not affect the knowledge graph topology (COSMETIC changes). As implemented in analyzeChanges, this classification ensures that only files with STRUCTURAL modifications trigger the expensive graph rebuilding process, reducing CPU and I/O costs from O(total files) to O(changed files) for incremental updates.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →