How the Understand-Anything Fingerprint System Effectively Detects Structural Changes in Code
The fingerprint system detects structural changes by parsing code into AST-based signatures, creating SHA-256 content hashes, and comparing public API surfaces to distinguish between cosmetic edits and changes that impact the knowledge graph.
The fingerprint system is the core mechanism that lets Understand‑Anything decide whether a file modification is merely cosmetic or truly structural (i.e., it can affect the knowledge graph). Unlike simple diff tools that compare line-by-line text, this system analyzes the semantic structure of code to determine if a change alters the public contract of functions, classes, or modules. By focusing on signatures rather than implementation details, it provides a reliable, low-overhead way to trigger partial or full knowledge graph updates only when necessary.
The Five-Stage Detection Pipeline
The system operates through five distinct stages, each implemented in specific source files within the packages/core/src/ directory.
Stage 1: Extracting Structural Fingerprints
In extractFileFingerprint (fingerprint.ts‑L122), Tree‑sitter parses each file to extract signatures including function names, parameters, return types, class members, and import/export statements. The resulting FileFingerprint object records only the shape of the code—its public contract—rather than the internal implementation logic.
Stage 2: Storing Content Hashes
The buildFingerprintStore function (fingerprint.ts‑L85) generates a SHA‑256 content hash for every file and collects all structural fingerprints into a FingerprintStore. This store is persisted alongside the git commit hash that generated it, creating a snapshot of the project's structural state at a specific point in time.
Stage 3: Comparing Old vs. New Fingerprints
When a new commit is analyzed, compareFingerprints (fingerprint.ts‑L45) performs a two-tier check. First, it compares content hashes for a fast path—identical hashes mean no analysis is needed. If hashes differ, it executes a deep signature diff that detects new or removed symbols, changed parameter lists, altered return types, modified export status, and significant line-count variations.
Stage 4: Generating the Change Report
The analyzeChanges function (fingerprint.ts‑L163) orchestrates the comparison across all files changed between commits. It aggregates results into categories: newFiles, deletedFiles, structurallyChangedFiles, cosmeticOnlyFiles, and unchangedFiles. Each structural change includes human‑readable details explaining exactly what signature changed.
Stage 5: Classifying the Update
Finally, classifyUpdate (change-classifier.ts‑L85) interprets the ChangeAnalysis output and selects an action: SKIP, PARTIAL_UPDATE, ARCHITECTURE_UPDATE, or FULL_UPDATE. Structural changes drive expensive rebuilds, while cosmetic-only edits allow the system to skip unnecessary processing.
Key Design Principles
Several architectural decisions make the fingerprint system both accurate and efficient:
- Signature-Only Focus: By storing only public contracts (function signatures, class interfaces, export lists), internal logic modifications like loop body changes are treated as cosmetic, avoiding unnecessary graph rebuilds.
- Fast Hash Shortcut: Identical SHA‑256 hashes bypass expensive structural diff entirely, returning
NONEimmediately. - Conservative Fallback: If Tree‑sitter cannot parse a file, the system assumes a structural change (
STRUCTURAL) to maintain safety. - Granular Diff Tracking: The comparison algorithm records specific details about which symbols changed, enabling precise downstream reporting.
- Threshold-Based Classification: The classifier uses counts and percentages of structural changes to decide between cheap partial updates and full re‑analysis, optimizing for large projects.
Implementation Examples
Building a Fingerprint Store
To create an initial fingerprint store for a project:
import { buildFingerprintStore } from "./fingerprint.js";
import { createPluginRegistry } from "./plugins/registry.js";
const projectDir = "/path/to/my/project";
const allFiles = ["src/index.ts", "src/utils.ts", "package.json"]; // normally discovered via glob
const registry = createPluginRegistry(); // provides tree‑sitter analysis
const gitCommitHash = "a1b2c3d4"; // obtain via `git rev-parse HEAD`
const store = buildFingerprintStore(projectDir, allFiles, registry, gitCommitHash);
// store can be written to `.understand-anything/fingerprint.json`
This invokes buildFingerprintStore from packages/core/src/fingerprint.ts to analyze each file and compute both structural fingerprints and content hashes.
Detecting Changes Between Commits
To analyze what changed between two git states:
import { getChangedFiles } from "./staleness.js";
import { analyzeChanges } from "./fingerprint.js";
import { createPluginRegistry } from "./plugins/registry.js";
const projectDir = "/path/to/my/project";
const oldStore = /* load previous FingerprintStore */;
const registry = createPluginRegistry();
const changedFiles = getChangedFiles(projectDir, oldStore.gitCommitHash);
const analysis = analyzeChanges(projectDir, changedFiles, oldStore, registry);
console.log("Structural changes:", analysis.structurallyChangedFiles);
console.log("Cosmetic only:", analysis.cosmeticOnlyFiles);
This leverages getChangedFiles from packages/core/src/staleness.ts‑L26) and analyzeChanges from fingerprint.ts‑L163) to categorize modifications.
Classifying Update Actions
To determine the required pipeline action based on detected changes:
import { classifyUpdate } from "./change-classifier.js";
const totalFiles = 1200; // total files currently in the graph
const decision = classifyUpdate(analysis, totalFiles);
switch (decision.action) {
case "SKIP":
console.log("No structural impact – skip re‑analysis");
break;
case "PARTIAL_UPDATE":
console.log("Re‑analyze only:", decision.filesToReanalyze);
break;
case "ARCHITECTURE_UPDATE":
console.log("Run architecture pipeline");
break;
case "FULL_UPDATE":
console.log("Full rebuild required");
}
The classifyUpdate function in packages/core/src/change-classifier.ts‑L85) uses the analysis results to minimize computational overhead.
Summary
- The fingerprint system uses Tree‑sitter to extract code signatures and SHA‑256 hashes to detect structural changes in the Understand‑Anything codebase.
- Five stages handle extraction, storage, comparison, reporting, and classification of code modifications.
- Cosmetic changes (internal logic updates) are filtered out by comparing only public API surfaces, preventing unnecessary knowledge graph rebuilds.
- Conservative defaults ensure that unparsable files are treated as structural changes to maintain graph integrity.
- Source files implementing this logic include
fingerprint.ts,change-classifier.ts, andstaleness.tsinpackages/core/src/.
Frequently Asked Questions
How does the fingerprint system distinguish between structural and cosmetic changes?
The system compares signatures—function names, parameters, return types, class members, and exports—rather than raw text. If only the function body changes while the signature remains identical, the change is classified as cosmetic. If the signature itself changes (new parameters, renamed functions, different return types), it is marked as structural because it could affect how other parts of the codebase interact with that symbol.
What happens if Tree‑sitter fails to parse a file?
If a file cannot be parsed by Tree‑sitter, the system defaults to STRUCTURAL in compareFingerprints (fingerprint.ts‑L45). This conservative approach ensures that any unknown code changes are treated as potentially impacting the knowledge graph, preventing stale or incorrect analysis results when the parser encounters unsupported syntax or binary files.
Can the fingerprint system handle partial or incremental updates?
Yes. The classifyUpdate function (change-classifier.ts‑L85) evaluates the ratio of structurally changed files to total files. Small numbers of structural changes trigger PARTIAL_UPDATE, which re‑analyzes only the affected files, while widespread changes initiate FULL_UPDATE or ARCHITECTURE_UPDATE pipelines. This incremental approach keeps analysis fast for large projects.
Where are fingerprints stored and how are they associated with commits?
Fingerprints are stored in a FingerprintStore object generated by buildFingerprintStore (fingerprint.ts‑L85). This collection includes each file's SHA‑256 hash and structural fingerprint, persisted alongside the git commit hash that produced it. When analyzing new commits, the system loads the previous store and compares it against fresh fingerprints to determine exactly what changed.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →