How to Debug Graph Generation Issues Using Intermediate Files in Understand Anything
The .understand-anything/intermediate/ directory contains staged JSON snapshots—including scan-result.json, batch-*.json, and assembled-graph.json—that let you trace exactly where the knowledge graph pipeline fails without modifying source code.
The Understand Anything repository (Lum1104/Understand-Anything) constructs knowledge graphs through a multi-phase agent pipeline. Each phase writes its output to a hidden intermediate directory, creating a permanent audit trail of read-only snapshots. By inspecting these files, you can pinpoint exactly which agent produced incorrect data, verify batch fragmentation, and manually re-run aggregation steps.
How the Intermediate Directory Structure Works
The pipeline generates files in <project-root>/.understand-anything/intermediate/ across five distinct phases. These files are intentionally read-only snapshots; they are never overwritten after creation and are not written back to source code, making them safe references for debugging.
Phase 1: Scan and Batch Creation
The project-scanner agent initializes the directory and writes scan-result.json, a complete inventory of discovered files and symbols. Subsequently, file-analyzer agents split this result into independent batches saved as batch-<index>.json or, for large projects, split parts like batch-<index>-part-<k>.json. This behavior is documented in understand-anything-plugin/agents/project-scanner.md.
Phase 2: Merge
The Python helper script merge-batch-graphs.py (located in understand-anything-plugin/skills/understand/) reads all batch-*.json fragments and produces a single assembled-graph.json. This merge step is described in understand-anything-plugin/skills/understand/SKILL.md.
Phase 3: Review
The graph-reviewer agent analyzes the assembled graph and writes diagnostics to review.json and a human-friendly summary to assemble-review.json, logging warnings about duplicate nodes or orphan edges.
Phase 4-5: Post-Processing and Cleanup
Additional agents may generate layers.json, tour.json, or fingerprint-input.json. After a successful run, the cleanup phase removes most intermediate files but deliberately retains scan-result.json to enable incremental rescans.
Step-by-Step Debugging Workflow
Follow this sequence to isolate graph generation failures using the intermediate files.
-
Locate the directory – Navigate to
<project-root>/.understand-anything/intermediate. The path creation logic resides inunderstand-anything-plugin/agents/project-scanner.md. -
Verify scan coverage – Open
scan-result.jsonto confirm every source file was discovered. The test fileunderstand-anything-plugin/src/__tests__/merge-recover-imports.test.mjscreates this file and asserts its shape, providing a reference for expected structure. -
Inspect batch fragments – Each
batch-*.jsoncontains a partial graph fragment. If a module is missing from the final graph, locate its specific fragment. The testmerge-recover-imports.test.mjswrites several batch files (lines 64‑72) that you can compare against your output. -
Re-run the merge manually – Execute
merge-batch-graphs.pyagainst your intermediate directory to reproduce the assembly step. Running this manually can reveal parsing errors or missing batch indices without re-running the entire pipeline. -
Validate the assembled graph – Check
assembled-graph.jsonagainst the schema defined inunderstand-anything-plugin/packages/core/src/schema.ts. The testschema.test.tsautomatically performs this validation and serves as a reference implementation. -
Check reviewer diagnostics – Examine
review.jsonfor warnings. These files are generated as described in the graph-reviewer section ofunderstand-anything-plugin/skills/understand/SKILL.md. -
Iterate locally – Edit the offending
batch-*.jsonfile (or the original source) and re-runmerge-batch-graphs.py. Because the pipeline reads directly from these intermediate files, you can debug without changing other repository logic.
Code Example: Reading Intermediate Files Programmatically
Use this TypeScript snippet to load and inspect intermediate files from a Node.js environment. This code reads existing snapshots without invoking external commands or creating new files.
import { readFileSync } from 'fs';
import { join } from 'path';
// Assume `projectRoot` is the absolute path to the target repo
const projectRoot = process.env.PROJECT_ROOT ?? '.';
const intermediate = join(projectRoot, '.understand-anything', 'intermediate');
// 1️⃣ Load the scan result
const scanResult = JSON.parse(
readFileSync(join(intermediate, 'scan-result.json'), 'utf-8')
);
console.log('Scanned files:', scanResult.files.length);
// 2️⃣ Load a specific batch (e.g., batch-0.json)
const batch0 = JSON.parse(
readFileSync(join(intermediate, 'batch-0.json'), 'utf-8')
);
console.log('Batch 0 nodes:', batch0.nodes?.length);
// 3️⃣ Load the final assembled graph
const assembled = JSON.parse(
readFileSync(join(intermediate, 'assembled-graph.json'), 'utf-8')
);
console.log('Assembled graph nodes:', assembled.nodes?.length);
// 4️⃣ Inspect reviewer diagnostics
const review = JSON.parse(
readFileSync(join(intermediate, 'review.json'), 'utf-8')
);
if (review.warnings?.length) {
console.warn('Reviewer warnings:', review.warnings);
}
Summary
- Intermediate files are staged in
.understand-anything/intermediate/and serve as read-only snapshots of each pipeline phase. - Key files include
scan-result.json(discovery),batch-*.json(fragments),assembled-graph.json(merged output), andreview.json(diagnostics). - Manual debugging involves inspecting JSON at each stage and optionally re-running
merge-batch-graphs.pywithout triggering the full pipeline. - Validation should be performed against the schema in
packages/core/src/schema.tsas demonstrated inschema.test.ts. - Incremental rescans are supported because
scan-result.jsonpersists after other intermediate files are cleaned up.
Frequently Asked Questions
Where exactly are the intermediate debug files stored?
The files are stored in a hidden directory at <project-root>/.understand-anything/intermediate/. This path is created automatically by the project-scanner agent during the initial scan phase, as documented in understand-anything-plugin/agents/project-scanner.md.
Can I safely modify intermediate files to fix graph errors?
Yes. The intermediate files are read-only snapshots that the pipeline never modifies after creation. You can edit a specific batch-*.json file to correct data issues and manually re-run merge-batch-graphs.py to generate a new assembled-graph.json without affecting the source repository or other pipeline stages.
How do I validate that my assembled graph follows the correct structure?
Validate assembled-graph.json against the TypeScript schema defined in understand-anything-plugin/packages/core/src/schema.ts. The repository includes automated validation in schema.test.ts, which you can reference to understand the expected node and edge structure.
Why does only scan-result.json remain after a successful pipeline run?
While the cleanup phase removes most intermediate files to save space, it retains scan-result.json specifically to enable incremental rescans. This allows subsequent runs to skip the full file discovery process if the project structure has not changed, significantly speeding up iterative debugging.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →