How the Graph Validation Review Process Works: Inline vs. LLM Validation
The Understand-Anything pipeline validates knowledge graphs through deterministic inline checks on every run, while optional LLM-powered review provides semantic analysis when users specify the --review flag.
The understand-anything repository constructs a knowledge graph from your codebase and validates it before delivering results to the dashboard. This graph validation review process operates through two distinct phases: a fast, deterministic inline check that runs automatically after graph assembly, and an optional deep review powered by an LLM sub-agent triggered by explicit user request.
How Inline Validation Works
Inline validation executes immediately after the graph builder produces the JSON structure. This deterministic phase ensures structural integrity before any expensive operations occur.
The orchestrator invokes validateGraph() from packages/core/src/schema.ts (line 499). This function applies Zod-style schema validation to verify node and edge types, required fields, and referential integrity. It checks that all edge targets exist as valid nodes and detects duplicate IDs within the graph structure.
If validation fails, the function returns a ValidationResult object containing specific errors and warnings. The orchestrator stores these in the $PHASE_WARNINGS variable and includes them directly in the final output report. You can optionally disable this check by passing { validate: false }, as implemented in packages/core/src/persistence/index.ts (lines 94-95), though this is rarely recommended.
How LLM-Powered Review Works
When you append the --review flag to commands like understand --full --review, the pipeline activates the graph-reviewer sub-agent for semantic analysis.
The orchestrator dispatches this agent (defined in agents/graph-reviewer.md) along with the assembled graph and accumulated $PHASE_WARNINGS. The sub-agent processes natural language instructions from skills/understand/graph-reviewer-prompt.md, analyzing the graph for semantic duplicates, missing documentation, and suspicious edge weights that static schemas cannot express.
The LLM generates a human-readable report that the orchestrator merges with inline warnings, presenting a comprehensive validation output that combines structural and semantic assessments.
Why the Validation Pipeline Uses Two Phases
The split between inline and LLM validation serves three critical architectural purposes:
- Performance and token cost. Inline validation runs in milliseconds using pure JavaScript and consumes zero LLM tokens. Transmitting graphs exceeding 500 nodes to an LLM can consume tens of thousands of tokens, making the gated approach essential for cost-effective default workflows.
- Deterministic baselines. The schema guarantees structural correctness on every execution, providing developers with a reliable, reproducible foundation that never varies between runs.
- Semantic insight. The LLM identifies subtle logical errors—such as duplicated services, incomplete documentation, or anomalous edge weights—that deterministic rules cannot practically express.
Code Implementation Examples
Implement inline validation using the core schema module:
import { validateGraph } from '@understand-anything/core/schema';
// After the graph is built:
const result = validateGraph(assembledGraph);
if (result.issues.length) {
// handle warnings or abort pipeline
console.error('Validation failed:', result.issues);
}
Conditionally trigger LLM review based on user arguments:
if (ARGUMENTS.includes('--review')) {
await dispatchSubagent('graph-reviewer', {
promptFile: 'skills/understand/graph-reviewer-prompt.md',
graph: assembledGraph,
priorWarnings: phaseWarnings,
});
}
Summary
- Inline validation always runs via
validateGraphinpackages/core/src/schema.tsto verify JSON structure, referential integrity, and ID uniqueness. - LLM review only activates with the
--reviewflag, utilizing thegraph-reviewersub-agent defined inagents/graph-reviewer.md. - The graph-reviewer prompt lives in
skills/understand/graph-reviewer-prompt.mdand handles semantic analysis beyond schema constraints. - Validation can be programmatically disabled via
{ validate: false }inpackages/core/src/persistence/index.ts(lines 94-95). - The two-phase approach balances execution speed and API costs against deep semantic insight.
Frequently Asked Questions
What is the difference between inline and LLM graph validation?
Inline validation uses deterministic TypeScript/Zod schema checks in packages/core/src/schema.ts to verify JSON structure and referential integrity. LLM validation sends the assembled graph to a sub-agent that runs skills/understand/graph-reviewer-prompt.md to detect semantic issues like duplicate services or missing documentation that static rules cannot express.
When does the LLM validation run?
LLM validation runs only when you explicitly include the --review flag in your command (for example, understand --full --review). Without this flag, the pipeline executes only the inline validation phase to minimize latency and token costs.
Can I disable the inline validation step?
Yes, though this is rarely recommended. You can pass { validate: false } to skip the schema check, as implemented in packages/core/src/persistence/index.ts lines 94-95. Disabling validation risks propagating structural errors downstream to the dashboard.
How does the LLM reviewer know what to check?
The orchestrator feeds the sub-agent instructions from skills/understand/graph-reviewer-prompt.md. This prompt template contains natural-language instructions guiding the LLM to examine edge weights, identify semantic duplicates, and flag documentation gaps that the deterministic validator cannot catch.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →