Architecture Layer Identification in Understand Anything: A Two-Stage Pipeline Explained
Architecture layer identification in Understand Anything assigns every file node to a logical layer through an initial heuristic scan of directory paths followed by an optional LLM-driven refinement that surfaces project-specific architectural groupings.
The Understand Anything engine transforms raw source code into a navigable knowledge graph where architecture layer identification plays a central role. By automatically categorizing files into logical layers such as "API" or "Service," the system creates a high-level architectural view that helps developers comprehend large codebases without manual documentation. This process is implemented in the layer-detector.ts module and operates in two distinct stages.
The Two-Stage Identification Pipeline
The system employs a deterministic-to-intelligent pipeline that balances speed with accuracy:
- Heuristic Detection — Matches directory names against curated patterns in
LAYER_PATTERNSfor instant, rule-based classification - LLM Refinement — Generates prompts for a language model to propose domain-specific layers based on the complete file structure
Files that fail heuristic matching default to the "Core" layer, while files unmatched by the LLM stage land in "Other".
Stage 1: Heuristic Layer Detection
The heuristic stage provides immediate, deterministic classification by analyzing file paths without external API calls.
Pattern Matching with LAYER_PATTERNS
At the heart of this stage lies the LAYER_PATTERNS array defined in packages/core/src/analyzer/layer-detector.ts. This constant maps directory keywords to layer names and descriptions, ordered by priority. The matchFileToLayer function (lines 105-144) normalizes each file path to lowercase, splits it into segments, and checks every segment—including plural forms—against the pattern table. The first matching pattern determines the layer assignment.
Graph Traversal and Default Assignment
The detectLayers function iterates over all GraphNode objects of type "file" in the KnowledgeGraph. For each node containing a valid filePath, it invokes matchFileToLayer. Files lacking a path or failing to match any pattern are automatically assigned to the "Core" layer. For each distinct layer name, the system constructs a Layer object (defined in packages/core/src/types.ts lines 63-68) with a kebab-case ID formatted as layer:<name>.
Stage 2: LLM-Driven Layer Refinement
When projects follow non-standard directory structures or require domain-specific terminology, the engine falls back to an LLM for nuanced classification.
Prompt Engineering for Layer Detection
The buildLayerDetectionPrompt function (lines 150-284) aggregates every file path from the knowledge graph into a formatted bullet list. It appends instructions requesting a JSON array containing 3-7 layer objects, each specifying a name, description, and array of filePatterns (path prefixes). This prompt structure ensures the LLM returns machine-parseable architectural recommendations tailored to the specific codebase.
Parsing and Applying LLM Responses
After receiving the LLM output, parseLayerDetectionResponse strips Markdown fences and validates the JSON structure against the expected schema. The applyLLMLayers function then creates a mapping of the LLM-defined layers and traverses the graph a second time, assigning each file to the first layer whose filePatterns prefix-match the normalized path. This stage can overwrite heuristic assignments, allowing the system to surface abstractions like "Authentication Layer" or "Event Sourcing Layer" that directory names alone cannot express.
Implementation Examples
Running Heuristic Detection
import { detectLayers } from './packages/core/src/analyzer/layer-detector.js';
import type { KnowledgeGraph } from './packages/core/src/types.js';
// Assume `graph` is the knowledge graph produced by the core analyzer
const heuristicLayers = detectLayers(graph);
console.log('Heuristic layers →', heuristicLayers);
This returns an array of Layer objects, each containing a generated ID (layer:api-layer), human-readable name, description, and the list of associated node IDs.
Executing the LLM Refinement Workflow
import {
buildLayerDetectionPrompt,
parseLayerDetectionResponse,
applyLLMLayers,
} from './packages/core/src/analyzer/layer-detector.js';
import type { KnowledgeGraph } from './packages/core/src/types.js';
// 1️⃣ Build prompt
const prompt = buildLayerDetectionPrompt(graph);
// 2️⃣ Send to your LLM provider
const llmResponse = await myLLM.call(prompt);
// 3️⃣ Parse the JSON response
const llmLayers = parseLayerDetectionResponse(llmResponse);
if (!llmLayers) throw new Error('Failed to parse LLM output');
// 4️⃣ Apply refined layers
const refinedLayers = applyLLMLayers(graph, llmLayers);
console.log('LLM-refined layers →', refinedLayers);
The refinedLayers collection reflects the LLM's domain-aware groupings, with any unmatched files automatically routed to the "Other" layer.
Summary
- Two-stage pipeline: Architecture layer identification begins with fast heuristics and optionally refines results using LLM intelligence.
- Pattern-based matching: The
LAYER_PATTERNStable inlayer-detector.tsdrives initial classification by scanning directory segments. - Default safety nets: Unmatched files default to "Core" (heuristic) or "Other" (LLM) layers to ensure complete graph coverage.
- Type-safe structures: The
Layerinterface intypes.tsstandardizes layer objects across the knowledge graph. - Extensible design: The LLM stage accepts arbitrary
filePatterns, enabling project-specific architectural vocabularies beyond conventional folder names.
Frequently Asked Questions
What happens if a file path matches multiple heuristic patterns?
The matchFileToLayer function implements a "first match wins" strategy. It evaluates directory segments in path order against the LAYER_PATTERNS array, assigning the layer associated with the first matching keyword. This deterministic approach ensures consistent classification, with earlier patterns in the array taking precedence over later ones.
Can the heuristic and LLM stages be used independently?
Yes. The detectLayers function operates autonomously without API dependencies, making it suitable for offline environments or rapid initial analysis. The LLM functions (buildLayerDetectionPrompt, parseLayerDetectionResponse, applyLLMLayers) can be invoked separately when nuanced, project-specific layering is required, or skipped entirely if heuristic classification suffices.
Where are the default layer patterns defined?
The default pattern mappings reside in the LAYER_PATTERNS constant at the top of packages/core/src/analyzer/layer-detector.ts. This array pairs directory keywords (e.g., routes, service, model) with layer metadata, and developers can modify this table to adjust heuristic behavior before compilation.
How does the system handle files without a filePath property?
During the detectLayers traversal, any GraphNode of type "file" lacking a filePath property is automatically forced into the "Core" layer. This safeguards the graph construction process, ensuring that every file node receives a valid layer assignment even when path information is incomplete or corrupted.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →