Architecture Layer Identification in Understand Anything: A Two-Stage Pipeline Explained

Architecture layer identification in Understand Anything assigns every file node to a logical layer through an initial heuristic scan of directory paths followed by an optional LLM-driven refinement that surfaces project-specific architectural groupings.

The Understand Anything engine transforms raw source code into a navigable knowledge graph where architecture layer identification plays a central role. By automatically categorizing files into logical layers such as "API" or "Service," the system creates a high-level architectural view that helps developers comprehend large codebases without manual documentation. This process is implemented in the layer-detector.ts module and operates in two distinct stages.

The Two-Stage Identification Pipeline

The system employs a deterministic-to-intelligent pipeline that balances speed with accuracy:

  1. Heuristic Detection — Matches directory names against curated patterns in LAYER_PATTERNS for instant, rule-based classification
  2. LLM Refinement — Generates prompts for a language model to propose domain-specific layers based on the complete file structure

Files that fail heuristic matching default to the "Core" layer, while files unmatched by the LLM stage land in "Other".

Stage 1: Heuristic Layer Detection

The heuristic stage provides immediate, deterministic classification by analyzing file paths without external API calls.

Pattern Matching with LAYER_PATTERNS

At the heart of this stage lies the LAYER_PATTERNS array defined in packages/core/src/analyzer/layer-detector.ts. This constant maps directory keywords to layer names and descriptions, ordered by priority. The matchFileToLayer function (lines 105-144) normalizes each file path to lowercase, splits it into segments, and checks every segment—including plural forms—against the pattern table. The first matching pattern determines the layer assignment.

Graph Traversal and Default Assignment

The detectLayers function iterates over all GraphNode objects of type "file" in the KnowledgeGraph. For each node containing a valid filePath, it invokes matchFileToLayer. Files lacking a path or failing to match any pattern are automatically assigned to the "Core" layer. For each distinct layer name, the system constructs a Layer object (defined in packages/core/src/types.ts lines 63-68) with a kebab-case ID formatted as layer:<name>.

Stage 2: LLM-Driven Layer Refinement

When projects follow non-standard directory structures or require domain-specific terminology, the engine falls back to an LLM for nuanced classification.

Prompt Engineering for Layer Detection

The buildLayerDetectionPrompt function (lines 150-284) aggregates every file path from the knowledge graph into a formatted bullet list. It appends instructions requesting a JSON array containing 3-7 layer objects, each specifying a name, description, and array of filePatterns (path prefixes). This prompt structure ensures the LLM returns machine-parseable architectural recommendations tailored to the specific codebase.

Parsing and Applying LLM Responses

After receiving the LLM output, parseLayerDetectionResponse strips Markdown fences and validates the JSON structure against the expected schema. The applyLLMLayers function then creates a mapping of the LLM-defined layers and traverses the graph a second time, assigning each file to the first layer whose filePatterns prefix-match the normalized path. This stage can overwrite heuristic assignments, allowing the system to surface abstractions like "Authentication Layer" or "Event Sourcing Layer" that directory names alone cannot express.

Implementation Examples

Running Heuristic Detection

import { detectLayers } from './packages/core/src/analyzer/layer-detector.js';
import type { KnowledgeGraph } from './packages/core/src/types.js';

// Assume `graph` is the knowledge graph produced by the core analyzer
const heuristicLayers = detectLayers(graph);

console.log('Heuristic layers →', heuristicLayers);

This returns an array of Layer objects, each containing a generated ID (layer:api-layer), human-readable name, description, and the list of associated node IDs.

Executing the LLM Refinement Workflow

import {
  buildLayerDetectionPrompt,
  parseLayerDetectionResponse,
  applyLLMLayers,
} from './packages/core/src/analyzer/layer-detector.js';
import type { KnowledgeGraph } from './packages/core/src/types.js';

// 1️⃣ Build prompt
const prompt = buildLayerDetectionPrompt(graph);

// 2️⃣ Send to your LLM provider
const llmResponse = await myLLM.call(prompt);

// 3️⃣ Parse the JSON response
const llmLayers = parseLayerDetectionResponse(llmResponse);
if (!llmLayers) throw new Error('Failed to parse LLM output');

// 4️⃣ Apply refined layers
const refinedLayers = applyLLMLayers(graph, llmLayers);
console.log('LLM-refined layers →', refinedLayers);

The refinedLayers collection reflects the LLM's domain-aware groupings, with any unmatched files automatically routed to the "Other" layer.

Summary

  • Two-stage pipeline: Architecture layer identification begins with fast heuristics and optionally refines results using LLM intelligence.
  • Pattern-based matching: The LAYER_PATTERNS table in layer-detector.ts drives initial classification by scanning directory segments.
  • Default safety nets: Unmatched files default to "Core" (heuristic) or "Other" (LLM) layers to ensure complete graph coverage.
  • Type-safe structures: The Layer interface in types.ts standardizes layer objects across the knowledge graph.
  • Extensible design: The LLM stage accepts arbitrary filePatterns, enabling project-specific architectural vocabularies beyond conventional folder names.

Frequently Asked Questions

What happens if a file path matches multiple heuristic patterns?

The matchFileToLayer function implements a "first match wins" strategy. It evaluates directory segments in path order against the LAYER_PATTERNS array, assigning the layer associated with the first matching keyword. This deterministic approach ensures consistent classification, with earlier patterns in the array taking precedence over later ones.

Can the heuristic and LLM stages be used independently?

Yes. The detectLayers function operates autonomously without API dependencies, making it suitable for offline environments or rapid initial analysis. The LLM functions (buildLayerDetectionPrompt, parseLayerDetectionResponse, applyLLMLayers) can be invoked separately when nuanced, project-specific layering is required, or skipped entirely if heuristic classification suffices.

Where are the default layer patterns defined?

The default pattern mappings reside in the LAYER_PATTERNS constant at the top of packages/core/src/analyzer/layer-detector.ts. This array pairs directory keywords (e.g., routes, service, model) with layer metadata, and developers can modify this table to adjust heuristic behavior before compilation.

How does the system handle files without a filePath property?

During the detectLayers traversal, any GraphNode of type "file" lacking a filePath property is automatically forced into the "Core" layer. This safeguards the graph construction process, ensuring that every file node receives a valid layer assignment even when path information is incomplete or corrupted.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →