# Architecture Layer Identification in Understand Anything: A Two-Stage Pipeline Explained

> Discover architecture layer identification in Understand Anything. Learn how a two-stage pipeline uses heuristics and LLMs to categorize project files for better understanding.

- Repository: [Yuxiang Lin/Understand-Anything](https://github.com/Lum1104/Understand-Anything)
- Tags: architecture
- Published: 2026-06-02

---

**Architecture layer identification in Understand Anything assigns every file node to a logical layer through an initial heuristic scan of directory paths followed by an optional LLM-driven refinement that surfaces project-specific architectural groupings.**

The Understand Anything engine transforms raw source code into a navigable knowledge graph where **architecture layer identification** plays a central role. By automatically categorizing files into logical layers such as "API" or "Service," the system creates a high-level architectural view that helps developers comprehend large codebases without manual documentation. This process is implemented in the [`layer-detector.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/layer-detector.ts) module and operates in two distinct stages.

## The Two-Stage Identification Pipeline

The system employs a deterministic-to-intelligent pipeline that balances speed with accuracy:

1. **Heuristic Detection** — Matches directory names against curated patterns in `LAYER_PATTERNS` for instant, rule-based classification
2. **LLM Refinement** — Generates prompts for a language model to propose domain-specific layers based on the complete file structure

Files that fail heuristic matching default to the **"Core"** layer, while files unmatched by the LLM stage land in **"Other"**.

## Stage 1: Heuristic Layer Detection

The heuristic stage provides immediate, deterministic classification by analyzing file paths without external API calls.

### Pattern Matching with LAYER_PATTERNS

At the heart of this stage lies the `LAYER_PATTERNS` array defined in [`packages/core/src/analyzer/layer-detector.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/core/src/analyzer/layer-detector.ts). This constant maps directory keywords to layer names and descriptions, ordered by priority. The `matchFileToLayer` function (lines 105-144) normalizes each file path to lowercase, splits it into segments, and checks every segment—including plural forms—against the pattern table. The first matching pattern determines the layer assignment.

### Graph Traversal and Default Assignment

The `detectLayers` function iterates over all `GraphNode` objects of type `"file"` in the `KnowledgeGraph`. For each node containing a valid `filePath`, it invokes `matchFileToLayer`. Files lacking a path or failing to match any pattern are automatically assigned to the `"Core"` layer. For each distinct layer name, the system constructs a `Layer` object (defined in [`packages/core/src/types.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/core/src/types.ts) lines 63-68) with a kebab-case ID formatted as `layer:<name>`.

## Stage 2: LLM-Driven Layer Refinement

When projects follow non-standard directory structures or require domain-specific terminology, the engine falls back to an LLM for nuanced classification.

### Prompt Engineering for Layer Detection

The `buildLayerDetectionPrompt` function (lines 150-284) aggregates every file path from the knowledge graph into a formatted bullet list. It appends instructions requesting a JSON array containing 3-7 layer objects, each specifying a `name`, `description`, and array of `filePatterns` (path prefixes). This prompt structure ensures the LLM returns machine-parseable architectural recommendations tailored to the specific codebase.

### Parsing and Applying LLM Responses

After receiving the LLM output, `parseLayerDetectionResponse` strips Markdown fences and validates the JSON structure against the expected schema. The `applyLLMLayers` function then creates a mapping of the LLM-defined layers and traverses the graph a second time, assigning each file to the first layer whose `filePatterns` prefix-match the normalized path. This stage can overwrite heuristic assignments, allowing the system to surface abstractions like "Authentication Layer" or "Event Sourcing Layer" that directory names alone cannot express.

## Implementation Examples

### Running Heuristic Detection

```typescript
import { detectLayers } from './packages/core/src/analyzer/layer-detector.js';
import type { KnowledgeGraph } from './packages/core/src/types.js';

// Assume `graph` is the knowledge graph produced by the core analyzer
const heuristicLayers = detectLayers(graph);

console.log('Heuristic layers →', heuristicLayers);

```

This returns an array of `Layer` objects, each containing a generated ID (`layer:api-layer`), human-readable name, description, and the list of associated node IDs.

### Executing the LLM Refinement Workflow

```typescript
import {
  buildLayerDetectionPrompt,
  parseLayerDetectionResponse,
  applyLLMLayers,
} from './packages/core/src/analyzer/layer-detector.js';
import type { KnowledgeGraph } from './packages/core/src/types.js';

// 1️⃣ Build prompt
const prompt = buildLayerDetectionPrompt(graph);

// 2️⃣ Send to your LLM provider
const llmResponse = await myLLM.call(prompt);

// 3️⃣ Parse the JSON response
const llmLayers = parseLayerDetectionResponse(llmResponse);
if (!llmLayers) throw new Error('Failed to parse LLM output');

// 4️⃣ Apply refined layers
const refinedLayers = applyLLMLayers(graph, llmLayers);
console.log('LLM-refined layers →', refinedLayers);

```

The `refinedLayers` collection reflects the LLM's domain-aware groupings, with any unmatched files automatically routed to the **"Other"** layer.

## Summary

- **Two-stage pipeline**: Architecture layer identification begins with fast heuristics and optionally refines results using LLM intelligence.
- **Pattern-based matching**: The `LAYER_PATTERNS` table in [`layer-detector.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/layer-detector.ts) drives initial classification by scanning directory segments.
- **Default safety nets**: Unmatched files default to **"Core"** (heuristic) or **"Other"** (LLM) layers to ensure complete graph coverage.
- **Type-safe structures**: The `Layer` interface in [`types.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/types.ts) standardizes layer objects across the knowledge graph.
- **Extensible design**: The LLM stage accepts arbitrary `filePatterns`, enabling project-specific architectural vocabularies beyond conventional folder names.

## Frequently Asked Questions

### What happens if a file path matches multiple heuristic patterns?

The `matchFileToLayer` function implements a "first match wins" strategy. It evaluates directory segments in path order against the `LAYER_PATTERNS` array, assigning the layer associated with the first matching keyword. This deterministic approach ensures consistent classification, with earlier patterns in the array taking precedence over later ones.

### Can the heuristic and LLM stages be used independently?

Yes. The `detectLayers` function operates autonomously without API dependencies, making it suitable for offline environments or rapid initial analysis. The LLM functions (`buildLayerDetectionPrompt`, `parseLayerDetectionResponse`, `applyLLMLayers`) can be invoked separately when nuanced, project-specific layering is required, or skipped entirely if heuristic classification suffices.

### Where are the default layer patterns defined?

The default pattern mappings reside in the `LAYER_PATTERNS` constant at the top of [`packages/core/src/analyzer/layer-detector.ts`](https://github.com/Lum1104/Understand-Anything/blob/main/packages/core/src/analyzer/layer-detector.ts). This array pairs directory keywords (e.g., `routes`, `service`, `model`) with layer metadata, and developers can modify this table to adjust heuristic behavior before compilation.

### How does the system handle files without a filePath property?

During the `detectLayers` traversal, any `GraphNode` of type `"file"` lacking a `filePath` property is automatically forced into the **"Core"** layer. This safeguards the graph construction process, ensuring that every file node receives a valid layer assignment even when path information is incomplete or corrupted.