What Is the SCAN Phase and What Does the Project-Scanner Do in Understand-Anything?

The SCAN phase is the first step of the /understand pipeline where the project-scanner agent recursively discovers files, detects programming languages, and builds an internal import map, writing the results to scan-result.json for downstream agents.

In the Lum1104/Understand-Anything repository, the SCAN phase serves as the entry point for the /understand skill pipeline. During this phase, the project-scanner agent inventories the entire target codebase, classifies each source file by language, and constructs a project-internal dependency graph. The resulting artifact, scan-result.json, provides the foundational data that later agents rely on to perform deep code analysis.

How the SCAN Phase Fits into the /understand Pipeline

The SCAN phase is explicitly defined as Phase 1 of the /understand pipeline. According to the pipeline diagram in docs/superpowers/specs/2026-05-24-semantic-batching-and-output-chunking-design.md (lines 42–50), the project-scanner is invoked before any batching or analysis occurs. The skill orchestrator defined in skills/understand/SKILL.md dispatches the scanner as a sub-agent at line 243, ensuring the repository structure is fully mapped before downstream processing begins.

Core Responsibilities of the Project-Scanner Agent

The agent's duties are documented in agents/project-scanner.md and can be grouped into four concrete tasks.

File Discovery and Ignore Patterns

The scanner recursively enumerates every file in the target directory while applying a hard-coded set of ignore patterns. As specified in agents/project-scanner.md (lines 2–20), paths matching node_modules/, .git/, dist/, build/, bin/, obj/, *.lock, or *.min.js are automatically excluded from the scan result.

Language and Framework Detection

For each discovered file, the agent runs a lightweight language detector that uses file extensions and simple heuristics to tag the file with its programming language and any applicable framework. This metadata helps later agents select the correct parsing strategy.

Import Resolution and Map Generation

By parsing import and require statements inside each source file, the scanner resolves project-internal references to absolute paths. External package imports are deliberately discarded. The final structure is an importMap object where every key is a file path and its value is an array of the project files it directly depends on. The merge logic for this map is described in agents/project-scanner.md at line 156.

Output Contract: scan-result.json

The SCAN phase emits a single JSON artifact with two top-level keys: files and importMap. The files array lists all discovered paths, while importMap encodes the internal dependency graph. This contract is consumed immediately by the file-analyzer, which uses the import map to create graph edges. The README table at line 291 confirms the project-scanner is responsible for discovering files, detecting languages, and building the import map.

{
  "files": [
    "src/index.ts",
    "src/utils/helpers.ts",
    "package.json"
  ],
  "importMap": {
    "src/index.ts": ["src/utils/helpers.ts"],
    "src/utils/helpers.ts": []
  }
}

Triggering the SCAN Phase from the /understand Skill

The pipeline does not run the scanner manually; it is dispatched as a sub-agent. The orchestration code in skills/understand/SKILL.md follows this pattern:

await dispatchSubagent({
  name: "project-scanner",
  prompt: await readFile("./agents/project-scanner.md"),
  // The sub-agent returns scan-result.json
});

This sub-agent invocation returns scan-result.json, which the skill later merges into the final scan result as noted at line 156 of agents/project-scanner.md.

Summary

  • The SCAN phase is Phase 1 of the /understand pipeline in the Lum1104/Understand-Anything repository.
  • The project-scanner is dispatched from skills/understand/SKILL.md (line 243) as a sub-agent.
  • It discovers files recursively while ignoring common build, dependency, and lock-file patterns defined in agents/project-scanner.md.
  • It detects languages via extension-based heuristics and builds an importMap of purely internal dependencies.
  • The final artifact, scan-result.json, contains a files list and an importMap consumed by downstream agents such as the file-analyzer.

Frequently Asked Questions

What files does the project-scanner ignore during the SCAN phase?

The scanner skips paths matching hard-coded patterns including node_modules/, .git/, dist/, build/, bin/, obj/, *.lock, and *.min.js. These defaults are defined in agents/project-scanner.md (lines 2–20).

Does the SCAN phase include external package dependencies in the import map?

No. The importMap produced by the project-scanner deliberately omits external packages and only records project-internal imports. External references are discarded during import resolution to keep the dependency graph focused on first-party code.

Which downstream agents consume the scan-result.json output?

The file-analyzer and architecture-analyzer are the primary consumers. The file-analyzer uses the importMap to construct graph edges, while later batch-computation stages rely on the files list and dependency map to schedule analysis.

How is the project-scanner triggered inside the /understand pipeline?

The /understand skill triggers the scanner by calling dispatchSubagent() with the agent name "project-scanner" and the prompt loaded from ./agents/project-scanner.md, as shown at line 243 of skills/understand/SKILL.md. The sub-agent executes autonomously and returns scan-result.json to the orchestrator.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →