How the Guided Tour Builder Generates Sequential Learning Paths in Understand-Anything
The guided tour builder creates educational step-by-step sequences by constructing an LLM prompt from the knowledge graph, parsing the AI response into validated tour steps, and falling back to a deterministic topological sort algorithm when needed.
The guided tour builder is a core component of the Egonex-AI/Understand-Anything platform that transforms static codebase analysis into interactive learning experiences. By combining large language model capabilities with graph theory algorithms, the system automatically produces dependency-ordered paths that help developers onboard to unfamiliar repositories. This process is implemented in the tour-generator.ts module within the core analyzer package.
Three-Stage Pipeline Architecture
The tour generation process follows a resilient three-stage pipeline designed to maximize success rates while maintaining educational coherence. Each stage handles a specific responsibility: prompt construction, response validation, and algorithmic fallback.
Stage 1: Prompt Construction with Graph Context
When a tour is requested, the buildTourGenerationPrompt function (located at understand-anything-plugin/packages/core/src/analyzer/tour-generator.ts, lines 7–45) assembles a comprehensive LLM prompt containing:
- Project metadata: Name, description, languages, and frameworks extracted from the knowledge graph's project section
- Node inventory: A compact list of all nodes including type, name, file path, and summary
- Dependency edges: Up to 50 edges describing relationships between components
- Architectural layers: Any detected layers with their names, descriptions, and contained node IDs
The function returns a plain-text prompt that explicitly instructs the LLM to "Generate a guided tour … Return a JSON object with a steps array …". This structured approach ensures the AI has sufficient context to create pedagogically sound sequences while respecting the codebase's actual architecture.
Stage 2: LLM Response Parsing and Validation
Raw LLM outputs often include markdown code fences or formatting artifacts. The parseTourGenerationResponse function (lines 66–100) handles extraction by:
- Stripping markdown code fences if present
- Parsing the resulting JSON
- Validating that each step contains required fields:
order,title,description, andnodeIds - Discarding malformed entries that fail validation
- Returning an array of
TourStepobjects as defined intypes.ts
This validation layer ensures that only properly structured steps reach the user, preventing partial or corrupted tours from disrupting the learning experience.
Stage 3: Deterministic Heuristic Fallback
If the LLM fails to produce a valid tour or returns empty results, the system falls back to generateHeuristicTour (lines 122–190), a deterministic algorithm that requires no external AI services:
- Node separation: Distinguishes concept nodes (documentation, README files) from code nodes (implementations)
- Graph analysis: Builds an adjacency map from the graph's edges and computes an in-degree table for all nodes
- Topological ordering: Performs a Kahn topological sort to obtain a dependency-ordered list of code nodes, ensuring learners encounter prerequisites before dependent components
- Step grouping: If architectural layers are present, groups nodes by layer according to the topological order; otherwise batches nodes three per step
- Concept integration: Adds a final "Key Concepts" step for any concept nodes identified earlier
- Sequential numbering: Assigns sequential
ordernumbers to all steps
This fallback guarantees that every analyzed repository produces a coherent tour, regardless of LLM availability or response quality.
Rendering Tours for End Users
Once generated, the TourStep[] array is attached to the KnowledgeGraph.tour field. Downstream consumers process this data through buildOnboardingGuide in understand-anything-plugin/src/onboard-builder.ts (lines 61–89). This renderer iterates over the tour array to produce markdown sections, displaying the files to examine and optional language tips for each step, creating a readable walkthrough suitable for documentation or IDE integration.
Implementation Example
import {
buildTourGenerationPrompt,
parseTourGenerationResponse,
generateHeuristicTour,
} from "@understand-anything/core";
// Assume `graph` is the KnowledgeGraph already built by the analyzer.
const prompt = buildTourGenerationPrompt(graph);
// Send `prompt` to the configured LLM (e.g., Claude, GPT‑4) and receive `rawResponse`.
const rawResponse = await llm.complete(prompt);
// Try to parse the LLM output.
let tour = parseTourGenerationResponse(rawResponse);
// If parsing failed, fall back to the heuristic generator.
if (tour.length === 0) {
tour = generateHeuristicTour(graph);
}
// Attach the tour to the graph for downstream rendering.
graph.tour = tour;
Summary
- The guided tour builder in Understand-Anything uses a hybrid approach combining LLM intelligence with graph algorithms to generate learning sequences.
- Prompt construction gathers project metadata, nodes, edges (up to 50), and architectural layers into a structured LLM request via
buildTourGenerationPrompt. - Response validation ensures data integrity through
parseTourGenerationResponse, which extracts JSON and verifies required fields likeorder,title,description, andnodeIds. - Heuristic fallback employs Kahn's topological sort algorithm in
generateHeuristicTourto create dependency-ordered steps when AI generation fails. - Tour structure is defined by the
TourStepinterface intypes.tsand rendered throughbuildOnboardingGuideinto markdown documentation.
Frequently Asked Questions
What happens if the LLM fails to generate a valid tour?
The system automatically falls back to generateHeuristicTour, which uses deterministic graph algorithms to create a valid learning sequence. This fallback separates concept nodes from code nodes, performs a topological sort based on dependency edges, and groups results into sequential steps, ensuring every repository produces a usable tour regardless of LLM response quality.
How does the heuristic fallback determine the order of steps?
The heuristic algorithm calculates an in-degree table from the graph's adjacency map and executes a Kahn topological sort. This produces a dependency-ordered list where prerequisite files appear before dependent ones. If the codebase contains detected architectural layers, nodes are grouped by layer; otherwise, the system batches three nodes per step to maintain digestible learning chunks.
What data structure represents a tour step in the codebase?
According to understand-anything-plugin/packages/core/src/types.ts, each tour step follows the TourStep interface containing: order (number), title (string), description (string), and nodeIds (array of strings referencing knowledge graph nodes). This structure is validated during the parsing stage and consumed by the onboarding guide renderer.
How are concept nodes handled differently from code nodes?
During heuristic generation, concept nodes (such as README files and documentation) are separated from implementation code nodes. While code nodes are ordered via topological sort based on dependencies, concept nodes are collected and presented in a final "Key Concepts" step. This ensures learners understand the codebase architecture before diving into implementation details, or review documentation after understanding the code structure.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →