How PAGETOPOLOGY.md Is Generated and Used in the AI Website Cloner Template
PAGETOPOLOGY.md is generated at runtime by the /clone-website skill during the Reconnaissance phase, capturing the DOM hierarchy of crawled pages to serve as a blueprint for builder agents and a human-readable reference.
The JCodesMore/ai-website-cloner-template repository implements a multi-agent pipeline for cloning websites. At the heart of this system lies PAGETOPOLOGY.md, a dynamically generated markdown document that bridges the gap between site reconnaissance and component generation.
Generation During the Reconnaissance Phase
The creation of PAGETOPOLOGY.md occurs during the Reconnaissance phase of the cloning pipeline. As detailed in the repository's README.md (lines 86-93), the AI agent crawls the target website, inspects the DOM hierarchy of each page, and serializes the structural relationships into a hierarchical text format.
The /clone-website Skill Execution
The /clone-website skill executes runtime logic to transform live page data into structured documentation. Unlike static template files, this skill performs live crawling and DOM analysis, producing the topology document on-the-fly rather than shipping it with the template.
Output Location and Naming Convention
The markdown file is written to docs/research/${slug(url)}-PAGETOPOLOGY.md within the repository root. This places generated research artifacts in a dedicated folder separate from source code, following the directory structure defined in AGENTS.md.
Document Structure and Content
The generated file contains a hierarchical text representation of the page's structural layout. According to the source analysis, it specifically documents:
- Sections and container elements
- Component nesting relationships
- Navigation flow between pages
Consumption by Builder Agents
Once generated, PAGETOPOLOGY.md serves as the authoritative source for subsequent pipeline stages, fulfilling two distinct purposes.
Blueprint for Parallel Generation
Builder agents read the topology to determine which sections require separate component specifications and code generation. The document defines the assembly order and nesting relationships, enabling parallel execution across multiple agents while maintaining structural integrity.
Human-Readable Verification
Developers can inspect the markdown file to validate that the AI correctly interpreted the page layout before the final assembly step. This manual verification capability allows for corrections and adjustments without requiring deep dives into agent internals.
Implementation Details
The following TypeScript pseudo-code illustrates the generation process executed by the /clone-website skill:
// Pseudo-code executed by the /clone-website skill
import { writeFileSync } from "fs";
import { getPageDOM, serializeDOMTree } from "./dom-utils";
async function generatePageTopology(url: string) {
const dom = await getPageDOM(url); // fetch & render page
const topology = serializeDOMTree(dom.documentElement); // turn DOM into hierarchical text
const markdown = `# Page Topology for ${url}\n\n${topology}`;
writeFileSync(`docs/research/${slug(url)}-PAGETOPOLOGY.md`, markdown);
}
Builder agents later consume this file using standard filesystem operations:
import { readFileSync } from "fs";
function loadPageTopology(pageSlug: string) {
const path = `docs/research/${pageSlug}-PAGETOPOLOGY.md`;
const content = readFileSync(path, "utf-8");
// Parse the markdown outline to drive component generation…
}
Key Files Supporting the Topology Workflow
Several files in the JCodesMore/ai-website-cloner-template repository govern how PAGETOPOLOGY.md is created and utilized:
README.md– Documents the multi-phase cloning pipeline, including the Reconnaissance step that produces the topology markdown.AGENTS.md– Serves as the single source of truth for agent instructions, directing agents to generate documentation underdocs/research/.scripts/sync-agent-rules.sh– Regenerates platform-specific instruction files, indirectly controlling where generated docs are placed.docs/research/– The runtime destination folder for generated markdown files such asPAGETOPOLOGY.md.
Summary
PAGETOPOLOGY.mdis generated dynamically by the/clone-websiteskill during the Reconnaissance phase, not shipped as a static template file.- The document captures the DOM hierarchy, sections, containers, and navigation flow of crawled pages.
- Builder agents use the topology as a blueprint to coordinate parallel component generation and assembly.
- Human developers can verify the AI's structural understanding by inspecting the markdown file before final assembly.
- Output is stored in
docs/research/according to instructions inAGENTS.mdand the pipeline documented inREADME.md.
Frequently Asked Questions
Is PAGETOPOLOGY.md included in the GitHub repository?
No. PAGETOPOLOGY.md is created at runtime during the cloning process. The docs/research/ directory remains empty in the template repository, as the file is generated fresh for each website cloning operation.
What specific information does the topology document capture?
The document records the structural layout of each page, including section hierarchies, container nesting, component relationships, and navigation flow between pages. This allows builder agents to understand exactly how UI elements relate to each other spatially and functionally.
How do builder agents access the topology data?
Builder agents read the markdown file from docs/research/${pageSlug}-PAGETOPOLOGY.md using standard filesystem operations. They parse the hierarchical outline to determine which components to generate and how to assemble them into the final page structure.
Can developers modify the topology after it is generated?
Yes. The markdown format serves as a human-readable reference that developers can edit to correct the AI's interpretation of the page layout. Modifications made to PAGETOPOLOGY.md will influence how subsequent builder agents generate components, allowing for manual refinement of the structural blueprint.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →