Data Generation Process for the AI Engineering Website: How `site/build.js` Works

The site/build.js script in the rohitg00/ai-engineering-from-scratch repository is a Node.js build pipeline that transforms human-editable curriculum sources into a machine-readable JavaScript module consumed by the static site.

This script serves as the single source of truth for the AI Engineering from Scratch curriculum, parsing Markdown files, discovering reusable artifacts, and serializing everything into site/data.js. The process ensures the website stays synchronized with every lesson, phase status, and glossary term defined in the repository.

Overview of the Build Pipeline

The build process orchestrates twelve distinct stages, from parsing source documents to generating SEO-ready auxiliary files. When executed via node site/build.js (typically triggered by GitHub Actions on every push), the script performs a complete synthesis of the curriculum structure.

The pipeline begins by defining absolute paths to key source files. According to the source code in Lines 14-21, the script initializes constants pointing to README.md, ROADMAP.md, glossary/terms.md, and the output location data.js, along with the GitHub base URL used for constructing absolute lesson links.

Parsing Source Documents

The script extracts structured data from three primary Markdown sources that serve as the curriculum's ground truth.

Processing ROADMAP.md for Status Tracking

The parseRoadmap function (Lines 30-61) scans the roadmap file to capture phase-level statuses and per-lesson completion emojis (✅, 🚧, ⬚). This produces a roadmapStatuses map that correlates lesson identifiers with their current implementation state, distinguishing between complete, in-progress, and planned lessons.

Extracting Lesson Metadata from README.md

The parseReadme function (Lines 63-132) walks through the main README to locate phase headers and lesson tables. For each lesson encountered, it extracts:

  • Lesson name and type (Build, Theory, Capstone)
  • Language list (Python, Node, Rust, etc.)
  • GitHub URL for the lesson directory (when present)
  • Cross-referenced status from the roadmap map

This function creates the foundational PHASES array that structures the entire curriculum navigation.

Building the Glossary Index

The parseGlossary function (Lines 71-103) reads glossary/terms.md and extracts terminology definitions. For each term, it identifies the what people say and what it actually means sections, constructing an array of term objects used for the website's searchable glossary feature.

Discovering Reusable Artifacts

A critical stage involves scanning the filesystem for curriculum artifacts using the discoverArtifacts function (Lines 112-199).

This function recursively traverses every phases/*/*/outputs directory, searching for Markdown files following specific naming conventions:

  • skill-*.md – Reusable skill implementations
  • prompt-*.md – Standardized prompt templates
  • prompt-*.md – Agent configurations

For each artifact found, the script extracts front-matter metadata, tags, and parent lesson references, yielding a flat ARTIFACTS array. This enables the UI to present reusable components alongside their originating lessons.

Enriching Lesson Content

After establishing the basic structure, the build process enhances lesson objects with content-derived metadata. Within the build() function, a dedicated loop (Lines 45-55) processes each lesson that has a valid URL:

  1. Reads the lesson's docs/en.md file
  2. Extracts the first blockquote as a one-line summary
  3. Parses all ### headings to compile keywords

This enrichment allows the website to display lesson previews and tag clouds without requiring runtime Markdown processing.

Generating Output Files

Serializing site/data.js

The core output occurs in Lines 74-84, where the script serializes three major data structures—PHASES, GLOSSARY, and ARTIFACTS—into a single JavaScript file. The generated site/data.js includes:

  • A header comment containing the build timestamp
  • The complete curriculum hierarchy with enriched metadata
  • Searchable glossary definitions
  • Reusable artifact registry
// Auto-generated by build.js — do not edit manually.
// Last built: 2026-07-30T12:34:56.789Z

const PHASES = [
  {
    "id": 0,
    "name": "Setup & Tooling",
    "status": "complete",
    "desc": "Get your environment ready …",
    "lessons": [
      {
        "name": "Dev Environment",
        "status": "complete",
        "type": "Build",
        "lang": "Python, Node, Rust",
        "url": "https://github.com/rohitg00/ai-engineering-from-scratch/tree/main/phases/00-setup-and-tooling/01-dev-environment/",
        "summary": "Configure a reproducible dev stack …",
        "keywords": "Virtualenv · Pip · Cargo · Rustup"
      }
    ]
  }
];

const GLOSSARY = [
  { "term":"Agent", "says":"A software entity that …", "means":"A program that …" }
];

const ARTIFACTS = [
  { "kind":"skill", "name":"Skill‑end‑to‑end‑safety‑gate", "description":"…", "tags":["safety"], "phase":19, "lesson":87 }
];

Synchronizing Auxiliary Assets

The build pipeline updates several auxiliary files to maintain SEO and documentation consistency (Lines 90-108):

  • README badges – Updates completion counts and lesson statistics
  • Static HTML pages – Rewrites count placeholders in template files
  • sitemap.xml – Generates search engine sitemap entries for every lesson URL
<url>
  <loc>https://aiengineeringfromscratch.com/lesson.html?path=phases/00-setup-and-tooling/01-dev-environment</loc>
  <lastmod>2026-07-30</lastmod>
  <changefreq>monthly</changefreq>
  <priority>0.6</priority>
</url>

Creating Machine-Readable Indexes

The writeLlms function (Lines 118-146) produces site/llms.txt, a human-readable index of all lessons, glossary terms, and artifacts formatted for AI agent consumption.

Build Metadata and Versioning

The script captures deployment context through the resolveRef function (Lines 100-122), which determines the current git ref (preferring the VERCEL_GIT_COMMIT_REF environment variable when available). This ref is written to build-meta.js, enabling the client to fetch raw Markdown files from the correct branch or tag.

Execution and Automation

The entry point (Lines 606-608) invokes the build() function when the script runs directly. In production, GitHub Actions triggers this automatically on every push, ensuring the website reflects the latest curriculum state without manual intervention.

Summary

  • site/build.js serves as the central build orchestrator for the AI Engineering from Scratch curriculum, transforming Markdown sources into structured data.
  • The pipeline parses ROADMAP.md for status emojis, README.md for lesson metadata, and glossary/terms.md for terminology definitions.
  • Artifact discovery recursively scans phases/*/*/outputs directories to catalog reusable skills, prompts, and agents.
  • The script enriches lesson objects by extracting summaries and keywords from individual docs/en.md files.
  • Output is serialized to site/data.js, accompanied by auxiliary files including sitemap.xml, llms.txt, and updated README badges.
  • Build metadata captures the git ref via VERCEL_GIT_COMMIT_REF to ensure version-aligned content fetching.

Frequently Asked Questions

What is the role of site/data.js in the website architecture?

site/data.js is the runtime bundle consumed by the client-side application. It contains the serialized PHASES, GLOSSARY, and ARTIFACTS arrays generated by build.js, enabling static rendering of lesson navigation, search functionality, and glossary lookups without requiring server-side Markdown parsing.

How does the build process track lesson completion status?

The parseRoadmap function scans ROADMAP.md for Unicode emojis (✅ for complete, 🚧 for in-progress, ⬚ for planned) and maps these to standardized status strings. When parseReadme processes lesson tables, it cross-references these mappings to assign each lesson its current implementation state.

Where does the website get its lesson summaries and keywords?

During the build loop (Lines 45-55), the script reads each lesson's docs/en.md file and extracts the first blockquote (treated as the summary) and all H3 headings (treated as keywords). This content is then embedded directly into the lesson objects within data.js.

What is llms.txt and why is it generated?

llms.txt is a machine-readable curriculum index created by the writeLlms function. It provides AI agents and large language models with a structured, link-rich overview of all lessons, glossary terms, and reusable artifacts, facilitating automated comprehension of the curriculum structure without requiring full repository traversal.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →