How the Book Generation Pipeline Works in AI Engineering from Scratch

The scripts/build_book.py script transforms Markdown lesson files into professional EPUB and PDF books by scanning the phases/ directory, processing content through a transformation pipeline, and rendering the final artifacts via Pandoc.

The AI Engineering from Scratch repository uses a custom automation pipeline to compile scattered curriculum markdown into distributable book formats. The entry point, scripts/build_book.py, located at the repository root, orchestrates the entire workflow—from discovering lessons in phase directories to generating finished EPUB files with properly formatted metadata.

Configuration and Volume Definitions

The pipeline starts by reading book/volumes.json to understand the six-book series structure. This JSON file defines which phases belong to each volume (Foundations, Deep Learning, Language, etc.) and provides essential metadata like the site URL and repository path.

CONFIG = json.loads((ROOT / "book" / "volumes.json").read_text())
SITE   = CONFIG["site"].rstrip("/")
REPO   = CONFIG["repo"].rstrip("/")

As defined in scripts/build_book.py lines 28-32, the script establishes four pathlib shortcuts that drive the entire build process:

  • ROOT — Repository root directory
  • PHASES — The phases/ directory containing curriculum content
  • BUILD — Temporary build folder at book/_build
  • DIST — Final distribution folder at dist/book

Discovering and Indexing Lessons

For each volume defined in the configuration, the pipeline must locate the actual lesson content. The lesson_dirs() function (lines 45-53 in scripts/build_book.py) scans each phase directory using the LESSON_DIR_RE regular expression imported from scripts/build_catalog.py.

def lesson_dirs(phase):
    base = PHASES / phase
    ...
    return [d for d in sorted(base.iterdir())
            if d.is_dir() and LESSON_DIR_RE.match(d.name)
            and (d / "docs" / "en.md").is_file()]

This function returns only directories that match the lesson-slug convention and contain a docs/en.md file, ensuring the pipeline processes valid lesson structures while ignoring auxiliary directories.

Before transforming content, the pipeline prepares navigation links. The urls_for() function (lines 60-66) constructs three essential URLs for every lesson:

def urls_for(phase, lesson):
    rel = f"phases/{phase}/{lesson}"
    return {
        "web": f"{SITE}/lesson.html?path={rel}",
        "code": f"{REPO}/tree/main/{rel}/code",
        "repo": f"{REPO}/tree/main/{rel}",
    }

These URLs populate the "continue online" boxes in the final book, allowing readers to access the interactive web edition, view raw code in the repository, or navigate directly to the lesson root.

Transforming Lesson Content

The transform_lesson() function serves as the core processing engine, reading each docs/en.md file line-by-line and applying specialized transformations. Located in the main script body, this function handles six distinct content types:

Code Fences and Special Blocks

The parser detects fenced code blocks (lines 89-101) and treats special fence types differently. When encountering ```figure blocks, the pipeline replaces them with "continue-online" boxes linking to interactive web figures. Standard code fences pass through with syntax highlighting preserved.

Mermaid Diagram Rendering

For ```mermaid fences, the pipeline calls render_mermaid() (lines 22-30) to generate SVG diagrams using the mmdc CLI tool. If Mermaid rendering fails, the system falls back to a placeholder box rather than failing the build, ensuring robustness across different environments.

Ship It Sections and Exercises

The transformer collapses ## Ship It markdown sections into single boxed notes pointing to artifact repositories (lines 35-44). For exercise sections, it appends links to starter code (lines 53-57), maintaining the connection between reading material and hands-on implementation.

Asset Path Rewriting

Relative image links pointing to ../assets/ get rewritten to absolute paths within the built book structure (phases/<phase>/<lesson>/assets/...) on lines 60-62. This ensures images resolve correctly in the final EPUB regardless of the reader's device.

Validation and Error Handling

If transform_lesson() encounters an unclosed fenced block, it raises a ValueError (lines 63-65) to surface formatting errors immediately rather than generating broken output.

Assembling Complete Volumes

Once individual lessons are processed, the assemble() function (main body of scripts/build_book.py) constructs complete volume manuscripts:

def assemble(vol):
    BUILD.mkdir(parents=True, exist_ok=True)
    parts = [how_to_use(vol)]
    for part_idx, phase in enumerate(vol["phases"]):
        title = clean_phase_title(phase_title(phase))
        parts.append(f"\n# Part {ROMAN[part_idx]} — {title} {{.unnumbered .part}}\n\n"

                     f"*Course phase {phase.split('-')[0]}. … <{SITE}/catalog.html>*\n")
        for lesson_dir in lesson_dirs(phase):
            parts.append("\n".join(transform_lesson(phase, lesson_dir)))
    text = "\n\n".join(parts)
    md   = BUILD / f"{vol['slug']}.md"
    md.write_text(text)
    return md, chapters, len(text.split())

This function:

  1. Creates the build directory structure
  2. Prepends a "How to Use" preface explaining the volume scope
  3. Generates Part headers using Roman numerals (ROMAN array) for each phase
  4. Concatenates all transformed lessons with proper markdown separation
  5. Writes the final file to book/_build/<volume-slug>.md

Metadata Generation and Rendering

Before invoking Pandoc, the metadata() function (lines 96-106) writes a titlepage.yaml file containing the series title, volume subtitle, author information, language settings, and table-of-contents title.

EPUB Generation

The render() function executes Pandoc with specific arguments (lines 15-21):

  • Input: The assembled markdown file and metadata YAML
  • Styling: Custom CSS from book/epub.css
  • Options: Table of contents inclusion and proper chapter hierarchy

PDF Generation (Optional)

When the --pdf flag is active, the pipeline generates a LaTeX title page using book/titlepage.tex, selects appropriate system fonts via pick_font() and font_families() (lines 70-86), and invokes Pandoc with --pdf-engine=xelatex. The font selection logic queries fc-list to find suitable serif and monospace fonts available on the build machine.

Both artifacts are written to dist/book/ with size reporting to the console.

CLI Orchestration and Build Flags

The main() function (lines 87-105) provides a command-line interface supporting three critical flags:

Flag Function
--volume <slug> Build only a specific volume instead of the entire series
--pdf Enable PDF generation (requires xelatex installation)
--assemble-only Stop after markdown generation, skipping Pandoc rendering

Before building, the script validates phase existence using check_phases() to ensure every referenced phase directory actually contains lessons, preventing mid-build failures.

Practical Usage Examples

Build the Entire Series (EPUB only)

python3 scripts/build_book.py

This scans book/volumes.json, assembles all six volumes, and deposits six .epub files into dist/book/.

Build Single Volume with PDF

python3 scripts/build_book.py --volume language --pdf

Produces only the Language volume (vol3-language.epub + vol3-language.pdf).

Fast Assembly Without Rendering

python3 scripts/build_book.py --assemble-only

Generates markdown files under book/_build/ for CI checks or custom rendering pipelines without invoking Pandoc.

Programmatic Integration

from scripts.build_book import assemble, render, CONFIG

# Select the Deep Learning volume (second in config)

vol = CONFIG["volumes"][1]
md_path, chapters, words = assemble(vol)
artifacts = render(vol, md_path, chapters, pdf=True)
print("Created:", artifacts)

Summary

  • scripts/build_book.py serves as the complete build orchestrator for the AI Engineering from Scratch curriculum, handling discovery, transformation, and rendering.
  • The pipeline relies on book/volumes.json to map phases to volumes and sets up path constants (ROOT, PHASES, BUILD, DIST) for directory navigation.
  • Content transformation handles special fences (Mermaid diagrams, figures), rewrites asset paths, and inserts navigation links via transform_lesson().
  • Assembly concatenates processed lessons into volume manuscripts with Roman-numeral part headers using the assemble() function.
  • Rendering uses Pandoc for EPUB generation and optionally XeLaTeX for PDF creation, with metadata managed through metadata() and render().

Frequently Asked Questions

What dependencies are required to run the book generation pipeline?

The script requires Python 3.x and Pandoc for EPUB generation. For PDF output, you need xelatex installed on your system. Mermaid diagram rendering requires the mmdc (Mermaid CLI) tool available in your PATH.

Can I build the book without installing LaTeX?

Yes. Omit the --pdf flag when running scripts/build_book.py. The pipeline generates EPUB files using only Pandoc, which has fewer system dependencies than the PDF toolchain requiring XeLaTeX and specific system fonts.

How does the pipeline handle broken markdown or unclosed code fences?

The transform_lesson() function validates that all fenced code blocks close properly. If it encounters an unclosed fence, it raises a ValueError immediately (lines 63-65), preventing the generation of malformed books and surfacing the specific file location requiring correction.

Where are the intermediate files stored during the build process?

Assembled markdown files land in book/_build/ before Pandoc processing. Final EPUB and PDF artifacts are written to dist/book/. The --assemble-only flag stops execution after populating book/_build/, allowing inspection of intermediate markdown before rendering.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →