How the Book Generation Pipeline Works in AI Engineering from Scratch
The scripts/build_book.py script transforms Markdown lesson files into professional EPUB and PDF books by scanning the phases/ directory, processing content through a transformation pipeline, and rendering the final artifacts via Pandoc.
The AI Engineering from Scratch repository uses a custom automation pipeline to compile scattered curriculum markdown into distributable book formats. The entry point, scripts/build_book.py, located at the repository root, orchestrates the entire workflow—from discovering lessons in phase directories to generating finished EPUB files with properly formatted metadata.
Configuration and Volume Definitions
The pipeline starts by reading book/volumes.json to understand the six-book series structure. This JSON file defines which phases belong to each volume (Foundations, Deep Learning, Language, etc.) and provides essential metadata like the site URL and repository path.
CONFIG = json.loads((ROOT / "book" / "volumes.json").read_text())
SITE = CONFIG["site"].rstrip("/")
REPO = CONFIG["repo"].rstrip("/")
As defined in scripts/build_book.py lines 28-32, the script establishes four pathlib shortcuts that drive the entire build process:
ROOT— Repository root directoryPHASES— Thephases/directory containing curriculum contentBUILD— Temporary build folder atbook/_buildDIST— Final distribution folder atdist/book
Discovering and Indexing Lessons
For each volume defined in the configuration, the pipeline must locate the actual lesson content. The lesson_dirs() function (lines 45-53 in scripts/build_book.py) scans each phase directory using the LESSON_DIR_RE regular expression imported from scripts/build_catalog.py.
def lesson_dirs(phase):
base = PHASES / phase
...
return [d for d in sorted(base.iterdir())
if d.is_dir() and LESSON_DIR_RE.match(d.name)
and (d / "docs" / "en.md").is_file()]
This function returns only directories that match the lesson-slug convention and contain a docs/en.md file, ensuring the pipeline processes valid lesson structures while ignoring auxiliary directories.
URL Generation and Link Construction
Before transforming content, the pipeline prepares navigation links. The urls_for() function (lines 60-66) constructs three essential URLs for every lesson:
def urls_for(phase, lesson):
rel = f"phases/{phase}/{lesson}"
return {
"web": f"{SITE}/lesson.html?path={rel}",
"code": f"{REPO}/tree/main/{rel}/code",
"repo": f"{REPO}/tree/main/{rel}",
}
These URLs populate the "continue online" boxes in the final book, allowing readers to access the interactive web edition, view raw code in the repository, or navigate directly to the lesson root.
Transforming Lesson Content
The transform_lesson() function serves as the core processing engine, reading each docs/en.md file line-by-line and applying specialized transformations. Located in the main script body, this function handles six distinct content types:
Code Fences and Special Blocks
The parser detects fenced code blocks (lines 89-101) and treats special fence types differently. When encountering ```figure blocks, the pipeline replaces them with "continue-online" boxes linking to interactive web figures. Standard code fences pass through with syntax highlighting preserved.
Mermaid Diagram Rendering
For ```mermaid fences, the pipeline calls render_mermaid() (lines 22-30) to generate SVG diagrams using the mmdc CLI tool. If Mermaid rendering fails, the system falls back to a placeholder box rather than failing the build, ensuring robustness across different environments.
Ship It Sections and Exercises
The transformer collapses ## Ship It markdown sections into single boxed notes pointing to artifact repositories (lines 35-44). For exercise sections, it appends links to starter code (lines 53-57), maintaining the connection between reading material and hands-on implementation.
Asset Path Rewriting
Relative image links pointing to ../assets/ get rewritten to absolute paths within the built book structure (phases/<phase>/<lesson>/assets/...) on lines 60-62. This ensures images resolve correctly in the final EPUB regardless of the reader's device.
Validation and Error Handling
If transform_lesson() encounters an unclosed fenced block, it raises a ValueError (lines 63-65) to surface formatting errors immediately rather than generating broken output.
Assembling Complete Volumes
Once individual lessons are processed, the assemble() function (main body of scripts/build_book.py) constructs complete volume manuscripts:
def assemble(vol):
BUILD.mkdir(parents=True, exist_ok=True)
parts = [how_to_use(vol)]
for part_idx, phase in enumerate(vol["phases"]):
title = clean_phase_title(phase_title(phase))
parts.append(f"\n# Part {ROMAN[part_idx]} — {title} {{.unnumbered .part}}\n\n"
f"*Course phase {phase.split('-')[0]}. … <{SITE}/catalog.html>*\n")
for lesson_dir in lesson_dirs(phase):
parts.append("\n".join(transform_lesson(phase, lesson_dir)))
text = "\n\n".join(parts)
md = BUILD / f"{vol['slug']}.md"
md.write_text(text)
return md, chapters, len(text.split())
This function:
- Creates the build directory structure
- Prepends a "How to Use" preface explaining the volume scope
- Generates Part headers using Roman numerals (
ROMANarray) for each phase - Concatenates all transformed lessons with proper markdown separation
- Writes the final file to
book/_build/<volume-slug>.md
Metadata Generation and Rendering
Before invoking Pandoc, the metadata() function (lines 96-106) writes a titlepage.yaml file containing the series title, volume subtitle, author information, language settings, and table-of-contents title.
EPUB Generation
The render() function executes Pandoc with specific arguments (lines 15-21):
- Input: The assembled markdown file and metadata YAML
- Styling: Custom CSS from
book/epub.css - Options: Table of contents inclusion and proper chapter hierarchy
PDF Generation (Optional)
When the --pdf flag is active, the pipeline generates a LaTeX title page using book/titlepage.tex, selects appropriate system fonts via pick_font() and font_families() (lines 70-86), and invokes Pandoc with --pdf-engine=xelatex. The font selection logic queries fc-list to find suitable serif and monospace fonts available on the build machine.
Both artifacts are written to dist/book/ with size reporting to the console.
CLI Orchestration and Build Flags
The main() function (lines 87-105) provides a command-line interface supporting three critical flags:
| Flag | Function |
|---|---|
--volume <slug> |
Build only a specific volume instead of the entire series |
--pdf |
Enable PDF generation (requires xelatex installation) |
--assemble-only |
Stop after markdown generation, skipping Pandoc rendering |
Before building, the script validates phase existence using check_phases() to ensure every referenced phase directory actually contains lessons, preventing mid-build failures.
Practical Usage Examples
Build the Entire Series (EPUB only)
python3 scripts/build_book.py
This scans book/volumes.json, assembles all six volumes, and deposits six .epub files into dist/book/.
Build Single Volume with PDF
python3 scripts/build_book.py --volume language --pdf
Produces only the Language volume (vol3-language.epub + vol3-language.pdf).
Fast Assembly Without Rendering
python3 scripts/build_book.py --assemble-only
Generates markdown files under book/_build/ for CI checks or custom rendering pipelines without invoking Pandoc.
Programmatic Integration
from scripts.build_book import assemble, render, CONFIG
# Select the Deep Learning volume (second in config)
vol = CONFIG["volumes"][1]
md_path, chapters, words = assemble(vol)
artifacts = render(vol, md_path, chapters, pdf=True)
print("Created:", artifacts)
Summary
scripts/build_book.pyserves as the complete build orchestrator for the AI Engineering from Scratch curriculum, handling discovery, transformation, and rendering.- The pipeline relies on
book/volumes.jsonto map phases to volumes and sets up path constants (ROOT,PHASES,BUILD,DIST) for directory navigation. - Content transformation handles special fences (Mermaid diagrams, figures), rewrites asset paths, and inserts navigation links via
transform_lesson(). - Assembly concatenates processed lessons into volume manuscripts with Roman-numeral part headers using the
assemble()function. - Rendering uses Pandoc for EPUB generation and optionally XeLaTeX for PDF creation, with metadata managed through
metadata()andrender().
Frequently Asked Questions
What dependencies are required to run the book generation pipeline?
The script requires Python 3.x and Pandoc for EPUB generation. For PDF output, you need xelatex installed on your system. Mermaid diagram rendering requires the mmdc (Mermaid CLI) tool available in your PATH.
Can I build the book without installing LaTeX?
Yes. Omit the --pdf flag when running scripts/build_book.py. The pipeline generates EPUB files using only Pandoc, which has fewer system dependencies than the PDF toolchain requiring XeLaTeX and specific system fonts.
How does the pipeline handle broken markdown or unclosed code fences?
The transform_lesson() function validates that all fenced code blocks close properly. If it encounters an unclosed fence, it raises a ValueError immediately (lines 63-65), preventing the generation of malformed books and surfacing the specific file location requiring correction.
Where are the intermediate files stored during the build process?
Assembled markdown files land in book/_build/ before Pandoc processing. Final EPUB and PDF artifacts are written to dist/book/. The --assemble-only flag stops execution after populating book/_build/, allowing inspection of intermediate markdown before rendering.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →