# How the Book Generation Pipeline Works in AI Engineering from Scratch

> Learn how the scripts/build_book.py script transforms Markdown into EPUB and PDF books by scanning phases/, processing content, and rendering with Pandoc.

- Repository: [Rohit Ghumare/ai-engineering-from-scratch](https://github.com/rohitg00/ai-engineering-from-scratch)
- Tags: how-to-guide
- Published: 2026-07-30

---

**The [`scripts/build_book.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/scripts/build_book.py) script transforms Markdown lesson files into professional EPUB and PDF books by scanning the `phases/` directory, processing content through a transformation pipeline, and rendering the final artifacts via Pandoc.**

The **AI Engineering from Scratch** repository uses a custom automation pipeline to compile scattered curriculum markdown into distributable book formats. The entry point, [`scripts/build_book.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/scripts/build_book.py), located at the repository root, orchestrates the entire workflow—from discovering lessons in phase directories to generating finished EPUB files with properly formatted metadata.

## Configuration and Volume Definitions

The pipeline starts by reading [`book/volumes.json`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/book/volumes.json) to understand the six-book series structure. This JSON file defines which phases belong to each volume (Foundations, Deep Learning, Language, etc.) and provides essential metadata like the site URL and repository path.

```python
CONFIG = json.loads((ROOT / "book" / "volumes.json").read_text())
SITE   = CONFIG["site"].rstrip("/")
REPO   = CONFIG["repo"].rstrip("/")

```

As defined in [`scripts/build_book.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/scripts/build_book.py) lines 28-32, the script establishes four pathlib shortcuts that drive the entire build process:

- `ROOT` — Repository root directory
- `PHASES` — The `phases/` directory containing curriculum content
- `BUILD` — Temporary build folder at `book/_build`
- `DIST` — Final distribution folder at `dist/book`

## Discovering and Indexing Lessons

For each volume defined in the configuration, the pipeline must locate the actual lesson content. The `lesson_dirs()` function (lines 45-53 in [`scripts/build_book.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/scripts/build_book.py)) scans each phase directory using the `LESSON_DIR_RE` regular expression imported from [`scripts/build_catalog.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/scripts/build_catalog.py).

```python
def lesson_dirs(phase):
    base = PHASES / phase
    ...
    return [d for d in sorted(base.iterdir())
            if d.is_dir() and LESSON_DIR_RE.match(d.name)
            and (d / "docs" / "en.md").is_file()]

```

This function returns only directories that match the lesson-slug convention and contain a [`docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/docs/en.md) file, ensuring the pipeline processes valid lesson structures while ignoring auxiliary directories.

## URL Generation and Link Construction

Before transforming content, the pipeline prepares navigation links. The `urls_for()` function (lines 60-66) constructs three essential URLs for every lesson:

```python
def urls_for(phase, lesson):
    rel = f"phases/{phase}/{lesson}"
    return {
        "web": f"{SITE}/lesson.html?path={rel}",
        "code": f"{REPO}/tree/main/{rel}/code",
        "repo": f"{REPO}/tree/main/{rel}",
    }

```

These URLs populate the "continue online" boxes in the final book, allowing readers to access the interactive web edition, view raw code in the repository, or navigate directly to the lesson root.

## Transforming Lesson Content

The `transform_lesson()` function serves as the core processing engine, reading each [`docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/docs/en.md) file line-by-line and applying specialized transformations. Located in the main script body, this function handles six distinct content types:

### Code Fences and Special Blocks

The parser detects fenced code blocks (lines 89-101) and treats special fence types differently. When encountering ` ```figure ` blocks, the pipeline replaces them with "continue-online" boxes linking to interactive web figures. Standard code fences pass through with syntax highlighting preserved.

### Mermaid Diagram Rendering

For ` ```mermaid ` fences, the pipeline calls `render_mermaid()` (lines 22-30) to generate SVG diagrams using the `mmdc` CLI tool. If Mermaid rendering fails, the system falls back to a placeholder box rather than failing the build, ensuring robustness across different environments.

### Ship It Sections and Exercises

The transformer collapses `## Ship It` markdown sections into single boxed notes pointing to artifact repositories (lines 35-44). For exercise sections, it appends links to starter code (lines 53-57), maintaining the connection between reading material and hands-on implementation.

### Asset Path Rewriting

Relative image links pointing to `../assets/` get rewritten to absolute paths within the built book structure (`phases/<phase>/<lesson>/assets/...`) on lines 60-62. This ensures images resolve correctly in the final EPUB regardless of the reader's device.

### Validation and Error Handling

If `transform_lesson()` encounters an unclosed fenced block, it raises a `ValueError` (lines 63-65) to surface formatting errors immediately rather than generating broken output.

## Assembling Complete Volumes

Once individual lessons are processed, the `assemble()` function (main body of [`scripts/build_book.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/scripts/build_book.py)) constructs complete volume manuscripts:

```python
def assemble(vol):
    BUILD.mkdir(parents=True, exist_ok=True)
    parts = [how_to_use(vol)]
    for part_idx, phase in enumerate(vol["phases"]):
        title = clean_phase_title(phase_title(phase))
        parts.append(f"\n# Part {ROMAN[part_idx]} — {title} {{.unnumbered .part}}\n\n"

                     f"*Course phase {phase.split('-')[0]}. … <{SITE}/catalog.html>*\n")
        for lesson_dir in lesson_dirs(phase):
            parts.append("\n".join(transform_lesson(phase, lesson_dir)))
    text = "\n\n".join(parts)
    md   = BUILD / f"{vol['slug']}.md"
    md.write_text(text)
    return md, chapters, len(text.split())

```

This function:

1. Creates the build directory structure
2. Prepends a "How to Use" preface explaining the volume scope
3. Generates Part headers using Roman numerals (`ROMAN` array) for each phase
4. Concatenates all transformed lessons with proper markdown separation
5. Writes the final file to `book/_build/<volume-slug>.md`

## Metadata Generation and Rendering

Before invoking Pandoc, the `metadata()` function (lines 96-106) writes a [`titlepage.yaml`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/titlepage.yaml) file containing the series title, volume subtitle, author information, language settings, and table-of-contents title.

### EPUB Generation

The `render()` function executes Pandoc with specific arguments (lines 15-21):

- Input: The assembled markdown file and metadata YAML
- Styling: Custom CSS from [`book/epub.css`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/book/epub.css)
- Options: Table of contents inclusion and proper chapter hierarchy

### PDF Generation (Optional)

When the `--pdf` flag is active, the pipeline generates a LaTeX title page using `book/titlepage.tex`, selects appropriate system fonts via `pick_font()` and `font_families()` (lines 70-86), and invokes Pandoc with `--pdf-engine=xelatex`. The font selection logic queries `fc-list` to find suitable serif and monospace fonts available on the build machine.

Both artifacts are written to `dist/book/` with size reporting to the console.

## CLI Orchestration and Build Flags

The `main()` function (lines 87-105) provides a command-line interface supporting three critical flags:

| Flag | Function |
|------|----------|
| `--volume <slug>` | Build only a specific volume instead of the entire series |
| `--pdf` | Enable PDF generation (requires `xelatex` installation) |
| `--assemble-only` | Stop after markdown generation, skipping Pandoc rendering |

Before building, the script validates phase existence using `check_phases()` to ensure every referenced phase directory actually contains lessons, preventing mid-build failures.

## Practical Usage Examples

### Build the Entire Series (EPUB only)

```bash
python3 scripts/build_book.py

```

This scans [`book/volumes.json`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/book/volumes.json), assembles all six volumes, and deposits six `.epub` files into `dist/book/`.

### Build Single Volume with PDF

```bash
python3 scripts/build_book.py --volume language --pdf

```

Produces only the Language volume (`vol3-language.epub` + `vol3-language.pdf`).

### Fast Assembly Without Rendering

```bash
python3 scripts/build_book.py --assemble-only

```

Generates markdown files under `book/_build/` for CI checks or custom rendering pipelines without invoking Pandoc.

### Programmatic Integration

```python
from scripts.build_book import assemble, render, CONFIG

# Select the Deep Learning volume (second in config)

vol = CONFIG["volumes"][1]
md_path, chapters, words = assemble(vol)
artifacts = render(vol, md_path, chapters, pdf=True)
print("Created:", artifacts)

```

## Summary

- **[`scripts/build_book.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/scripts/build_book.py)** serves as the complete build orchestrator for the AI Engineering from Scratch curriculum, handling discovery, transformation, and rendering.
- The pipeline relies on **[`book/volumes.json`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/book/volumes.json)** to map phases to volumes and sets up path constants (`ROOT`, `PHASES`, `BUILD`, `DIST`) for directory navigation.
- **Content transformation** handles special fences (Mermaid diagrams, figures), rewrites asset paths, and inserts navigation links via `transform_lesson()`.
- **Assembly** concatenates processed lessons into volume manuscripts with Roman-numeral part headers using the `assemble()` function.
- **Rendering** uses Pandoc for EPUB generation and optionally XeLaTeX for PDF creation, with metadata managed through `metadata()` and `render()`.

## Frequently Asked Questions

### What dependencies are required to run the book generation pipeline?

The script requires Python 3.x and Pandoc for EPUB generation. For PDF output, you need `xelatex` installed on your system. Mermaid diagram rendering requires the `mmdc` (Mermaid CLI) tool available in your PATH.

### Can I build the book without installing LaTeX?

Yes. Omit the `--pdf` flag when running [`scripts/build_book.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/scripts/build_book.py). The pipeline generates EPUB files using only Pandoc, which has fewer system dependencies than the PDF toolchain requiring XeLaTeX and specific system fonts.

### How does the pipeline handle broken markdown or unclosed code fences?

The `transform_lesson()` function validates that all fenced code blocks close properly. If it encounters an unclosed fence, it raises a `ValueError` immediately (lines 63-65), preventing the generation of malformed books and surfacing the specific file location requiring correction.

### Where are the intermediate files stored during the build process?

Assembled markdown files land in `book/_build/` before Pandoc processing. Final EPUB and PDF artifacts are written to `dist/book/`. The `--assemble-only` flag stops execution after populating `book/_build/`, allowing inspection of intermediate markdown before rendering.