# How Lesson Prerequisites Are Defined and Connected Across Phases in rohitg00/ai-engineering-from-scratch

> Discover how rohitg00/ai-engineering-from-scratch defines lesson prerequisites using directed graph metadata to ensure valid ordering and generate navigation links across its roadmap.

- Repository: [Rohit Ghumare/ai-engineering-from-scratch](https://github.com/rohitg00/ai-engineering-from-scratch)
- Tags: internals
- Published: 2026-07-30

---

**The curriculum defines lesson prerequisites as directed graph dependencies within each lesson’s [`docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/docs/en.md) metadata, parsing them with regular expressions to validate ordering, generate navigation links, and detect cyclic dependencies across the 435-lesson roadmap.**

The rohitg00/ai-engineering-from-scratch repository implements a sophisticated dependency management system for its AI engineering curriculum. Each lesson declares its prerequisites in plain text within markdown metadata, enabling complex cross-phase connections that ensure learners master foundational concepts before advancing to complex topics like LLM engineering and safety systems.

## Prerequisite Declaration in Lesson Metadata

Every lesson’s metadata lives in [`docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/docs/en.md) alongside the lesson code and tests. Within this file, the **Prerequisites** field uses a concise notation to declare knowledge dependencies.

### The docs/en.md Structure

The prerequisite line appears as standard bold text in the markdown:

```markdown
**Prerequisites:** <description>

```

This pattern is scanned by the validation script using the regular expression `r'^\*\*Prerequisites:\*\*\s*(.+)$'` to extract the raw dependency string. The description references **phases** and **lesson numbers** (or track blocks) using a human-readable syntax that the build system later resolves into concrete lesson paths.

## Prerequisite Notation and Cross-Phase Referencing

The curriculum supports several reference patterns to accommodate the multi-phase structure:

- **Range notation**: `Phase 19 Track A lessons 20‑29` indicates that every lesson numbered 20 through 29 in Phase 19 Track A must be completed first.
- **Single lesson**: `Phase 18 · 16 (Llama Guard / Garak / PyRIT)` references a specific lesson in Phase 18.
- **Multiple selective lessons**: `Phase 11 lessons 04 (embeddings), 06 (RAG)` lists specific lessons within the same phase.

Because prerequisites are expressed in plain text, they can reference any previous phase. For example, a capstone lesson in Phase 19 may require foundational lessons from Phase 11 (LLM engineering) and safety lessons from Phase 18, creating a cross-phase dependency that guarantees learners possess the necessary background before tackling advanced material.

## Parsing and Validation with scripts/audit_lessons.py

The repository’s validation logic resides in [`scripts/audit_lessons.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/scripts/audit_lessons.py), which scans every [`docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/docs/en.md) file to build a complete dependency graph.

### Extracting Prerequisites with Regular Expressions

The audit script extracts prerequisite strings using pattern matching against the bold markdown syntax:

```python
import re
from pathlib import Path

PREREQ_RE = re.compile(r'^\*\*Prerequisites:\*\*\s*(.+)$', re.MULTILINE)

def extract_prereqs(md_path: Path) -> list[str]:
    """Return a list of raw prerequisite strings from a lesson’s docs/en.md."""
    text = md_path.read_text()
    match = PREREQ_RE.search(text)
    if not match:
        return []
    # Split on commas and strip whitespace

    raw = match.group(1)
    return [s.strip() for s in raw.split(',')]

```

This function returns a list of raw prerequisite strings that reference phases and lesson numbers.

### Building the Dependency Graph

The validation pipeline constructs a directed graph using **NetworkX** to model lesson dependencies:

```python
import networkx as nx

def build_dependency_graph(root: Path) -> nx.DiGraph:
    G = nx.DiGraph()
    for md in root.rglob('docs/en.md'):
        lesson = md.parent.relative_to(root).as_posix()
        G.add_node(lesson)
        for prereq in extract_prereqs(md):
            # Normalize prerequisite text to lesson identifiers

            for dep in resolve_prereq_to_lessons(prereq, root):
                G.add_edge(dep, lesson)
    return G

```

The helper function `resolve_prereq_to_lessons` interprets strings such as `Phase 19 Track A lessons 20‑29` into concrete lesson paths (for example, `phases/19-capstone-projects/20-task-spec-format`).

### Cycle Detection in CI

The continuous integration pipeline validates the curriculum by detecting circular dependencies that would break the learning flow:

```python
def check_for_cycles(G: nx.DiGraph) -> list[list[str]]:
    return list(nx.simple_cycles(G))

# In CI

if cycles := check_for_cycles(dep_graph):
    raise RuntimeError(f"Prerequisite cycles detected: {cycles}")

```

The CI `audit` job runs this check and fails the pull request if any cycle exists, ensuring the curriculum remains a valid directed acyclic graph.

## Site Generation and Navigation

The site builder at [`site/build.js`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/site/build.js) consumes the dependency graph to create navigation links between lessons. When prerequisites resolve to specific lesson identifiers, the builder generates forward and backward links in the static site, allowing learners to jump through the curriculum according to dependency relationships rather than just sequential ordering.

## Key Files in the Prerequisite System

Several files work together to enforce the learning path:

- **[`phases/19-capstone-projects/87-end-to-end-safety-gate/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/19-capstone-projects/87-end-to-end-safety-gate/docs/en.md)**: Contains the **Prerequisites** line demonstrating the metadata format used across the curriculum.
- **[`scripts/audit_lessons.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/scripts/audit_lessons.py)**: Scans all lesson directories, parses prerequisite strings with regular expressions, and builds the dependency graph for validation.
- **[`site/build.js`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/site/build.js)**: Reads the resolved prerequisite relationships to generate inter-lesson navigation in the static site output.
- **[`ROADMAP.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/ROADMAP.md)**: Reflects prerequisite relationships by ordering lessons appropriately according to their dependency constraints.
- **[`README.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/README.md)**: Contains the lesson table with markdown links that must align with the prerequisite graph validated by the audit script.

## Summary

- **Prerequisites are declared** in [`docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/docs/en.md) using bold markdown syntax (`**Prerequisites:**`) that lists required phases and lesson numbers.
- **The validation script** ([`scripts/audit_lessons.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/scripts/audit_lessons.py)) parses these strings with regular expressions and constructs a NetworkX directed graph to model dependencies.
- **Cross-phase connections** allow advanced lessons to require foundational knowledge from any earlier phase in the 435-lesson curriculum.
- **CI pipeline protection** detects cyclic dependencies using `nx.simple_cycles()` and fails builds if circular references exist.
- **Site navigation** is generated automatically from the dependency graph, creating links that respect the curriculum's knowledge prerequisites.

## Frequently Asked Questions

### How are lesson prerequisites formatted in the markdown files?

Each lesson includes a line in [`docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/docs/en.md) formatted as `**Prerequisites:**` followed by a description using phase and lesson identifiers. For example, `Phase 19 Track A lessons 20‑29` or `Phase 11 lessons 04 (embeddings), 06 (RAG)`. The validation script treats this bold text as structured metadata to build the dependency graph.

### What happens if a prerequisite cycle is detected?

The CI pipeline runs `check_for_cycles()` from [`scripts/audit_lessons.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/scripts/audit_lessons.py), which uses NetworkX's `simple_cycles()` function to find circular dependencies. If any cycles exist, the script raises a `RuntimeError` with the detected cycle paths and fails the build, preventing the curriculum from containing logical loops that would trap learners in circular requirements.

### Can lessons depend on content from multiple phases?

Yes. The plain-text prerequisite notation supports cross-phase references, allowing a single lesson to require knowledge from any combination of earlier phases. For instance, a Phase 19 capstone can list prerequisites from Phase 11 (embeddings/RAG) and Phase 18 (safety tooling), ensuring comprehensive preparation before advanced study.

### How does the site builder use prerequisite data?

The [`site/build.js`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/site/build.js) script consumes the dependency graph built by [`scripts/audit_lessons.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/scripts/audit_lessons.py) to generate navigation links between lessons. When the builder creates the static site, it uses the prerequisite relationships to insert forward links (lessons that depend on the current one) and backward links (prerequisites), enabling non-linear traversal of the curriculum based on knowledge dependencies rather than just file system ordering.