# Building Complete LLM Pre‑Training Pipelines from Scratch: A Deep Dive into the AI‑Engineering‑From‑Scratch Curriculum

> Learn to build complete LLM pre-training pipelines from scratch with this 503-lesson curriculum. Implement every component in pure Python, Rust, or TypeScript before using production frameworks.

- Repository: [Rohit Ghumare/ai-engineering-from-scratch](https://github.com/rohitg00/ai-engineering-from-scratch)
- Tags: deep-dive
- Published: 2026-07-26

---

**The AI‑Engineering‑From‑Scratch repository provides a 503‑lesson curriculum that teaches you to construct end‑to‑end LLM pre‑training pipelines—from raw tokenizers to RLHF—by implementing every component first in pure Python (or Rust/TypeScript) before using production frameworks like PyTorch.**

The rohitg00/ai-engineering-from-scratch repository is a self‑contained, language‑agnostic educational framework designed to take you from linear‑algebra foundations to production‑grade LLM engineering. By following its strict "Build It / Use It" methodology, you construct complete LLM pre‑training pipelines from scratch without relying on black‑box abstractions you do not understand.

## Repository Architecture and Lesson Structure

The repository organizes content into a deterministic filesystem that separates curriculum content from automation tooling.

### The Phase‑Lesson Hierarchy

Every lesson lives in `phases/<NN>-<phase‑name>/<NN>-<lesson‑name>/` and contains three mandatory subdirectories:

- `code/` – Runnable implementations in Python, TypeScript, Rust, or Julia
- [`docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/docs/en.md) – Narrative explanation and learning objectives  
- `outputs/` – Exported artifacts (prompts, skills, agents, MCP servers)

This structure is enforced by [`scripts/audit_lessons.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/scripts/audit_lessons.py), which runs static checks to ensure each lesson follows the required schema and commit conventions.

### The Six‑Beat "Build It / Use It" Split

Each lesson follows a rigid pedagogical pattern defined in the repository’s [`README.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/README.md):

1. **MOTTO** – One‑line core idea capturing the lesson’s essence
2. **PROBLEM** – Concrete pain point that motivates the solution  
3. **CONCEPT** – Intuition and diagrams explaining the theory
4. **BUILD IT** – Raw‑math implementation with zero frameworks
5. **USE IT** – Equivalent implementation using production libraries (e.g., PyTorch)
6. **SHIP IT** – A reusable artifact exported to `outputs/`

This guarantees that every artifact you ship is **understood** because you authored the low‑level version yourself.

## Curriculum Flow: From Math to Production

The curriculum is organized as a directed phase graph (visualized via Mermaid in [`README.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/README.md)) that ensures prerequisite knowledge is always satisfied before advancing.

- **Phases 0‑1:** Linear algebra and deep‑learning foundations  
- **Phases 2‑5:** Vision, NLP, and transformer architectures  
- **Phase 10:** LLMs from Scratch (the core pre‑training pipeline)
- **Phases 11‑14:** RL, tools, protocols, and agent engineering

Learners may skip ahead if they already possess lower‑layer knowledge, but the graph structure ensures that lessons like `phases/10-llms-from-scratch/04-pre-training-mini-gpt/` assume mastery of earlier tensor operations and attention mechanisms.

## Automation and CI Pipeline

The repository maintains quality through automated tooling located in `scripts/`:

- **[`scripts/audit_lessons.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/scripts/audit_lessons.py)** – Enforces hard rules such as one‑commit‑per‑lesson and valid [`quiz.json`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/quiz.json) schema
- **[`scripts/build_catalog.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/scripts/build_catalog.py)** – Parses the entire curriculum to generate [`site/data.js`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/site/data.js), which powers the static website’s lesson table  
- **[`scripts/build_book.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/scripts/build_book.py)** – Assembles the six‑volume PDF/EPUB book from lesson markdown  
- **[`scripts/install_skills.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/scripts/install_skills.py)** – Registers every artifact in `outputs/` directories into a running LLM agent

GitHub Actions (defined in [`.github/workflows/curriculum.yml`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/.github/workflows/curriculum.yml)) run the audit, regenerate [`site/data.js`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/site/data.js), and update README statistics on every push.

## Constructing the LLM Pre‑Training Pipeline (Phase 10)

Phase 10 is the culmination of the curriculum, where you assemble a complete GPT‑style pre‑training pipeline lesson‑by‑lesson:

### Lesson 01 – Tokenizers

You implement **Byte‑Pair Encoding (BPE)**, WordPiece, and SentencePiece in pure Python within [`phases/10-llms-from-scratch/01-tokenizers/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/10-llms-from-scratch/01-tokenizers/code/main.py). An optional Rust version demonstrates performance optimization.

### Lesson 02 – Building a Tokenizer Pipeline

Constructs an end‑to‑end tokenization pipeline that reads raw text, builds a vocabulary, and serializes to [`tokenizer.json`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/tokenizer.json).

### Lesson 03 – Data Pipelines  

Implements streaming and sharding for pre‑training datasets. You build a `datasets`‑style iterator that yields token IDs in memory‑efficient batches without loading the entire corpus into RAM.

### Lesson 04 – Mini‑GPT Pre‑Training  

The core implementation in [`phases/10-llms-from-scratch/04-pre-training-mini-gpt/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/10-llms-from-scratch/04-pre-training-mini-gpt/code/main.py) trains a **124 M‑parameter transformer** from scratch. You write the transformer blocks, masked‑language‑model loss, and training loop using only NumPy‑like primitives before porting to PyTorch.

### Lesson 05 – Distributed Training  

Wraps the mini‑GPT using **FSDP** and **DeepSpeed** wrappers to scale across multiple GPUs, teaching gradient sharding and mixed‑precision mechanics.

### Lesson 06 – Instruction Tuning (SFT)  

Implements **LoRA**‑based fine‑tuning, adding low‑rank adapters to the pretrained model rather than updating full weights.

### Lesson 07 – RLHF (PPO)  

Builds the complete RLHF stack: a reward model trained on human preferences, followed by the **PPO** (Proximal Policy Optimization) loop that refines the fine‑tuned model against that reward signal.

By the end of Phase 10, you possess a **complete LLM pipeline** capable of pre‑training, fine‑tuning, evaluation, and quantization—every component of which you wrote first by hand.

## Running Lessons and Validating Code

Each lesson is immediately executable. To run the mini‑GPT pre‑training loop:

```bash
git clone https://github.com/rohitg00/ai-engineering-from-scratch.git
cd ai-engineering-from-scratch
python phases/10-llms-from-scratch/04-pre-training-mini-gpt/code/main.py

```

Unit tests reside in `code/tests/` within each lesson directory. The same execution pattern works for any lesson; the [`README.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/README.md) table lists exact entry points for all 503 lessons.

To generate the static site locally:

```bash
node site/build.js  # Parses README/ROADMAP/glossary → site/data.js

```

## Exporting Reusable Artifacts

Every lesson’s `outputs/` directory contains a markdown artifact with frontmatter that can be imported directly into LLM agents:

```markdown
---
name: token-loss-debugger
description: Diagnose perplexity spikes in transformer training
phase: 10
lesson: 04
---
Analyze the gradient norms and identify...

```

Install all 503 skills into a running agent using:

```bash
python scripts/install_skills.py

```

This transforms the curriculum into a **production toolbox** of prompts and agents that you both built and understand.

## Summary

- The **AI‑Engineering‑From‑Scratch** repository contains **503 lessons** across **20 phases**, teaching LLM engineering from linear algebra to RLHF.
- The **"Build It / Use It"** methodology requires implementing algorithms in raw math before using PyTorch or similar frameworks.
- **Phase 10** specifically constructs a complete pre‑training pipeline: tokenizers, data loaders, a 124 M‑parameter GPT, distributed training wrappers, LoRA fine‑tuning, and PPO‑based RLHF.
- Automation scripts ([`audit_lessons.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/audit_lessons.py), [`build_catalog.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/build_catalog.py), [`install_skills.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/install_skills.py)) enforce quality and export reusable artifacts.
- Every lesson is runnable locally and produces **shippable artifacts** (skills, prompts, agents) stored in `outputs/` directories.

## Frequently Asked Questions

### What prerequisites are required to start building LLM pre‑training pipelines from scratch?

You need basic Python proficiency and high‑school linear algebra. The curriculum starts at Phase 0 with tensor fundamentals, so you can begin immediately even if you have never implemented a neural network. However, you may skip early phases if you already understand matrix operations and backpropagation.

### How does the "Build It / Use It" methodology work in practice?

First, you implement the algorithm using only NumPy or raw Python (the **BUILD IT** step) to understand the underlying math. Then, you refactor the same logic using PyTorch or Hugging Face Transformers (the **USE IT** step) to learn production APIs. Finally, you export a reusable artifact (the **SHIP IT** step) to `outputs/`, ensuring you can apply the knowledge without relying on framework magic you did not write.

### Can I skip ahead to Phase 10 if I already know deep‑learning fundamentals?

Yes. The curriculum is organized as a phase graph, and the [`README.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/README.md) explicitly lists prerequisites for each lesson. If you already understand transformers and attention mechanisms from earlier phases, you can jump directly to `phases/10-llms-from-scratch/` and run the pre‑training pipeline code.

### What hardware is needed for the mini‑GPT pre‑training lesson?

The 124 M‑parameter model in `phases/10-llms-from-scratch/04-pre-training-mini-gpt/` can train on a single GPU with 16 GB VRAM using mixed precision. Lesson 05 adds distributed training support via FSDP and DeepSpeed for multi‑GPU scaling, but the base implementation is designed to run on consumer hardware for educational purposes.