Building Complete LLM Pre‑Training Pipelines from Scratch: A Deep Dive into the AI‑Engineering‑From‑Scratch Curriculum

The AI‑Engineering‑From‑Scratch repository provides a 503‑lesson curriculum that teaches you to construct end‑to‑end LLM pre‑training pipelines—from raw tokenizers to RLHF—by implementing every component first in pure Python (or Rust/TypeScript) before using production frameworks like PyTorch.

The rohitg00/ai-engineering-from-scratch repository is a self‑contained, language‑agnostic educational framework designed to take you from linear‑algebra foundations to production‑grade LLM engineering. By following its strict "Build It / Use It" methodology, you construct complete LLM pre‑training pipelines from scratch without relying on black‑box abstractions you do not understand.

Repository Architecture and Lesson Structure

The repository organizes content into a deterministic filesystem that separates curriculum content from automation tooling.

The Phase‑Lesson Hierarchy

Every lesson lives in phases/<NN>-<phase‑name>/<NN>-<lesson‑name>/ and contains three mandatory subdirectories:

  • code/ – Runnable implementations in Python, TypeScript, Rust, or Julia
  • docs/en.md – Narrative explanation and learning objectives
  • outputs/ – Exported artifacts (prompts, skills, agents, MCP servers)

This structure is enforced by scripts/audit_lessons.py, which runs static checks to ensure each lesson follows the required schema and commit conventions.

The Six‑Beat "Build It / Use It" Split

Each lesson follows a rigid pedagogical pattern defined in the repository’s README.md:

  1. MOTTO – One‑line core idea capturing the lesson’s essence
  2. PROBLEM – Concrete pain point that motivates the solution
  3. CONCEPT – Intuition and diagrams explaining the theory
  4. BUILD IT – Raw‑math implementation with zero frameworks
  5. USE IT – Equivalent implementation using production libraries (e.g., PyTorch)
  6. SHIP IT – A reusable artifact exported to outputs/

This guarantees that every artifact you ship is understood because you authored the low‑level version yourself.

Curriculum Flow: From Math to Production

The curriculum is organized as a directed phase graph (visualized via Mermaid in README.md) that ensures prerequisite knowledge is always satisfied before advancing.

  • Phases 0‑1: Linear algebra and deep‑learning foundations
  • Phases 2‑5: Vision, NLP, and transformer architectures
  • Phase 10: LLMs from Scratch (the core pre‑training pipeline)
  • Phases 11‑14: RL, tools, protocols, and agent engineering

Learners may skip ahead if they already possess lower‑layer knowledge, but the graph structure ensures that lessons like phases/10-llms-from-scratch/04-pre-training-mini-gpt/ assume mastery of earlier tensor operations and attention mechanisms.

Automation and CI Pipeline

The repository maintains quality through automated tooling located in scripts/:

GitHub Actions (defined in .github/workflows/curriculum.yml) run the audit, regenerate site/data.js, and update README statistics on every push.

Constructing the LLM Pre‑Training Pipeline (Phase 10)

Phase 10 is the culmination of the curriculum, where you assemble a complete GPT‑style pre‑training pipeline lesson‑by‑lesson:

Lesson 01 – Tokenizers

You implement Byte‑Pair Encoding (BPE), WordPiece, and SentencePiece in pure Python within phases/10-llms-from-scratch/01-tokenizers/code/main.py. An optional Rust version demonstrates performance optimization.

Lesson 02 – Building a Tokenizer Pipeline

Constructs an end‑to‑end tokenization pipeline that reads raw text, builds a vocabulary, and serializes to tokenizer.json.

Lesson 03 – Data Pipelines

Implements streaming and sharding for pre‑training datasets. You build a datasets‑style iterator that yields token IDs in memory‑efficient batches without loading the entire corpus into RAM.

Lesson 04 – Mini‑GPT Pre‑Training

The core implementation in phases/10-llms-from-scratch/04-pre-training-mini-gpt/code/main.py trains a 124 M‑parameter transformer from scratch. You write the transformer blocks, masked‑language‑model loss, and training loop using only NumPy‑like primitives before porting to PyTorch.

Lesson 05 – Distributed Training

Wraps the mini‑GPT using FSDP and DeepSpeed wrappers to scale across multiple GPUs, teaching gradient sharding and mixed‑precision mechanics.

Lesson 06 – Instruction Tuning (SFT)

Implements LoRA‑based fine‑tuning, adding low‑rank adapters to the pretrained model rather than updating full weights.

Lesson 07 – RLHF (PPO)

Builds the complete RLHF stack: a reward model trained on human preferences, followed by the PPO (Proximal Policy Optimization) loop that refines the fine‑tuned model against that reward signal.

By the end of Phase 10, you possess a complete LLM pipeline capable of pre‑training, fine‑tuning, evaluation, and quantization—every component of which you wrote first by hand.

Running Lessons and Validating Code

Each lesson is immediately executable. To run the mini‑GPT pre‑training loop:

git clone https://github.com/rohitg00/ai-engineering-from-scratch.git
cd ai-engineering-from-scratch
python phases/10-llms-from-scratch/04-pre-training-mini-gpt/code/main.py

Unit tests reside in code/tests/ within each lesson directory. The same execution pattern works for any lesson; the README.md table lists exact entry points for all 503 lessons.

To generate the static site locally:

node site/build.js  # Parses README/ROADMAP/glossary → site/data.js

Exporting Reusable Artifacts

Every lesson’s outputs/ directory contains a markdown artifact with frontmatter that can be imported directly into LLM agents:

---
name: token-loss-debugger
description: Diagnose perplexity spikes in transformer training
phase: 10
lesson: 04
---
Analyze the gradient norms and identify...

Install all 503 skills into a running agent using:

python scripts/install_skills.py

This transforms the curriculum into a production toolbox of prompts and agents that you both built and understand.

Summary

  • The AI‑Engineering‑From‑Scratch repository contains 503 lessons across 20 phases, teaching LLM engineering from linear algebra to RLHF.
  • The "Build It / Use It" methodology requires implementing algorithms in raw math before using PyTorch or similar frameworks.
  • Phase 10 specifically constructs a complete pre‑training pipeline: tokenizers, data loaders, a 124 M‑parameter GPT, distributed training wrappers, LoRA fine‑tuning, and PPO‑based RLHF.
  • Automation scripts (audit_lessons.py, build_catalog.py, install_skills.py) enforce quality and export reusable artifacts.
  • Every lesson is runnable locally and produces shippable artifacts (skills, prompts, agents) stored in outputs/ directories.

Frequently Asked Questions

What prerequisites are required to start building LLM pre‑training pipelines from scratch?

You need basic Python proficiency and high‑school linear algebra. The curriculum starts at Phase 0 with tensor fundamentals, so you can begin immediately even if you have never implemented a neural network. However, you may skip early phases if you already understand matrix operations and backpropagation.

How does the "Build It / Use It" methodology work in practice?

First, you implement the algorithm using only NumPy or raw Python (the BUILD IT step) to understand the underlying math. Then, you refactor the same logic using PyTorch or Hugging Face Transformers (the USE IT step) to learn production APIs. Finally, you export a reusable artifact (the SHIP IT step) to outputs/, ensuring you can apply the knowledge without relying on framework magic you did not write.

Can I skip ahead to Phase 10 if I already know deep‑learning fundamentals?

Yes. The curriculum is organized as a phase graph, and the README.md explicitly lists prerequisites for each lesson. If you already understand transformers and attention mechanisms from earlier phases, you can jump directly to phases/10-llms-from-scratch/ and run the pre‑training pipeline code.

What hardware is needed for the mini‑GPT pre‑training lesson?

The 124 M‑parameter model in phases/10-llms-from-scratch/04-pre-training-mini-gpt/ can train on a single GPU with 16 GB VRAM using mixed precision. Lesson 05 adds distributed training support via FSDP and DeepSpeed for multi‑GPU scaling, but the base implementation is designed to run on consumer hardware for educational purposes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →