How Phase 7 Transformers Build the Foundation for Phase 10 LLMs from Scratch
Phase 7 supplies the core transformer primitives—self-attention mechanisms, multi-head attention layers, and positional encodings—that Phase 10 imports and stacks to create a complete large language model from scratch.
The rohitg00/ai-engineering-from-scratch repository structures its curriculum as a progressive build pipeline, where Phase 7 (Transformers Deep Dive) focuses on implementing the mathematical building blocks of the transformer architecture, and Phase 10 (LLMs from Scratch) leverages these exact components to construct a trainable language model. Understanding this relationship reveals how modern LLMs inherit their power from the reusable, composable transformer blocks developed in earlier phases.
What Phase 7 Provides: The Transformer Primitives
Phase 7 implements the fundamental operations required for sequence modeling in pure Python and PyTorch, creating a library of validated components that handle tensor operations.
Self-Attention and Multi-Head Mechanisms
At the heart of Phase 7 lies the self-attention kernel implemented in phases/07-transformers-deep-dive/code/attention.py, which computes scaled dot-product attention between query, key, and value tensors. The repository wraps this in phases/07-transformers-deep-dive/code/multi_head.py to enable multi-head attention, allowing the model to jointly attend to information from different representation subspaces at different positions.
Positional Encodings and Feed-Forward Networks
To inject sequence order information, phases/07-transformers-deep-dive/code/positional_encoding.py implements sinusoidal position encodings that are added to input embeddings. These feed into phases/07-transformers-deep-dive/code/transformer_block.py, which combines the multi-head attention with a position-wise feed-forward network and residual connections, creating a complete, reusable transformer block.
How Phase 10 Consumes Phase 7: Assembling the LLM
Phase 10 does not reimplement these mathematical operations; instead, it treats Phase 7 as a dependency, importing the validated transformer blocks and adding LLM-specific infrastructure for text processing and generation.
Importing the Transformer Block
The LLM implementation imports the TransformerBlock class directly from Phase 7. This block becomes the repeating unit in a deep stack:
# Conceptual import path as implemented in Phase 10
from phases.07_transformers_deep_dive.code.transformer_block import TransformerBlock
class LLM:
def __init__(self, vocab_size: int, n_layers: int = 12):
self.layers = [
TransformerBlock(d_model=512, n_head=8)
for _ in range(n_layers)
]
Adding Tokenization and Language Modeling Heads
While Phase 7 operates on abstract tensors, Phase 10 introduces phases/io-llm-from-scratch/code/tokenizer.py to map raw text to token IDs. After processing through the imported transformer blocks, the model applies a language modeling head—a linear projection from hidden states to vocabulary logits—to predict the next token in the sequence.
The Training Loop and Scaling
Phase 10 implements the training infrastructure—data loaders, loss computation, and optimization loops—that feeds massive text corpora through the Phase 7 transformer blocks. It also adds scaling optimizations like gradient checkpointing, but these still operate on the same TransformerBlock API established in Phase 7.
Code Architecture: From Block to Model
The dependency chain between Phase 7 and Phase 10 demonstrates a clean separation between algorithmic implementation and systems engineering. Phase 7 focuses on mathematical correctness via unit tests in phases/07-transformers-deep-dive/tests/, while Phase 10 focuses on assembly and scale.
Here is the architectural relationship in practice, showing how Phase 10 composes Phase 7 components:
import torch
from phases.07_transformers_deep_dive.code.transformer_block import TransformerBlock
from phases.io_llm_from_scratch.code.tokenizer import SimpleTokenizer
class LLM(torch.nn.Module):
def __init__(self, vocab_size: int, d_model: int = 512, n_layers: int = 12):
super().__init__()
self.tokenizer = SimpleTokenizer(vocab_size)
self.embed = torch.nn.Embedding(vocab_size, d_model)
# Phase 7 components stacked here
self.transformer_layers = torch.nn.ModuleList([
TransformerBlock(d_model=d_model, n_head=8)
for _ in range(n_layers)
])
# Phase 10 specific: Language modeling head
self.lm_head = torch.nn.Linear(d_model, vocab_size)
def forward(self, text: str):
# Phase 10: Text to tokens
ids = self.tokenizer.encode(text)
x = self.embed(ids) # (seq_len, d_model)
# Phase 7: Process through transformer blocks
for layer in self.transformer_layers:
x = layer(x) # Uses Phase 7 self-attention
# Phase 10: Project to vocabulary
logits = self.lm_head(x) # (seq_len, vocab_size)
return logits
This architecture ensures that improvements to the attention mechanism in phases/07-transformers-deep-dive/code/attention.py automatically propagate to the full LLM in Phase 10, maintaining a single source of truth for the core transformer logic.
Summary
- Phase 7 implements the mathematical foundation:
TransformerBlock, multi-head attention, and positional encodings inphases/07-transformers-deep-dive/code/. - Phase 10 imports these blocks via
phases/io-llm-from-scratch/code/llm.pyand adds tokenization, language modeling heads, and training loops. - The relationship is compositional: Phase 10 stacks Phase 7 blocks into deep architectures without modifying the underlying attention implementations.
- Unit tests in Phase 7 validate the primitives before they are scaled up in Phase 10.
Frequently Asked Questions
Why doesn't Phase 10 reimplement the attention mechanism?
Phase 10 treats the transformer as a solved primitive. By importing the validated TransformerBlock from phases/07-transformers-deep-dive/code/transformer_block.py, Phase 10 avoids code duplication and inherits the mathematical guarantees established by Phase 7's unit tests. This separation allows Phase 10 to focus on systems challenges like tokenization and distributed training.
Can I use Phase 7 blocks for tasks other than language modeling?
Yes. The TransformerBlock class in Phase 7 is agnostic to the specific task. It processes tensors of shape (seq_len, d_model) and can be used for machine translation, time-series forecasting, or any sequence-to-sequence task. Phase 10 specifically adds the language modeling head and causal masking required for autoregressive text generation.
How does the tokenizer in Phase 10 connect to Phase 7's requirements?
Phase 7 expects numerical input tensors, while raw text requires tokenization. Phase 10's phases/io-llm-from-scratch/code/tokenizer.py bridges this gap by converting strings to integer indices. These indices are then embedded into vectors that match the d_model dimension expected by Phase 7's transformer blocks, creating a clean interface between text processing and tensor computation.
What happens if I modify the attention mechanism in Phase 7?
Any changes to phases/07-transformers-deep-dive/code/attention.py or multi_head.py automatically affect Phase 10's LLM because Phase 10 imports these modules directly. This design enforces consistency across the curriculum, ensuring that the LLM always uses the most rigorously tested version of the transformer primitives.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →