Transformer Architecture for LLM Pretraining: Core Components Explained
A Transformer architecture for LLM pretraining comprises token embeddings, positional encodings, multi-head self-attention, feed-forward networks, layer normalization, and residual connections, organized into stacked blocks and finalized with a language modeling head, as implemented in the train-llm-from-scratch repository.
The train-llm-from-scratch repository provides a ground-up implementation of the Transformer architecture for LLM pretraining. Understanding these fundamental building blocks—from the initial embedding layers to the optional reinforcement learning heads—is essential for training large language models effectively and debugging the training pipeline.
Core Building Blocks
The Transformer class in src/models/transformer.py implements the canonical architecture described in Attention Is All You Need. Each component serves a specific function in processing token sequences and is configurable via the constructor parameters.
Token and Positional Embeddings
The architecture begins with an embedding layer that maps discrete token IDs to dense vectors of dimension n_embed. This repository implements learned positional embeddings rather than fixed sinusoidal encodings, injecting sequence order information through a trainable matrix initialized within the Transformer class. During the forward pass, token and positional embeddings are summed to create the initial input representations fed into the first Transformer block.
Multi-Head Self-Attention (MHSA)
The multi-head self-attention mechanism allows each token to attend to every other token in the sequence, capturing long-range dependencies in parallel. The number of attention heads is controlled by the n_head parameter. As referenced in tests/verify_rl_optimizes.py, test configurations typically use n_head=4, while production training in scripts/train_transformer.py scales this to n_head=16 for increased model capacity.
Feed-Forward Networks (FFN)
After attention, each token representation passes through a feed-forward network consisting of two linear transformations with a non-linear activation. The hidden dimension is typically set to 4 × n_embed, following standard Transformer design. This network processes each position independently, adding non-linearity to the model's representations before passing them to the next block.
Layer Normalization and Residual Connections
Layer normalization stabilizes training by normalizing inputs across the feature dimension, applied at the start of each sub-layer within the Transformer blocks. Residual connections wrap around both the attention and feed-forward sub-layers, enabling gradient flow by adding the sub-layer input to its output. These mechanisms are critical for training deep networks with many blocks without suffering from vanishing gradients.
Stacked Transformer Blocks and the Language Model Head
The core architecture repeats the attention and feed-forward units n_blocks times, creating a deep stack of identical layers. For example, tests/verify_rl_optimizes.py constructs models with n_blocks=2 for rapid testing, while scripts/train_transformer.py configures deeper stacks such as n_blocks=24 for full-scale pretraining. The final language model head projects the last hidden state back to the vocabulary space, producing logits for next-token prediction as utilized in scripts/generate_text.py.
Optional RL Extensions
Beyond standard pretraining, the repository supports reinforcement learning fine-tuning through additional heads attached to the frozen or fine-tuned base Transformer.
Value Heads for PPO
The TransformerWithValueHead class in src/post_training/value_head.py wraps the base Transformer and adds a scalar value head. This outputs a state value estimate required for Proximal Policy Optimization (PPO) training, returning both the language modeling logits and a value scalar during the forward pass to compute advantages.
Reward Models for Preference Optimization
For preference-based fine-tuning methods like DPO (Direct Preference Optimization) or GRPO, the repository implements reward modeling capabilities in src/post_training/reward_model.py. These scalar outputs evaluate response quality, enabling the model to learn from human preferences or automated reward signals attached to the base Transformer backbone.
Instantiating the Transformer
According to src/post_training/utils.py, the model is constructed by passing a configuration dictionary to the Transformer constructor, defining the architecture's hyperparameters.
# Excerpt from src/post_training/utils.py
from src.models.transformer import Transformer
model = Transformer(
n_head=cfg["n_head"],
n_embed=cfg["n_embed"],
context_length=cfg["context_length"],
vocab_size=cfg["vocab_size"],
n_blocks=cfg["n_blocks"],
)
Training and Generation Examples
Pretraining utilizes the standard language modeling objective with cross-entropy loss. The scripts/train_transformer.py file demonstrates creating a full-scale model and optimizing it with AdamW.
# From scripts/train_transformer.py
from src.models.transformer import Transformer
import torch
cfg = {
"n_head": 16,
"n_embed": 1024,
"context_length": 1024,
"vocab_size": 256,
"n_blocks": 24,
}
model = Transformer(**cfg).to("cuda")
optimizer = torch.optim.AdamW(model.parameters(), lr=1e-4)
For inference, scripts/generate_text.py loads a checkpoint and generates text using the language model head.
# From scripts/generate_text.py
from src.models.transformer import Transformer
import torch
checkpoint = torch.load("ckpt.pt")
cfg = checkpoint["config"]
model = Transformer(**cfg).to("cpu")
model.load_state_dict(checkpoint["model"])
# Generate tokens using the LM head
prompt = torch.tensor([[1, 2, 3]]) # tokenized input
generated = model.generate(prompt, max_new_tokens=50)
Summary
- The Transformer class in
src/models/transformer.pyimplements the core architecture with embeddings, multi-head self-attention, feed-forward networks, and layer normalization. - Residual connections and stacked blocks (
n_blocks) enable deep architectures ranging from 2 layers in unit tests to 24+ in production training. - The language model head projects hidden states to vocabulary logits for next-token prediction during pretraining and text generation.
- Optional value heads (
src/post_training/value_head.py) and reward models (src/post_training/reward_model.py) extend the architecture for reinforcement learning fine-tuning. - Configuration parameters including
n_head,n_embed, andvocab_sizedefine the model's capacity and are passed during instantiation viasrc/post_training/utils.py.
Frequently Asked Questions
What do the n_head and n_embed parameters control?
The n_embed parameter defines the dimensionality of the embedding vectors and hidden states throughout the model, while n_head specifies the number of parallel attention heads in the multi-head self-attention mechanism. The repository configures these via dictionaries passed to the Transformer constructor, with typical values being n_embed=1024 and n_head=16 for training runs in scripts/train_transformer.py.
How does the repository handle positional encoding?
The implementation uses learned positional embeddings rather than fixed sinusoidal encodings. These embeddings are added to the token embeddings at the input layer, allowing the model to learn positional information during training. This approach is integrated within the Transformer class initialization and requires no additional input during the forward pass.
What is the purpose of the TransformerWithValueHead class?
Located in src/post_training/value_head.py, this wrapper class attaches a scalar value head to the base Transformer backbone. It enables Proximal Policy Optimization (PPO) by returning both the language modeling logits and a value estimate for the current state, which is necessary for computing advantages and updating the policy during RL fine-tuning.
Can I modify the vocabulary size for domain-specific training?
Yes, the vocab_size parameter passed to the Transformer constructor determines the size of the embedding table and the output dimension of the language model head. The repository demonstrates this with vocab_size=256 in training scripts, but this can be adjusted to match any custom tokenizer's vocabulary size for domain-specific pretraining.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →