# Transformer Architecture for LLM Pretraining: Core Components Explained

> Explore the core components of Transformer architecture for LLM pretraining including self-attention and feed-forward networks. Learn how they power large language models.

- Repository: [Fareed Khan/train-llm-from-scratch](https://github.com/FareedKhan-dev/train-llm-from-scratch)
- Tags: deep-dive
- Published: 2026-06-11

---

**A Transformer architecture for LLM pretraining comprises token embeddings, positional encodings, multi-head self-attention, feed-forward networks, layer normalization, and residual connections, organized into stacked blocks and finalized with a language modeling head, as implemented in the `train-llm-from-scratch` repository.**

The `train-llm-from-scratch` repository provides a ground-up implementation of the Transformer architecture for LLM pretraining. Understanding these fundamental building blocks—from the initial embedding layers to the optional reinforcement learning heads—is essential for training large language models effectively and debugging the training pipeline.

## Core Building Blocks

The `Transformer` class in [`src/models/transformer.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/models/transformer.py) implements the canonical architecture described in *Attention Is All You Need*. Each component serves a specific function in processing token sequences and is configurable via the constructor parameters.

### Token and Positional Embeddings

The architecture begins with an **embedding layer** that maps discrete token IDs to dense vectors of dimension `n_embed`. This repository implements **learned positional embeddings** rather than fixed sinusoidal encodings, injecting sequence order information through a trainable matrix initialized within the `Transformer` class. During the forward pass, token and positional embeddings are summed to create the initial input representations fed into the first Transformer block.

### Multi-Head Self-Attention (MHSA)

The **multi-head self-attention** mechanism allows each token to attend to every other token in the sequence, capturing long-range dependencies in parallel. The number of attention heads is controlled by the `n_head` parameter. As referenced in [`tests/verify_rl_optimizes.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/tests/verify_rl_optimizes.py), test configurations typically use `n_head=4`, while production training in [`scripts/train_transformer.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/scripts/train_transformer.py) scales this to `n_head=16` for increased model capacity.

### Feed-Forward Networks (FFN)

After attention, each token representation passes through a **feed-forward network** consisting of two linear transformations with a non-linear activation. The hidden dimension is typically set to `4 × n_embed`, following standard Transformer design. This network processes each position independently, adding non-linearity to the model's representations before passing them to the next block.

### Layer Normalization and Residual Connections

**Layer normalization** stabilizes training by normalizing inputs across the feature dimension, applied at the start of each sub-layer within the `Transformer` blocks. **Residual connections** wrap around both the attention and feed-forward sub-layers, enabling gradient flow by adding the sub-layer input to its output. These mechanisms are critical for training deep networks with many blocks without suffering from vanishing gradients.

### Stacked Transformer Blocks and the Language Model Head

The core architecture repeats the attention and feed-forward units `n_blocks` times, creating a deep stack of identical layers. For example, [`tests/verify_rl_optimizes.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/tests/verify_rl_optimizes.py) constructs models with `n_blocks=2` for rapid testing, while [`scripts/train_transformer.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/scripts/train_transformer.py) configures deeper stacks such as `n_blocks=24` for full-scale pretraining. The final **language model head** projects the last hidden state back to the vocabulary space, producing logits for next-token prediction as utilized in [`scripts/generate_text.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/scripts/generate_text.py).

## Optional RL Extensions

Beyond standard pretraining, the repository supports reinforcement learning fine-tuning through additional heads attached to the frozen or fine-tuned base Transformer.

### Value Heads for PPO

The `TransformerWithValueHead` class in [`src/post_training/value_head.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/value_head.py) wraps the base Transformer and adds a scalar value head. This outputs a state value estimate required for Proximal Policy Optimization (PPO) training, returning both the language modeling logits and a value scalar during the forward pass to compute advantages.

### Reward Models for Preference Optimization

For preference-based fine-tuning methods like DPO (Direct Preference Optimization) or GRPO, the repository implements reward modeling capabilities in [`src/post_training/reward_model.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/reward_model.py). These scalar outputs evaluate response quality, enabling the model to learn from human preferences or automated reward signals attached to the base Transformer backbone.

## Instantiating the Transformer

According to [`src/post_training/utils.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/utils.py), the model is constructed by passing a configuration dictionary to the `Transformer` constructor, defining the architecture's hyperparameters.

```python

# Excerpt from src/post_training/utils.py

from src.models.transformer import Transformer

model = Transformer(
    n_head=cfg["n_head"],
    n_embed=cfg["n_embed"],
    context_length=cfg["context_length"],
    vocab_size=cfg["vocab_size"],
    n_blocks=cfg["n_blocks"],
)

```

## Training and Generation Examples

Pretraining utilizes the standard language modeling objective with cross-entropy loss. The [`scripts/train_transformer.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/scripts/train_transformer.py) file demonstrates creating a full-scale model and optimizing it with AdamW.

```python

# From scripts/train_transformer.py

from src.models.transformer import Transformer
import torch

cfg = {
    "n_head": 16,
    "n_embed": 1024,
    "context_length": 1024,
    "vocab_size": 256,
    "n_blocks": 24,
}
model = Transformer(**cfg).to("cuda")
optimizer = torch.optim.AdamW(model.parameters(), lr=1e-4)

```

For inference, [`scripts/generate_text.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/scripts/generate_text.py) loads a checkpoint and generates text using the language model head.

```python

# From scripts/generate_text.py

from src.models.transformer import Transformer
import torch

checkpoint = torch.load("ckpt.pt")
cfg = checkpoint["config"]
model = Transformer(**cfg).to("cpu")
model.load_state_dict(checkpoint["model"])

# Generate tokens using the LM head

prompt = torch.tensor([[1, 2, 3]])  # tokenized input

generated = model.generate(prompt, max_new_tokens=50)

```

## Summary

- The **Transformer** class in [`src/models/transformer.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/models/transformer.py) implements the core architecture with embeddings, multi-head self-attention, feed-forward networks, and layer normalization.
- **Residual connections** and **stacked blocks** (`n_blocks`) enable deep architectures ranging from 2 layers in unit tests to 24+ in production training.
- The **language model head** projects hidden states to vocabulary logits for next-token prediction during pretraining and text generation.
- Optional **value heads** ([`src/post_training/value_head.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/value_head.py)) and **reward models** ([`src/post_training/reward_model.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/reward_model.py)) extend the architecture for reinforcement learning fine-tuning.
- Configuration parameters including `n_head`, `n_embed`, and `vocab_size` define the model's capacity and are passed during instantiation via [`src/post_training/utils.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/utils.py).

## Frequently Asked Questions

### What do the n_head and n_embed parameters control?

The `n_embed` parameter defines the dimensionality of the embedding vectors and hidden states throughout the model, while `n_head` specifies the number of parallel attention heads in the multi-head self-attention mechanism. The repository configures these via dictionaries passed to the `Transformer` constructor, with typical values being `n_embed=1024` and `n_head=16` for training runs in [`scripts/train_transformer.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/scripts/train_transformer.py).

### How does the repository handle positional encoding?

The implementation uses **learned positional embeddings** rather than fixed sinusoidal encodings. These embeddings are added to the token embeddings at the input layer, allowing the model to learn positional information during training. This approach is integrated within the `Transformer` class initialization and requires no additional input during the forward pass.

### What is the purpose of the TransformerWithValueHead class?

Located in [`src/post_training/value_head.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/post_training/value_head.py), this wrapper class attaches a scalar value head to the base Transformer backbone. It enables Proximal Policy Optimization (PPO) by returning both the language modeling logits and a value estimate for the current state, which is necessary for computing advantages and updating the policy during RL fine-tuning.

### Can I modify the vocabulary size for domain-specific training?

Yes, the `vocab_size` parameter passed to the `Transformer` constructor determines the size of the embedding table and the output dimension of the language model head. The repository demonstrates this with `vocab_size=256` in training scripts, but this can be adjusted to match any custom tokenizer's vocabulary size for domain-specific pretraining.