# How to Configure Transformer Hyperparameters for Different Model Sizes in train-llm-from-scratch

> Learn to configure transformer hyperparameters for various model sizes in train-llm-from-scratch. Adjust embedding dimensions, attention heads, and layer counts for optimal scaling from small to large transformers.

- Repository: [Fareed Khan/train-llm-from-scratch](https://github.com/FareedKhan-dev/train-llm-from-scratch)
- Tags: how-to-guide
- Published: 2026-05-31

---

**All architectural and training hyperparameters in the train-llm-from-scratch repository are centralized in [`config/config.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/config/config.py), allowing you to scale from tiny 10M parameter models to large 1B+ parameter transformers by adjusting embedding dimensions, attention heads, layer counts, and corresponding training settings.**

The *train-llm-from-scratch* repository by FareedKhan-dev provides a minimal, educational implementation of a GPT-style transformer where every aspect of model sizing is controlled through a single configuration dictionary. Understanding how to tune these parameters lets you smoothly transition from quick CPU prototyping to large-scale GPU training. This guide walks you through the specific configuration keys, scaling rules, and code patterns used to customize transformer hyperparameters for different model sizes.

## Core Architectural Hyperparameters

The transformer architecture is defined by five critical parameters stored in [`config/config.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/config/config.py) lines **4‑8**. These values are passed directly to the `Transformer` class constructor in [`src/models/transformer.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/models/transformer.py) (lines **20‑37**) and determine the model's capacity and memory footprint.

### Embedding Dimensions and Attention Heads

**`N_EMBED`** sets the dimensionality of token and positional embeddings, while **`N_HEAD`** controls the number of parallel attention heads per layer.

- **Scaling rule**: Double `N_EMBED` when moving between model tiers: 512 for tiny, 1024 for medium, 2048 for large.
- **Divisibility constraint**: `N_HEAD` must divide `N_EMBED` evenly. The repository follows the convention `N_HEAD = N_EMBED // 64`, resulting in 64-dimensional head sizes.
- **Parameter impact**: Doubling `N_EMBED` roughly quadruples parameters due to the feed-forward network (hidden dimension `4 * N_EMBED`) and projection matrices.

### Model Depth and Context Length

**`N_BLOCKS`** (referenced as `n_blocks` in the config dict) defines the number of transformer layers, while **`CONTEXT_LENGTH`** sets the maximum sequence length.

- **Depth scaling**: Small models use 6‑12 blocks, medium models 24‑36, and large models 48‑64. Each block adds approximately `12 * N_EMBED^2` parameters.
- **Context considerations**: The repository defaults to 512, but you can increase to 1024‑2048 for long-document tasks. Note that memory scales quadratically with context length due to the attention mechanism.

In [`config/config.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/config/config.py), these architectural values are defined at module level:

```python

# config/config.py (lines 4-8)

VOCAB_SIZE = 50304      # Fixed by tokenizer

CONTEXT_LENGTH = 512
N_EMBED = 1024
N_HEAD = 16             # 1024 / 16 = 64 per head

N_BLOCKS = 24

```

## Training Hyperparameters for Scale

Training dynamics are controlled separately from architecture via parameters defined in [`config/config.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/config/config.py) lines **14‑23**. These are consumed by the training loop in [`scripts/train_transformer.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/scripts/train_transformer.py) (lines **27‑33**, **90‑118**) and must scale inversely with model size to maintain stable optimization.

### Batch Size and Training Steps

**`T_BATCH_SIZE`** and **`T_TRAIN_STEPS`** require inverse scaling: larger models need smaller per-device batches but more optimization steps.

- **Tiny models**: Batch size 16‑32, 50k steps
- **Medium models**: Batch size 32‑64, 200k steps  
- **Large models**: Batch size 64+ (with gradient accumulation), 500k‑1M steps

**`T_CONTEXT_LENGTH`** (the actual sequence length fed during training) can be set lower than `CONTEXT_LENGTH` for initial experiments to reduce memory pressure, then increased to the full context for final training runs.

### Learning Rate Scheduling

The repository implements step-wise decay via **`T_LR`** and **`T_LR_DECAYED`**, with the drop point controlled by **`T_LR_DECAY_STEP`**.

- **Initial rate**: `5e-4` works across most model sizes.
- **Decay timing**: Set `T_LR_DECAY_STEP` to ½‑¾ of `T_TRAIN_STEPS` (e.g., 100k for a 200k step run).
- **Final rate**: Decay to `5e-5` for fine-grained convergence on large models.

## Configuration Recipes for Different Model Sizes

Use these concrete scaling recipes when editing [`config/config.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/config/config.py) to target specific parameter counts:

| Model Size | `N_EMBED` | `N_HEAD` | `N_BLOCKS` | `T_BATCH_SIZE` | `T_TRAIN_STEPS` | `T_LR_DECAY_STEP` | Est. Parameters |
|------------|-----------|----------|------------|----------------|-----------------|-------------------|-----------------|
| **Tiny** | 512 | 8 | 6 | 16 | 50,000 | 25,000 | ~10M |
| **Medium** | 1024 | 16 | 24 | 32 | 200,000 | 100,000 | ~300M |
| **Large** | 2048 | 32 | 48 | 64* | 500,000 | 250,000 | ~1B+ |

\*Use gradient accumulation if GPU memory limits force you below the target batch size.

To implement the medium configuration, modify [`config/config.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/config/config.py) as follows:

```python

# config/config.py - Medium Model Configuration

N_EMBED = 1024
N_HEAD = 16
N_BLOCKS = 24
CONTEXT_LENGTH = 1024

# Training settings

T_BATCH_SIZE = 32
T_CONTEXT_LENGTH = 1024
T_TRAIN_STEPS = 200_000
T_LR = 5e-4
T_LR_DECAY_STEP = 100_000
T_LR_DECAYED = 5e-5

```

The training script automatically picks up these values when constructing the model in [`scripts/train_transformer.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/scripts/train_transformer.py) lines **13‑19**:

```python

# scripts/train_transformer.py

from config.config import default_config as config

model = Transformer(
    n_head=config['n_head'],
    n_embed=config['n_embed'],
    context_length=config['context_length'],
    vocab_size=config['vocab_size'],
    N_BLOCKS=config['n_blocks']
).to(config['device'])

```

## Programmatic Configuration Management

Rather than editing [`config.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/config.py) directly, create derived configuration dictionaries for different experiments. This preserves the base defaults while enabling A/B testing between model sizes:

```python

# experiment_configs.py

from config.config import default_config as base_cfg

def get_config(size: str):
    cfg = base_cfg.copy()
    
    if size == "tiny":
        cfg.update({
            'n_embed': 512,
            'n_head': 8,
            'n_blocks': 6,
            't_batch_size': 16,
            't_train_steps': 50_000,
            't_lr_decay_step': 25_000
        })
    elif size == "large":
        cfg.update({
            'n_embed': 2048,
            'n_head': 32,
            'n_blocks': 48,
            't_batch_size': 64,
            't_train_steps': 500_000,
            't_lr_decay_step': 250_000,
            'context_length': 1024
        })
    
    return cfg

# Usage in training script

config = get_config("large")
model = Transformer(
    n_head=config['n_head'],
    n_embed=config['n_embed'],
    context_length=config['context_length'],
    vocab_size=config['vocab_size'],
    N_BLOCKS=config['n_blocks']
).to(config['device'])

```

## Summary

- **Central configuration**: All hyperparameters live in [`config/config.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/config/config.py), separating architecture definition from training logic.
- **Scaling dimensions**: Increase `N_EMBED` (512→1024→2048), `N_HEAD` (8→16→32), and `N_BLOCKS` (6→24→48) to scale from 10M to 1B+ parameters.
- **Training adjustments**: Larger models require more steps (50k→500k) and careful learning rate decay timing, with batch sizes adjusted for GPU memory constraints.
- **Implementation pattern**: Import `default_config`, modify specific keys, and pass the dictionary to the `Transformer` constructor in [`src/models/transformer.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/models/transformer.py).

## Frequently Asked Questions

### What file contains all transformer hyperparameters in train-llm-from-scratch?

All hyperparameters are defined in **[`config/config.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/config/config.py)**. Lines **4‑8** control architectural dimensions (`N_EMBED`, `N_HEAD`, `N_BLOCKS`, `CONTEXT_LENGTH`), while lines **14‑23** define training settings (`T_BATCH_SIZE`, `T_TRAIN_STEPS`, learning rate schedules). The training script imports this dictionary at runtime to construct the model and configure the optimizer.

### How do I scale from a 10M to a 300M parameter model?

Increase `N_EMBED` from 512 to 1024, `N_HEAD` from 8 to 16 (maintaining 64-dimension per head), and `N_BLOCKS` from 6 to 24. Update training parameters: set `T_TRAIN_STEPS` to 200,000, `T_BATCH_SIZE` to 32, and `T_LR_DECAY_STEP` to 100,000. These changes, implemented in [`config/config.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/config/config.py), expand the model width and depth while extending training time for the larger capacity.

### Why must `N_HEAD` divide `N_EMBED` evenly?

The transformer splits the embedding dimension into equal chunks for each attention head. In the repository's implementation ([`src/models/transformer.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/models/transformer.py)), the head dimension is calculated as `n_embed // n_head`. If `N_EMBED` is not divisible by `N_HEAD`, the matrix reshaping operations in the multi-head attention mechanism will raise runtime errors. Standard configurations use `N_HEAD = N_EMBED // 64`, ensuring head dimensions of 64.

### How do I handle memory constraints when increasing model size?

Reduce **`T_BATCH_SIZE`** or use gradient accumulation to maintain an effective batch size while fitting in GPU memory. Additionally, you can temporarily lower **`T_CONTEXT_LENGTH`** below `CONTEXT_LENGTH` during initial training phases. For the largest configurations, consider reducing `N_BLOCKS` slightly or using mixed-precision training, though the repository defaults to standard float32 tensors.