How to Configure Transformer Hyperparameters for Different Model Sizes in train-llm-from-scratch
All architectural and training hyperparameters in the train-llm-from-scratch repository are centralized in config/config.py, allowing you to scale from tiny 10M parameter models to large 1B+ parameter transformers by adjusting embedding dimensions, attention heads, layer counts, and corresponding training settings.
The train-llm-from-scratch repository by FareedKhan-dev provides a minimal, educational implementation of a GPT-style transformer where every aspect of model sizing is controlled through a single configuration dictionary. Understanding how to tune these parameters lets you smoothly transition from quick CPU prototyping to large-scale GPU training. This guide walks you through the specific configuration keys, scaling rules, and code patterns used to customize transformer hyperparameters for different model sizes.
Core Architectural Hyperparameters
The transformer architecture is defined by five critical parameters stored in config/config.py lines 4‑8. These values are passed directly to the Transformer class constructor in src/models/transformer.py (lines 20‑37) and determine the model's capacity and memory footprint.
Embedding Dimensions and Attention Heads
N_EMBED sets the dimensionality of token and positional embeddings, while N_HEAD controls the number of parallel attention heads per layer.
- Scaling rule: Double
N_EMBEDwhen moving between model tiers: 512 for tiny, 1024 for medium, 2048 for large. - Divisibility constraint:
N_HEADmust divideN_EMBEDevenly. The repository follows the conventionN_HEAD = N_EMBED // 64, resulting in 64-dimensional head sizes. - Parameter impact: Doubling
N_EMBEDroughly quadruples parameters due to the feed-forward network (hidden dimension4 * N_EMBED) and projection matrices.
Model Depth and Context Length
N_BLOCKS (referenced as n_blocks in the config dict) defines the number of transformer layers, while CONTEXT_LENGTH sets the maximum sequence length.
- Depth scaling: Small models use 6‑12 blocks, medium models 24‑36, and large models 48‑64. Each block adds approximately
12 * N_EMBED^2parameters. - Context considerations: The repository defaults to 512, but you can increase to 1024‑2048 for long-document tasks. Note that memory scales quadratically with context length due to the attention mechanism.
In config/config.py, these architectural values are defined at module level:
# config/config.py (lines 4-8)
VOCAB_SIZE = 50304 # Fixed by tokenizer
CONTEXT_LENGTH = 512
N_EMBED = 1024
N_HEAD = 16 # 1024 / 16 = 64 per head
N_BLOCKS = 24
Training Hyperparameters for Scale
Training dynamics are controlled separately from architecture via parameters defined in config/config.py lines 14‑23. These are consumed by the training loop in scripts/train_transformer.py (lines 27‑33, 90‑118) and must scale inversely with model size to maintain stable optimization.
Batch Size and Training Steps
T_BATCH_SIZE and T_TRAIN_STEPS require inverse scaling: larger models need smaller per-device batches but more optimization steps.
- Tiny models: Batch size 16‑32, 50k steps
- Medium models: Batch size 32‑64, 200k steps
- Large models: Batch size 64+ (with gradient accumulation), 500k‑1M steps
T_CONTEXT_LENGTH (the actual sequence length fed during training) can be set lower than CONTEXT_LENGTH for initial experiments to reduce memory pressure, then increased to the full context for final training runs.
Learning Rate Scheduling
The repository implements step-wise decay via T_LR and T_LR_DECAYED, with the drop point controlled by T_LR_DECAY_STEP.
- Initial rate:
5e-4works across most model sizes. - Decay timing: Set
T_LR_DECAY_STEPto ½‑¾ ofT_TRAIN_STEPS(e.g., 100k for a 200k step run). - Final rate: Decay to
5e-5for fine-grained convergence on large models.
Configuration Recipes for Different Model Sizes
Use these concrete scaling recipes when editing config/config.py to target specific parameter counts:
| Model Size | N_EMBED |
N_HEAD |
N_BLOCKS |
T_BATCH_SIZE |
T_TRAIN_STEPS |
T_LR_DECAY_STEP |
Est. Parameters |
|---|---|---|---|---|---|---|---|
| Tiny | 512 | 8 | 6 | 16 | 50,000 | 25,000 | ~10M |
| Medium | 1024 | 16 | 24 | 32 | 200,000 | 100,000 | ~300M |
| Large | 2048 | 32 | 48 | 64* | 500,000 | 250,000 | ~1B+ |
*Use gradient accumulation if GPU memory limits force you below the target batch size.
To implement the medium configuration, modify config/config.py as follows:
# config/config.py - Medium Model Configuration
N_EMBED = 1024
N_HEAD = 16
N_BLOCKS = 24
CONTEXT_LENGTH = 1024
# Training settings
T_BATCH_SIZE = 32
T_CONTEXT_LENGTH = 1024
T_TRAIN_STEPS = 200_000
T_LR = 5e-4
T_LR_DECAY_STEP = 100_000
T_LR_DECAYED = 5e-5
The training script automatically picks up these values when constructing the model in scripts/train_transformer.py lines 13‑19:
# scripts/train_transformer.py
from config.config import default_config as config
model = Transformer(
n_head=config['n_head'],
n_embed=config['n_embed'],
context_length=config['context_length'],
vocab_size=config['vocab_size'],
N_BLOCKS=config['n_blocks']
).to(config['device'])
Programmatic Configuration Management
Rather than editing config.py directly, create derived configuration dictionaries for different experiments. This preserves the base defaults while enabling A/B testing between model sizes:
# experiment_configs.py
from config.config import default_config as base_cfg
def get_config(size: str):
cfg = base_cfg.copy()
if size == "tiny":
cfg.update({
'n_embed': 512,
'n_head': 8,
'n_blocks': 6,
't_batch_size': 16,
't_train_steps': 50_000,
't_lr_decay_step': 25_000
})
elif size == "large":
cfg.update({
'n_embed': 2048,
'n_head': 32,
'n_blocks': 48,
't_batch_size': 64,
't_train_steps': 500_000,
't_lr_decay_step': 250_000,
'context_length': 1024
})
return cfg
# Usage in training script
config = get_config("large")
model = Transformer(
n_head=config['n_head'],
n_embed=config['n_embed'],
context_length=config['context_length'],
vocab_size=config['vocab_size'],
N_BLOCKS=config['n_blocks']
).to(config['device'])
Summary
- Central configuration: All hyperparameters live in
config/config.py, separating architecture definition from training logic. - Scaling dimensions: Increase
N_EMBED(512→1024→2048),N_HEAD(8→16→32), andN_BLOCKS(6→24→48) to scale from 10M to 1B+ parameters. - Training adjustments: Larger models require more steps (50k→500k) and careful learning rate decay timing, with batch sizes adjusted for GPU memory constraints.
- Implementation pattern: Import
default_config, modify specific keys, and pass the dictionary to theTransformerconstructor insrc/models/transformer.py.
Frequently Asked Questions
What file contains all transformer hyperparameters in train-llm-from-scratch?
All hyperparameters are defined in config/config.py. Lines 4‑8 control architectural dimensions (N_EMBED, N_HEAD, N_BLOCKS, CONTEXT_LENGTH), while lines 14‑23 define training settings (T_BATCH_SIZE, T_TRAIN_STEPS, learning rate schedules). The training script imports this dictionary at runtime to construct the model and configure the optimizer.
How do I scale from a 10M to a 300M parameter model?
Increase N_EMBED from 512 to 1024, N_HEAD from 8 to 16 (maintaining 64-dimension per head), and N_BLOCKS from 6 to 24. Update training parameters: set T_TRAIN_STEPS to 200,000, T_BATCH_SIZE to 32, and T_LR_DECAY_STEP to 100,000. These changes, implemented in config/config.py, expand the model width and depth while extending training time for the larger capacity.
Why must N_HEAD divide N_EMBED evenly?
The transformer splits the embedding dimension into equal chunks for each attention head. In the repository's implementation (src/models/transformer.py), the head dimension is calculated as n_embed // n_head. If N_EMBED is not divisible by N_HEAD, the matrix reshaping operations in the multi-head attention mechanism will raise runtime errors. Standard configurations use N_HEAD = N_EMBED // 64, ensuring head dimensions of 64.
How do I handle memory constraints when increasing model size?
Reduce T_BATCH_SIZE or use gradient accumulation to maintain an effective batch size while fitting in GPU memory. Additionally, you can temporarily lower T_CONTEXT_LENGTH below CONTEXT_LENGTH during initial training phases. For the largest configurations, consider reducing N_BLOCKS slightly or using mixed-precision training, though the repository defaults to standard float32 tensors.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →