How to Configure Different LLM Sizes Using Hyperparameters: n_embed, n_head, and n_blocks

You can configure LLM model sizes in the train-llm-from-scratch repository by adjusting three core hyperparameters—n_embed (embedding dimension), n_head (attention heads), and n_blocks (transformer layers)—through Python dictionaries, JSON configuration files, or command-line overrides.

The FareedKhan-dev/train-llm-from-scratch repository separates pre-training and post-training pipelines, but both stages expose identical architectural size knobs. Understanding how to manipulate n_embed, n_head, and n_blocks allows you to scale models from tiny CPU-testable configurations to large production-scale transformers without modifying the underlying architecture code.

Understanding the Core Hyperparameters

Three values control the transformer architecture's capacity and computational requirements. Each parameter directly impacts the model's parameter count and memory footprint.

n_embed (Embedding Dimension)

n_embed defines the hidden size of the model, controlling the width of token embeddings and the dimensionality of each transformer's feed-forward network. In config/config.py, this is defined as N_EMBED (default 2048 for pre-training), while config/post_training_config.py uses n_embed (default 1024 for post-training stages). Larger values increase model capacity but require significantly more memory and compute.

n_head (Number of Attention Heads)

n_head specifies the number of parallel attention mechanisms operating in each transformer block. The repository defaults to 16 heads across both configurations. A critical constraint is that n_embed must be divisible by n_head to ensure even splitting of the embedding dimension across heads (e.g., 1024 dimensions with 16 heads yields 64 dimensions per head).

n_blocks (Number of Transformer Blocks)

n_blocks determines the model's depth by controlling how many transformer layers are stacked. Pre-training defaults to 24 blocks in config/config.py, while post-training configurations inherit this value through BaseModelConfig.n_blocks. Increasing this value improves the model's ability to capture long-range dependencies but linearly increases training time.

Where These Hyperparameters Live in the Codebase

The repository maintains separate configuration systems for pre-training and post-training stages, though both expose the same three size parameters.

Pre-training Configuration (config/config.py)

For pre-training, hyperparameters reside in a Python dictionary within config/config.py. The scripts/train_transformer.py file instantiates the model by passing these values directly to the Transformer class:


# From scripts/train_transformer.py

model = Transformer(
    n_head=config['n_head'],
    n_embed=config['n_embed'],
    context_length=config['context_length'],
    vocab_size=config['vocab_size'],
    N_BLOCKS=config['n_blocks']
)

The default_config dictionary in config/config.py contains the default assignments for N_EMBED, N_HEAD, and N_BLOCKS.

Post-training Configuration (post_training_config.py)

Post-training stages (SFT, DPO, PPO, GRPO) use a dataclass-based configuration system defined in config/post_training_config.py. The BaseModelConfig class holds the architectural parameters:

@dataclass
class BaseModelConfig:
    n_embed: int = 1024
    n_head: int = 16
    n_blocks: int = 24
    # ... other fields

Configuration resolution follows a hierarchy: dataclass defaults merge with configs/base.json, then stage-specific JSON files (e.g., configs/sft.json), and finally CLI overrides.

How to Modify LLM Sizes

You can adjust model sizes through four distinct methods depending on your workflow stage.

Editing Pre-training Parameters

To change the model size for pre-training, modify the constants in config/config.py:


# config/config.py

N_EMBED = 1536      # Must be divisible by N_HEAD

N_HEAD = 12         # Number of attention heads

N_BLOCKS = 18       # Number of transformer layers

These changes take effect immediately when running scripts/train_transformer.py.

Configuring Post-training via JSON

For post-training stages, create or edit JSON files in the configs/ directory. The loader.load_config function merges these with base settings:

{
  "n_embed": 1536,
  "n_head": 12,
  "n_blocks": 18,
  "lr": 2e-5,
  "batch_size": 32
}

Save this as configs/sft.json or similar, and the training script will resolve these values through the configuration inheritance chain.

Using CLI Overrides

When launching post-training scripts, override any architectural parameter using the --field flag:

python -m src.post_training.cli train_sft \
    --config configs/sft.json \
    --field n_embed=2048 \
    --field n_head=16 \
    --field n_blocks=24

The src/post_training/cli.py entry point passes these overrides to loader.load_config, which gives precedence to command-line values over JSON configurations.

Programmatic Configuration with the Smoke Helper

For rapid experimentation or CPU testing, use the smoke helper function to generate minimal configurations:

from config.post_training_config import SFTConfig, smoke

# Creates a tiny model: n_embed=64, n_head=4, n_blocks=2

tiny_cfg = smoke(SFTConfig)
print(tiny_cfg.n_embed)  # Output: 64

This approach is ideal for debugging pipeline logic without loading multi-billion-parameter models.

Parameter Count Estimation

The total parameter count scales approximately according to the formula:


params ≈ n_blocks * (4 * n_embed^2 + 2 * n_head * (n_embed / n_head)^2) + vocab * n_embed

This quadratic relationship with n_embed and linear relationship with n_blocks explains why increasing embedding dimensions yields more powerful models but requires substantially more GPU memory. Always ensure n_embed is divisible by n_head to avoid runtime errors during tensor reshaping operations.

Summary

  • Three core knobs control LLM size: n_embed (width), n_head (parallel attention), and n_blocks (depth).
  • Pre-training uses config/config.py with dictionary keys N_EMBED, N_HEAD, and N_BLOCKS consumed by scripts/train_transformer.py.
  • Post-training uses dataclass-based configs in config/post_training_config.py, resolved via JSON files and CLI overrides through loader.load_config.
  • CLI overrides allow rapid experimentation using --field parameter=value syntax when launching training scripts.
  • Validation rule: n_embed must be divisible by n_head to ensure proper head dimension allocation.

Frequently Asked Questions

What happens if n_embed is not divisible by n_head?

The model will raise a runtime error during the attention mechanism's forward pass. The train-llm-from-scratch implementation assumes even division to split the embedding dimension into n_head separate attention heads. For example, with n_embed=1024, valid n_head values include 2, 4, 8, 16, 32, or 64.

Can I change model size between pre-training and post-training stages?

Yes, but you must ensure the checkpoint conversion handles dimension mismatches. Typically, you would pre-train with one configuration (e.g., n_embed=2048) and then adjust the post-training configuration JSON to match the saved checkpoint dimensions, or implement interpolation/trimming logic if switching sizes between stages.

How do I quickly test my pipeline with a small model?

Import the smoke function from config.post_training_config and pass your stage configuration class (e.g., PPOConfig, SFTConfig). This generates a tiny configuration with n_embed=64, n_head=4, and n_blocks=2, allowing you to verify data loading and training logic on CPU without loading large weights.

Where are the default values for post-training defined?

Default architectural values for post-training reside in BaseModelConfig within config/post_training_config.py. These are then potentially overridden by configs/base.json and stage-specific files like configs/sft.json. The loader.load_config function manages this merge hierarchy, with CLI arguments taking final precedence.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →