How to Calculate Total Transformer Model Parameters: A Complete Guide for PyTorch
To calculate total transformer model parameters, sum the token embedding, positional embedding, final layer‑norm, and language‑model head weights, then add N_BLOCKS multiplied by the per‑block count of attention projections (3×n_embed²), MLP layers (8×n_embed² + 5×n_embed), and layer‑norm parameters (4×n_embed).
Managing GPU memory and estimating training costs requires knowing exactly how many trainable weights your model contains. In the FareedKhan-dev/train-llm-from-scratch repository, the vanilla Transformer is implemented from scratch in PyTorch, making it straightforward to derive a closed‑form parameter count. This guide breaks down every weight matrix and bias vector in src/models/transformer.py and its submodules so you can compute the exact parameter budget for any configuration.
Anatomy of the Transformer Parameter Count
The implementation in src/models/transformer.py assembles the model from three macro‑components: embedding tables, a stack of identical transformer blocks, and final projection layers.
Embedding Layers
The input representation combines token indices with positional information:
- Token embedding —
self.token_embed = nn.Embedding(vocab_size, n_embed)at lines 34‑35 ofsrc/models/transformer.py. This contributes vocab_size × n_embed parameters. - Positional embedding —
self.position_embed = nn.Embedding(context_length, n_embed)at lines 34‑36. This adds context_length × n_embed parameters.
Both embeddings are simple lookup tables without bias vectors.
Inside a Transformer Block
Each block (defined in src/models/transformer_block.py, lines 18‑32) contains three sub‑components whose parameters must be counted:
-
Two LayerNorms — Each
nn.LayerNorm(n_embed)stores a weight and bias vector of sizen_embed. Across the block’s pre‑attention and pre‑MLP norms, this yields 4 × n_embed parameters. -
Multi‑Head Attention — As implemented in
src/models/attention.py(lines 29‑31), every head uses three bias‑free linear projections (query,key,value). Withhead_size = n_embed // n_head, the math simplifies to 3 × n_embed² parameters per block because then_headterms cancel out. -
Feed‑Forward MLP — The MLP (located in
src/models/mlp.py) expands the hidden dimension to4 × n_embedand projects back. With defaultbias=True, the two linear layers contribute:- Hidden projection:
n_embed × (4·n_embed)weights +4·n_embedbias - Output projection:
(4·n_embed) × n_embedweights +n_embedbias
Summed together, the MLP adds 8 × n_embed² + 5 × n_embed parameters.
- Hidden projection:
Total per block:
3·n_embed² (attention) + 8·n_embed² + 5·n_embed (MLP) + 4·n_embed (norms) = 11·n_embed² + 9·n_embed.
Final Projection and Normalization
After the block stack, the model applies:
- Final LayerNorm —
self.layer_norm = nn.LayerNorm(n_embed)at lines 37‑38 contributes 2 × n_embed (weight + bias). - Language‑Model Head —
self.lm_head = nn.Linear(n_embed, vocab_size)at lines 38‑39 adds vocab_size × n_embed weights plus vocab_size bias terms.
The Exact Parameter Formula
Let V = vocab_size, E = n_embed, C = context_length, and B = N_BLOCKS. The total number of trainable parameters P is:
P = 2·V·E + V + (C + 2)·E + B·(11·E² + 9·E)
Breaking down the terms:
2·V·Ecounts the token embedding weights and the LM head weight matrix (both sizeV × E).Vaccounts for the LM head bias.(C + 2)·Ecaptures positional embeddings (C·E) plus the final layer‑norm weight and bias (2·E).B·(11·E² + 9·E)scales the per‑block parameter count across allN_BLOCKSrepetitions.
Step‑by‑Step Implementation in Python
You can implement this formula directly without instantiating the model:
def transformer_param_count(
n_head: int,
n_embed: int,
context_length: int,
vocab_size: int,
n_blocks: int,
) -> int:
"""Return the total number of trainable parameters for the model."""
# Embeddings
token_embed = vocab_size * n_embed
pos_embed = context_length * n_embed
# Final layers
final_ln = 2 * n_embed # weight + bias
lm_head_w = vocab_size * n_embed
lm_head_b = vocab_size
# One Transformer block
block_params = (
3 * n_embed ** 2 # attention (bias‑free)
+ 8 * n_embed ** 2 + 5 * n_embed # MLP (hidden + projection)
+ 4 * n_embed # two LayerNorms (weight+bias each)
)
total = (
token_embed + pos_embed
+ final_ln + lm_head_w + lm_head_b
+ n_blocks * block_params
)
return total
# Example – the hyper‑parameters used in the repo’s demo:
if __name__ == "__main__":
n_head = 4
n_embed = 32
context_len = 5
vocab = 100
n_blocks = 2
print("Total parameters:", transformer_param_count(
n_head, n_embed, context_len, vocab, n_blocks))
Running the snippet with the example values prints:
Total parameters: 115388
Summary
- Embedding tables contribute
(vocab_size + context_length) × n_embedparameters. - Each transformer block adds
11·n_embed² + 9·n_embedtrainable weights, derived from attention projections, MLP expansion, and layer‑norms. - Final layers contribute
(2 + vocab_size) × n_embed + vocab_sizeparameters for the last layer‑norm and the output projection. - Use the closed‑form formula
P = 2·V·E + V + (C + 2)·E + B·(11·E² + 9·E)to instantly estimate memory usage without loading the model into GPU RAM.
Frequently Asked Questions
Why do the attention projections use bias=False and how does that affect the count?
In src/models/attention.py (lines 29‑31), the query, key, and value projections are initialized with bias=False, which is standard for scaled dot‑product attention. This removes 3 × n_embed bias parameters per block that would otherwise appear in a typical linear layer, reducing the total count by exactly 3 × n_embed × N_BLOCKS.
How is the MLP parameter count derived as 8·n_embed² + 5·n_embed?
The MLP in src/models/mlp.py uses two linear layers with bias=True. The first expands from n_embed to 4·n_embed, contributing n_embed × 4·n_embed weights and 4·n_embed bias. The second projects back, adding 4·n_embed × n_embed weights and n_embed bias. Summing these terms gives 8·n_embed² weights + 5·n_embed bias.
Can I verify this calculation against PyTorch’s built‑in counter?
Yes. After instantiating the Transformer class from src/models/transformer.py, run:
total = sum(p.numel() for p in model.parameters() if p.requires_grad)
This should match the output of transformer_param_count() exactly when provided with the same hyper‑parameters (n_embed, vocab_size, etc.).
Does the formula change if I use weight tying between input and output embeddings?
The repository does not implement weight tying; self.lm_head is independent of self.token_embed. If you tie the weights (setting lm_head.weight = token_embed.weight), subtract vocab_size × n_embed from the total and remove the LM‑head bias term, yielding P = V·E + (C + 2)·E + B·(11·E² + 9·E).
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →