How to Calculate Total Transformer Model Parameters: A Complete Guide for PyTorch

To calculate total transformer model parameters, sum the token embedding, positional embedding, final layer‑norm, and language‑model head weights, then add N_BLOCKS multiplied by the per‑block count of attention projections (3×n_embed²), MLP layers (8×n_embed² + 5×n_embed), and layer‑norm parameters (4×n_embed).

Managing GPU memory and estimating training costs requires knowing exactly how many trainable weights your model contains. In the FareedKhan-dev/train-llm-from-scratch repository, the vanilla Transformer is implemented from scratch in PyTorch, making it straightforward to derive a closed‑form parameter count. This guide breaks down every weight matrix and bias vector in src/models/transformer.py and its submodules so you can compute the exact parameter budget for any configuration.

Anatomy of the Transformer Parameter Count

The implementation in src/models/transformer.py assembles the model from three macro‑components: embedding tables, a stack of identical transformer blocks, and final projection layers.

Embedding Layers

The input representation combines token indices with positional information:

  • Token embedding — self.token_embed = nn.Embedding(vocab_size, n_embed) at lines 34‑35 of src/models/transformer.py. This contributes vocab_size × n_embed parameters.
  • Positional embedding — self.position_embed = nn.Embedding(context_length, n_embed) at lines 34‑36. This adds context_length × n_embed parameters.

Both embeddings are simple lookup tables without bias vectors.

Inside a Transformer Block

Each block (defined in src/models/transformer_block.py, lines 18‑32) contains three sub‑components whose parameters must be counted:

  1. Two LayerNorms — Each nn.LayerNorm(n_embed) stores a weight and bias vector of size n_embed. Across the block’s pre‑attention and pre‑MLP norms, this yields 4 × n_embed parameters.

  2. Multi‑Head Attention — As implemented in src/models/attention.py (lines 29‑31), every head uses three bias‑free linear projections (query, key, value). With head_size = n_embed // n_head, the math simplifies to 3 × n_embed² parameters per block because the n_head terms cancel out.

  3. Feed‑Forward MLP — The MLP (located in src/models/mlp.py) expands the hidden dimension to 4 × n_embed and projects back. With default bias=True, the two linear layers contribute:

    • Hidden projection: n_embed × (4·n_embed) weights + 4·n_embed bias
    • Output projection: (4·n_embed) × n_embed weights + n_embed bias

    Summed together, the MLP adds 8 × n_embed² + 5 × n_embed parameters.

Total per block:
3·n_embed² (attention) + 8·n_embed² + 5·n_embed (MLP) + 4·n_embed (norms) = 11·n_embed² + 9·n_embed.

Final Projection and Normalization

After the block stack, the model applies:

  • Final LayerNorm — self.layer_norm = nn.LayerNorm(n_embed) at lines 37‑38 contributes 2 × n_embed (weight + bias).
  • Language‑Model Head — self.lm_head = nn.Linear(n_embed, vocab_size) at lines 38‑39 adds vocab_size × n_embed weights plus vocab_size bias terms.

The Exact Parameter Formula

Let V = vocab_size, E = n_embed, C = context_length, and B = N_BLOCKS. The total number of trainable parameters P is:

P = 2·V·E + V + (C + 2)·E + B·(11·E² + 9·E)

Breaking down the terms:

  • 2·V·E counts the token embedding weights and the LM head weight matrix (both size V × E).
  • V accounts for the LM head bias.
  • (C + 2)·E captures positional embeddings (C·E) plus the final layer‑norm weight and bias (2·E).
  • B·(11·E² + 9·E) scales the per‑block parameter count across all N_BLOCKS repetitions.

Step‑by‑Step Implementation in Python

You can implement this formula directly without instantiating the model:

def transformer_param_count(
    n_head: int,
    n_embed: int,
    context_length: int,
    vocab_size: int,
    n_blocks: int,
) -> int:
    """Return the total number of trainable parameters for the model."""
    # Embeddings

    token_embed = vocab_size * n_embed
    pos_embed   = context_length * n_embed

    # Final layers

    final_ln    = 2 * n_embed                # weight + bias

    lm_head_w   = vocab_size * n_embed
    lm_head_b   = vocab_size

    # One Transformer block

    block_params = (
        3 * n_embed ** 2                     # attention (bias‑free)

        + 8 * n_embed ** 2 + 5 * n_embed     # MLP (hidden + projection)

        + 4 * n_embed                         # two LayerNorms (weight+bias each)

    )
    total = (
        token_embed + pos_embed
        + final_ln + lm_head_w + lm_head_b
        + n_blocks * block_params
    )
    return total


# Example – the hyper‑parameters used in the repo’s demo:

if __name__ == "__main__":
    n_head = 4
    n_embed = 32
    context_len = 5
    vocab = 100
    n_blocks = 2

    print("Total parameters:", transformer_param_count(
        n_head, n_embed, context_len, vocab, n_blocks))

Running the snippet with the example values prints:

Total parameters: 115388

Summary

  • Embedding tables contribute (vocab_size + context_length) × n_embed parameters.
  • Each transformer block adds 11·n_embed² + 9·n_embed trainable weights, derived from attention projections, MLP expansion, and layer‑norms.
  • Final layers contribute (2 + vocab_size) × n_embed + vocab_size parameters for the last layer‑norm and the output projection.
  • Use the closed‑form formula P = 2·V·E + V + (C + 2)·E + B·(11·E² + 9·E) to instantly estimate memory usage without loading the model into GPU RAM.

Frequently Asked Questions

Why do the attention projections use bias=False and how does that affect the count?

In src/models/attention.py (lines 29‑31), the query, key, and value projections are initialized with bias=False, which is standard for scaled dot‑product attention. This removes 3 × n_embed bias parameters per block that would otherwise appear in a typical linear layer, reducing the total count by exactly 3 × n_embed × N_BLOCKS.

How is the MLP parameter count derived as 8·n_embed² + 5·n_embed?

The MLP in src/models/mlp.py uses two linear layers with bias=True. The first expands from n_embed to 4·n_embed, contributing n_embed × 4·n_embed weights and 4·n_embed bias. The second projects back, adding 4·n_embed × n_embed weights and n_embed bias. Summing these terms gives 8·n_embed² weights + 5·n_embed bias.

Can I verify this calculation against PyTorch’s built‑in counter?

Yes. After instantiating the Transformer class from src/models/transformer.py, run:

total = sum(p.numel() for p in model.parameters() if p.requires_grad)

This should match the output of transformer_param_count() exactly when provided with the same hyper‑parameters (n_embed, vocab_size, etc.).

Does the formula change if I use weight tying between input and output embeddings?

The repository does not implement weight tying; self.lm_head is independent of self.token_embed. If you tie the weights (setting lm_head.weight = token_embed.weight), subtract vocab_size × n_embed from the total and remove the LM‑head bias term, yielding P = V·E + (C + 2)·E + B·(11·E² + 9·E).

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →