How GPT Architecture Performs Autoregressive Language Modeling: Inside nn-zero-to-hero

The GPT architecture performs autoregressive language modeling by applying causal (masked) self-attention to predict the next token based only on previous tokens, training with next-token cross-entropy loss, and generating sequences through iterative sampling.

The karpathy/nn-zero-to-hero repository provides a ground-up educational implementation of how the GPT architecture performs autoregressive language modeling. This codebase demonstrates the complete transition from simple character-level bigram models to full transformer-based causal language models. Understanding these mechanics reveals precisely why GPT models generate coherent text one token at a time while strictly respecting left-to-right sequence order.

Core Components of GPT Autoregressive Modeling

The GPT architecture treats language modeling as a causal prediction task across four core architectural layers. Each component enforces the autoregressive constraint that future tokens remain invisible during prediction.

Token Embedding and Positional Encoding

Each input token first maps to a dense vector via an embedding matrix. Since the Transformer contains no recurrence, deterministic positional encodings (learned or sinusoidal) are added to distinguish sequence positions. In nn-zero-to-hero, this concept builds upon earlier lessons in lectures/makemore/makemore_part2_mlp.ipynb, which introduces embedding layers before advancing to transformer blocks.

Causal (Masked) Self-Attention

The multi-head self-attention layer attends only to earlier positions through an upper-triangular attention mask. This causal mask forces each token to base its representation exclusively on tokens to its left, implementing the fundamental autoregressive constraint. The mask prevents the model from "cheating" by looking at future target tokens during training.

Transformer Block Stack

The model stacks several identical blocks, each containing:

  • A masked self-attention sub-layer respecting the causal constraint
  • A position-wise feed-forward network (typically a two-layer MLP)
  • Residual connections and layer normalization

Because every layer maintains the causal mask, the entire stack preserves autoregressive properties. The repository explores these normalization techniques in lectures/makemore/makemore_part3_bn.ipynb, which implements batch normalization utilities that later appear in GPT block architectures.

Output Projection and Next-Token Prediction

The final hidden state of the last context token projects back to vocabulary size, producing logits. Applying softmax yields a probability distribution over candidate next tokens. This projection layer converts transformer representations into vocabulary predictions, completing the single-step prediction process.

Training Objective for Autoregressive Language Modeling

GPT trains with a next-token cross-entropy loss (negative log-likelihood) across all positions in the training data. The model maximizes the likelihood of the true next token at every sequence position, learning to predict P(token[t+1] | token[1..t]).

The training loop implements this objective by comparing model logits against ground-truth targets:

model = MiniGPT(vocab_sz=5000).to(device)
opt  = torch.optim.AdamW(model.parameters(), lr=3e-4)

for epoch in range(num_epochs):
    for xb, yb in dataloader:               # xb: (B, T), yb: (B, T)

        logits = model(xb)                  # (B, T, vocab_sz)

        loss = nn.functional.cross_entropy(
            logits.view(-1, logits.size(-1)), yb.view(-1))
        opt.zero_grad()
        loss.backward()
        opt.step()

This objective ensures the model learns probability distributions over the vocabulary conditioned strictly on preceding context.

Autoregressive Generation and Sampling

At inference time, GPT generates text through iterative sampling. The model receives an initial context (often starting with a special token), samples from the output distribution, appends the result to the context, and repeats until reaching an end-of-sequence token.

The repository demonstrates this exact sampling logic in lectures/makemore/makemore_part5_cnn1.ipynb (lines 14-33), showing the fundamental loop that builds context token-by-token:

def sample(model, start_idx, max_new_tokens=100):
    model.eval()
    idx = start_idx.clone()          # (1, T) seed sequence

    for _ in range(max_new_tokens):
        logits = model(idx)[:, -1, :]            # logits for last token only

        probs = torch.softmax(logits, dim=-1)
        nxt = torch.multinomial(probs, num_samples=1)
        idx = torch.cat([idx, nxt], dim=1)       # append sampled token

        if nxt.item() == eos_token_id:
            break
    return idx

This loop implements the core autoregressive property: each new token depends only on previously generated tokens and the original context.

Minimal PyTorch Implementation

A complete GPT-style model implementing these principles requires only causal masking and proper embedding layers. The architecture avoids complex recurrence by relying entirely on the causal self-attention mechanism:

import torch
import torch.nn as nn

class MiniGPT(nn.Module):
    def __init__(self, vocab_sz, n_embd=128, n_head=4, n_layer=4, block_sz=128):
        super().__init__()
        self.tok_emb = nn.Embedding(vocab_sz, n_embd)
        self.pos_emb = nn.Parameter(torch.randn(1, block_sz, n_embd))
        decoder_layer = nn.TransformerDecoderLayer(
            d_model=n_embd,
            nhead=n_head,
            dim_feedforward=4 * n_embd,
            activation='gelu'
        )
        self.transformer = nn.TransformerDecoder(decoder_layer, num_layers=n_layer)
        self.ln_f = nn.LayerNorm(n_embd)
        self.head = nn.Linear(n_embd, vocab_sz, bias=False)

    def forward(self, idx):
        # idx: (B, T) integer token IDs

        B, T = idx.size()
        tok = self.tok_emb(idx)                     # (B, T, n_embd)

        pos = self.pos_emb[:, :T, :]                # (1, T, n_embd)

        x = tok + pos                               # add positional info

        # causal mask: allow each position to see only earlier ones

        mask = torch.triu(torch.ones(T, T, device=idx.device), diagonal=1).bool()
        x = self.transformer(x.permute(1, 0, 2), tgt_mask=mask)  # (T, B, n_embd)

        x = self.ln_f(x.permute(1, 0, 2))           # (B, T, n_embd)

        logits = self.head(x)                       # (B, T, vocab_sz)

        return logits

Key implementation details mirroring the nn-zero-to-hero approach:

  • Learned positional embeddings (pos_emb) inject sequence position information
  • Upper-triangular causal mask (torch.triu) enforces the autoregressive constraint
  • Final linear projection (head) maps hidden states to vocabulary logits

Summary

  • Causal masking enforces the autoregressive property by preventing attention to future tokens via upper-triangular masks
  • Next-token prediction trains the model to maximize likelihood of subsequent tokens using cross-entropy loss across all positions
  • Iterative sampling generates text by repeatedly predicting one token, appending it to context, and feeding the extended sequence back into the model
  • Transformer blocks stack masked self-attention and feed-forward layers with residual connections, maintaining causal constraints throughout the depth
  • Embedding layers combine token representations with positional encodings to handle sequential data without recurrence

Frequently Asked Questions

What makes GPT "autoregressive" rather than autoencoding?

GPT is autoregressive because it predicts future tokens based only on past context, enforced by causal masking that blocks attention to right-side tokens. Autoencoding models like BERT use bidirectional attention to reconstruct masked tokens from surrounding context in both directions, violating the left-to-right constraint required for sequential generation.

Why does GPT use an upper-triangular mask in the attention mechanism?

The upper-triangular mask (implemented via torch.triu with diagonal=1) sets attention scores to negative infinity for future positions, forcing the softmax output to zero for those tokens. This ensures position i can only attend to positions 0 through i-1, maintaining the autoregressive property that predictions depend solely on preceding tokens.

How does the training objective differ between autoregressive and masked language modeling?

Autoregressive language modeling uses a next-token prediction objective where the model learns P(x_t | x_{<t}) and is trained to minimize cross-entropy between predicted logits and the actual next token across all sequence positions simultaneously. Masked language modeling (as in BERT) predicts randomly masked tokens within the sequence using bidirectional context, requiring a different training setup where some input tokens are hidden from the model.

Where in nn-zero-to-hero does the sampling loop first appear?

The fundamental sampling loop for autoregressive generation appears in lectures/makemore/makemore_part5_cnn1.ipynb (lines 14-33), demonstrating how to iteratively build sequences by sampling from model outputs and appending results to the growing context. This precursor to GPT establishes the same generation pattern later used in transformer-based implementations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →