# How to Implement Manual Backpropagation Without PyTorch Autograd

> Learn to implement manual backpropagation without PyTorch autograd. Compute gradients using the chain rule and apply SGD updates for deeper neural network understanding.

- Repository: [Andrej/nn-zero-to-hero](https://github.com/karpathy/nn-zero-to-hero)
- Tags: how-to-guide
- Published: 2026-05-23

---

**You can train neural networks without `torch.autograd` by explicitly computing gradients using the chain rule, storing them in tensor `.grad` attributes, and applying SGD updates manually.**

The *nn-zero-to-hero* repository demonstrates exactly how to implement manual backpropagation without PyTorch autograd in the Makemore lecture series. By disabling automatic differentiation and deriving gradients step-by-step through the MLP architecture, you gain foundational understanding of backpropagation mechanics while maintaining numerical correctness validated against PyTorch's built-in gradients.

## The Core Strategy for Manual Backpropagation

To implement manual backpropagation in PyTorch while keeping the framework's tensor operations, you must prevent automatic gradient tracking while retaining the ability to store computed gradients.

**Keep `requires_grad=True`** on parameters solely to allocate the `.grad` field for storage. Never call `.backward()`. Instead, wrap the forward pass in `torch.no_grad()` to prevent computational graph construction, then manually derive each derivative using the chain rule.

According to the source code in `lectures/makemore/makemore_part4_backprop.ipynb`, this approach involves computing six distinct gradient components: `dC` (embeddings), `dW1`, `db1` (first linear layer), `dhpreact` (batch norm), `dW2`, `db2` (second linear layer), and `dlogits` (output).

## Step-by-Step Gradient Computation

### Output Layer and Cross-Entropy Loss

The gradient of the cross-entropy loss with respect to the logits (pre-softmax outputs) follows a simple pattern. As implemented in notebook cells L98-L104:

```python
dlogits = F.softmax(logits, dim=1)
dlogits[torch.arange(batch_size), Yb] -= 1
dlogits /= batch_size

```

This formula represents `softmax(logits) - one_hot(Y)` normalized by batch size, where `Yb` contains the target indices.

### Second Linear Layer Gradients

Backpropagating through the final linear layer `logits = h @ W2 + b2` requires computing gradients for the weights, bias, and hidden activations. From cells L108-L115:

```python
dh = dlogits @ W2.t()
dW2 = h.t() @ dlogits
db2 = dlogits.sum(0)

```

**`dh`** flows backward to the preceding activation function, while **`dW2`** and **`db2`** accumulate parameter updates.

### Tanh Activation Backward

For the `tanh` activation `h = torch.tanh(hpreact)`, the derivative follows the elementary rule `d/dx tanh(x) = 1 - tanh²(x)`:

```python
dhpreact = (1 - h**2) * dh

```

This element-wise multiplication propagates the gradient through the nonlinearity to the batch normalization output.

### Batch Normalization Derivatives

Batch normalization requires careful handling due to its statistics-dependent computation. The manual implementation in cells L447-L470 provides a compact formula for `dhprebn` (gradient before batch norm):

```python
dbngain = (bnraw * dhpreact).sum(0, keepdim=True)
dbnbias = dhpreact.sum(0, keepdim=True)
dhprebn = bngain * bnvar_inv / batch_size * (
    batch_size * dhpreact
    - dhpreact.sum(0)
    - batch_size/(batch_size-1) * bnraw * (dhpreact*bnraw).sum(0)
)

```

Here, `bnraw` represents the normalized pre-activations `(hprebn - bnmean) * bnvar_inv`. The `dbngain` and `dbnbias` terms compute gradients for the learnable scale and shift parameters.

### First Linear Layer and Embeddings

Gradients for the first linear layer `hprebn = embcat @ W1 + b1` follow the same pattern as the second layer:

```python
dW1 = embcat.t() @ dhprebn
db1 = dhprebn.sum(0)

```

For the embedding lookup `emb = C[Xb]`, gradients must be accumulated into the embedding matrix `C` using a double loop, as shown in cells L332-L337:

```python
demb = dhprebn.view(emb.shape)  # Reshape to (batch, block, hidden)

dC = torch.zeros_like(C)
for b in range(batch_size):
    for t in range(block_size):
        ix = Xb[b, t]
        dC[ix] += demb[b, t]

```

This loops over each token in the context window and accumulates the gradient into the appropriate embedding vector.

## Validating Manual Gradients Against Autograd

The notebook includes a validation mechanism in cells L380-L440 that confirms bit-exact correctness. The `cmp()` function compares each manual gradient tensor against `torch.autograd` output:

```python
def cmp(s, dt, t):
    ex = torch.all(dt == t.grad).item()
    app = torch.allclose(dt, t.grad)
    maxdiff = (dt - t.grad).abs().max().item()
    print(f'{s:15s} | exact: {str(ex):5s} | approximate: {str(app):5s} | maxdiff: {maxdiff}')

```

When properly implemented, **maxdiff ≈ 0**, proving your manual backpropagation matches PyTorch's automatic differentiation numerically.

## Complete Manual Backpropagation Training Loop

Below is the complete implementation combining all gradient computations into a training loop. This follows the exact sequence from `lectures/makemore/makemore_part4_backprop.ipynb`:

```python
import torch
import torch.nn.functional as F

# Parameters (initialized as in notebook cells L64-L84)

g = torch.Generator().manual_seed(2147483647)
C = torch.randn((vocab_size, n_embd), generator=g)
W1 = torch.randn((n_embd*block_size, n_hidden), generator=g) * (5/3)/((n_embd*block_size)**0.5)
b1 = torch.randn(n_hidden, generator=g)*0.1
W2 = torch.randn((n_hidden, vocab_size), generator=g)*0.1
b2 = torch.randn(vocab_size, generator=g)*0.1
bngain = torch.randn((1, n_hidden))*0.1 + 1.0
bnbias = torch.randn((1, n_hidden))*0.1

params = [C, W1, b1, W2, b2, bngain, bnbias]
for p in params:
    p.requires_grad = True  # Only to hold .grad field

# Training loop

batch_size = 32
for step in range(200000):
    # Minibatch

    idx = torch.randint(0, Xtr.shape[0], (batch_size,), generator=g)
    Xb, Yb = Xtr[idx], Ytr[idx]
    
    with torch.no_grad():
        # Forward pass (cells L165-L171 for batch norm)

        emb = C[Xb]
        embcat = emb.view(batch_size, -1)
        hprebn = embcat @ W1 + b1
        
        # Batch norm forward

        bnmean = hprebn.mean(0, keepdim=True)
        bnvar = hprebn.var(0, keepdim=True, unbiased=True)
        bnvar_inv = (bnvar + 1e-5).rsqrt()
        bnraw = (hprebn - bnmean) * bnvar_inv
        hpreact = bngain * bnraw + bnbias
        
        h = torch.tanh(hpreact)
        logits = h @ W2 + b2
        loss = F.cross_entropy(logits, Yb)
        
        # Manual backward pass

        # 1) Cross-entropy (cells L98-L104)

        dlogits = F.softmax(logits, dim=1)
        dlogits[torch.arange(batch_size), Yb] -= 1
        dlogits /= batch_size
        
        # 2) Second linear (cells L108-L115)

        dh = dlogits @ W2.t()
        dW2 = h.t() @ dlogits
        db2 = dlogits.sum(0)
        
        # 3) Tanh

        dhpreact = (1 - h**2) * dh
        
        # 4) Batch norm (compact formula, cells L447-L470)

        dbngain = (bnraw * dhpreact).sum(0, keepdim=True)
        dbnbias = dhpreact.sum(0, keepdim=True)
        dhprebn = bngain * bnvar_inv / batch_size * (
            batch_size * dhpreact
            - dhpreact.sum(0)
            - batch_size/(batch_size-1) * bnraw * (dhpreact*bnraw).sum(0)
        )
        
        # 5) First linear

        dW1 = embcat.t() @ dhprebn
        db1 = dhprebn.sum(0)
        
        # 6) Embedding (double loop, cells L332-L337)

        demb = dhprebn.view(emb.shape)
        dC = torch.zeros_like(C)
        for b in range(batch_size):
            for t in range(block_size):
                dC[Xb[b, t]] += demb[b, t]
        
        grads = [dC, dW1, db1, dW2, db2, dbngain, dbnbias]
        
        # Manual SGD update

        lr = 0.1 if step < 100000 else 0.01
        for p, g in zip(params, grads):
            p.data += -lr * g

```

## Summary

- **Manual backpropagation** requires computing gradients explicitly using the chain rule for each operation in your computational graph.
- The *nn-zero-to-hero* repository provides validated implementations in `lectures/makemore/makemore_part4_backprop.ipynb` with cell-level references for cross-entropy, linear layers, batch norm, and embeddings.
- Keep `requires_grad=True` only to store gradients in `.grad` fields, but use `torch.no_grad()` to prevent automatic graph construction during the forward pass.
- Key formulas include `dlogits = softmax(logits) - one_hot(Y)` for the output and compact batch norm derivatives for `dhprebn`.
- Validate implementation using the `cmp()` function (cells L380-L440) to ensure bit-exact match with PyTorch autograd before deploying.

## Frequently Asked Questions

### Do I need to set `requires_grad=False` to disable autograd?

No. You should keep `requires_grad=True` on parameters so they can store gradient values in their `.grad` attributes. Instead, disable autograd by wrapping your forward pass in `with torch.no_grad():`. This prevents PyTorch from building the computational graph while still allowing you to manually assign values to `p.grad` later.

### How do I verify my manual gradients are correct?

Use the `cmp()` function provided in notebook cells L380-L440. This utility compares your manually computed tensors against PyTorch's autograd output by checking for exact equality or close numerical agreement. When implemented correctly, the maximum difference between manual and automatic gradients should be approximately zero.

### Can I apply manual backpropagation to architectures beyond the Makemore MLP?

Yes. The principles apply to any differentiable computation. For understanding how to extend this to other operations, reference `lectures/micrograd/micrograd_lecture_first_half_roughly.ipynb`, which builds a tiny autodiff engine from scratch. The chain rule remains the same regardless of architecture complexity.

### Is manual backpropagation slower than using PyTorch autograd?

Generally yes. Manual implementation in Python lacks the optimized C++ kernels and memory efficiency of PyTorch's built-in autograd engine. However, the educational value and debugging capability of seeing exact gradient flow through each layer often outweighs the performance cost during development and learning.