# The Role of Activation Functions Like tanh and ReLU in Neural Networks: A Deep Dive into karpathy/nn-zero-to-hero

> Explore the crucial role of activation functions like tanh and ReLU in neural networks. Understand how they introduce non-linearity, manage gradients, and enable complex function approximation in deep learning architectures.

- Repository: [Andrej/nn-zero-to-hero](https://github.com/karpathy/nn-zero-to-hero)
- Tags: deep-dive
- Published: 2026-05-23

---

**Activation functions like tanh and ReLU introduce the non-linearity essential for neural networks to approximate complex functions, with tanh providing smooth, zero-centered outputs bounded between -1 and 1, while ReLU offers computationally efficient, sparse activations that mitigate vanishing gradients in deep architectures.**

Activation functions serve as the non-linear "glue" that prevents stacked linear layers from collapsing into a single linear transformation, regardless of network depth. In the karpathy/nn-zero-to-hero educational repository, Andrej Karpathy demonstrates how these functions shape both forward computation and backpropagation through handcrafted scalar implementations. Understanding the distinct characteristics of **tanh** (hyperbolic tangent) and **ReLU** (Rectified Linear Unit) is fundamental to designing trainable deep learning architectures.

## Why Non-Linearity Matters in Neural Networks

Without activation functions, a deep stack of linear layers would mathematically reduce to a single linear mapping, rendering additional layers useless. Non-linearities allow each layer to reshape the data manifold, enabling the network to learn intricate decision boundaries and approximate arbitrary functions. The derivative of the chosen activation determines how error signals propagate backward during backpropagation, directly impacting training stability and convergence speed.

## tanh – Smooth, Bounded, and Zero-Centered

### Mathematical Properties and Gradient Behavior

The **tanh** function follows the form `tanh(x) = (e^x - e^-x) / (e^x + e^-x)`, producing outputs strictly bounded in **(-1, 1)** with **odd symmetry** (`tanh(-x) = -tanh(x)`). Its smooth derivative, `1 - tanh²(x)`, provides well-behaved gradients for moderate inputs. Because outputs center around zero, tanh helps mitigate weight symmetry issues and can accelerate convergence compared to sigmoid activations. However, for large magnitude inputs, the gradient approaches zero, creating the **vanishing gradient problem** that hampers training in very deep networks.

### Implementation in nn-zero-to-hero

The repository illustrates tanh mechanics in `micrograd_lecture_first_half_roughly.ipynb`, where a handcrafted `Value` class defines a `tanh()` method for scalar automatic differentiation. The companion notebook `micrograd_lecture_second_half_roughly.ipynb` demonstrates a forward pass via `o = n.tanh()` and visualizes the resulting computational graph with explicit backward paths. For PyTorch implementations, the repository utilizes `torch.tanh(tensor)` or `torch.nn.Tanh()` layers, particularly visible in `makemore_part4_backprop.ipynb` and `makemore_part5_cnn1.ipynb` where tanh activations process hidden layer outputs in character-level language models.

## ReLU – Sparse, Efficient, and Gradient-Robust

### Mathematical Properties and Advantages

**ReLU** applies the element-wise operation `ReLU(x) = max(0, x)`, outputting zero for negative inputs and preserving positive values unchanged. This piecewise linearity means gradients remain constant (1) for all positive activations, effectively eliminating the vanishing gradient problem for active neurons. The operation requires only a comparison, making it computationally cheaper than exponential activations like tanh. Additionally, ReLU induces **sparse activation patterns** where many neurons output exactly zero, reducing computational load and often improving generalization through implicit regularization.

### Dead Neurons and Practical Considerations

Despite its advantages, ReLU suffers from the **"dead neurons"** problem: if a unit's weights push its pre-activation negative for all training inputs, the gradient becomes permanently zero, preventing further parameter updates. ReLU outputs are also not zero-centered, which can cause a bias shift in subsequent layers requiring careful learning rate tuning. While the nn-zero-to-hero notebooks focus on tanh for pedagogical clarity regarding smooth derivatives, replacing `tanh` with `torch.relu` in the provided architectures immediately demonstrates ReLU's impact on training dynamics and convergence speed.

## Comparing tanh vs ReLU for Deep Learning

When selecting between these activations, consider these architectural implications:

- **tanh**: Best suited for shallow networks, recurrent neural networks (RNNs), or scenarios requiring bounded, zero-centered outputs. The smooth, symmetric gradients work well when the vanishing gradient risk is manageable through moderate depth.

- **ReLU**: Dominates modern convolutional neural networks (CNNs) and deep feed-forward architectures. The constant positive gradient allows training networks with hundreds of layers, though practitioners must monitor for dead neurons through proper initialization and learning rate selection.

In production PyTorch code, this choice manifests as `nn.Tanh()` versus `nn.ReLU()` layer definitions, or functional equivalents `torch.tanh()` versus `torch.relu()`.

## Practical Code Examples

The following implementation demonstrates identical two-layer architectures using tanh and ReLU activations, illustrating their behavioral differences during training:

```python
import torch
import torch.nn as nn
import torch.nn.functional as F

# Dummy data

x = torch.randn(64, 10)          # batch of 64, 10-dimensional inputs

y = torch.randint(0, 2, (64,))   # binary targets

# Simple two-layer network with tanh

class NetTanh(nn.Module):
    def __init__(self):
        super().__init__()
        self.fc1 = nn.Linear(10, 32)
        self.fc2 = nn.Linear(32, 1)

    def forward(self, x):
        h = torch.tanh(self.fc1(x))      # ← tanh activation

        out = self.fc2(h)
        return out.squeeze()

# Same architecture with ReLU

class NetReLU(nn.Module):
    def __init__(self):
        super().__init__()
        self.fc1 = nn.Linear(10, 32)
        self.fc2 = nn.Linear(32, 1)

    def forward(self, x):
        h = F.relu(self.fc1(x))          # ← ReLU activation

        out = self.fc2(h)
        return out.squeeze()

# Training loop (shared)

def train(model, epochs=200, lr=0.01):
    opt = torch.optim.SGD(model.parameters(), lr=lr)
    loss_fn = nn.BCEWithLogitsLoss()
    for epoch in range(epochs):
        opt.zero_grad()
        logits = model(x)
        loss = loss_fn(logits, y.float())
        loss.backward()
        opt.step()
    return loss.item()

print("Final loss with tanh :", train(NetTanh()))
print("Final loss with ReLU :", train(NetReLU()))

```

Running this comparison typically shows the **ReLU** version converging faster and achieving lower final loss within the same epoch budget, demonstrating its gradient-friendly properties in practice.

## Summary

- Activation functions like **tanh** and **ReLU** introduce essential non-linearity that prevents deep networks from collapsing into linear transformations.
- **tanh** provides smooth, zero-centered outputs bounded between -1 and 1, making it ideal for shallow networks and RNNs where output symmetry matters, though it risks vanishing gradients in deep architectures.
- **ReLU** offers computationally efficient, sparse activations with robust gradient flow for positive inputs, enabling the training of very deep CNNs and modern architectures despite potential dead neuron issues.
- The karpathy/nn-zero-to-hero repository demonstrates these concepts through handcrafted `Value.tanh()` implementations in `micrograd_lecture_first_half_roughly.ipynb` and practical PyTorch usage in the makemore lecture series.
- Swapping `torch.tanh` for `torch.relu` in the repository's examples provides immediate empirical insight into how activation choice affects convergence speed and gradient propagation.

## Frequently Asked Questions

### What causes the vanishing gradient problem in tanh activation functions?

The vanishing gradient problem occurs because tanh saturates at -1 and 1 for large negative and positive inputs respectively. As the absolute value of the input increases, the derivative `1 - tanh²(x)` approaches zero, causing error signals to diminish exponentially during backpropagation through deep layers. This phenomenon prevents early layers from receiving meaningful weight updates in deep networks, effectively halting learning.

### Why do deep convolutional networks prefer ReLU over tanh?

Modern CNNs prefer ReLU because its gradient remains constant (1) for all positive activations, regardless of input magnitude. This prevents the vanishing gradient problem that plagues tanh in deep architectures, allowing networks with hundreds of layers to train effectively. Additionally, ReLU's computational simplicity—requiring only a threshold comparison rather than expensive exponential calculations—accelerates both forward and backward passes during training.

### What are dead neurons in ReLU and how can they be avoided?

Dead neurons occur when a ReLU unit consistently receives negative pre-activations, causing it to output zero and receive zero gradients permanently. To prevent this, practitioners use techniques like Leaky ReLU (which allows small negative gradients), proper weight initialization schemes (He initialization), and careful learning rate selection. Monitoring the ratio of "active" neurons during training helps identify when dead neurons are impacting model capacity.

### How does karpathy/nn-zero-to-hero demonstrate activation function mechanics?

The repository demonstrates activation mechanics through scalar-level automatic differentiation in `micrograd_lecture_first_half_roughly.ipynb`, where a `Value` class implements `tanh()` with explicit backward pass calculations. The notebooks visualize how non-linearities transform gradients during backpropagation, and the makemore series (`makemore_part4_backprop.ipynb`) shows these principles scaled to neural network layers using PyTorch's `torch.tanh` operations.