The Role of Activation Functions Like tanh and ReLU in Neural Networks: A Deep Dive into karpathy/nn-zero-to-hero

Activation functions like tanh and ReLU introduce the non-linearity essential for neural networks to approximate complex functions, with tanh providing smooth, zero-centered outputs bounded between -1 and 1, while ReLU offers computationally efficient, sparse activations that mitigate vanishing gradients in deep architectures.

Activation functions serve as the non-linear "glue" that prevents stacked linear layers from collapsing into a single linear transformation, regardless of network depth. In the karpathy/nn-zero-to-hero educational repository, Andrej Karpathy demonstrates how these functions shape both forward computation and backpropagation through handcrafted scalar implementations. Understanding the distinct characteristics of tanh (hyperbolic tangent) and ReLU (Rectified Linear Unit) is fundamental to designing trainable deep learning architectures.

Why Non-Linearity Matters in Neural Networks

Without activation functions, a deep stack of linear layers would mathematically reduce to a single linear mapping, rendering additional layers useless. Non-linearities allow each layer to reshape the data manifold, enabling the network to learn intricate decision boundaries and approximate arbitrary functions. The derivative of the chosen activation determines how error signals propagate backward during backpropagation, directly impacting training stability and convergence speed.

tanh – Smooth, Bounded, and Zero-Centered

Mathematical Properties and Gradient Behavior

The tanh function follows the form tanh(x) = (e^x - e^-x) / (e^x + e^-x), producing outputs strictly bounded in (-1, 1) with odd symmetry (tanh(-x) = -tanh(x)). Its smooth derivative, 1 - tanh²(x), provides well-behaved gradients for moderate inputs. Because outputs center around zero, tanh helps mitigate weight symmetry issues and can accelerate convergence compared to sigmoid activations. However, for large magnitude inputs, the gradient approaches zero, creating the vanishing gradient problem that hampers training in very deep networks.

Implementation in nn-zero-to-hero

The repository illustrates tanh mechanics in micrograd_lecture_first_half_roughly.ipynb, where a handcrafted Value class defines a tanh() method for scalar automatic differentiation. The companion notebook micrograd_lecture_second_half_roughly.ipynb demonstrates a forward pass via o = n.tanh() and visualizes the resulting computational graph with explicit backward paths. For PyTorch implementations, the repository utilizes torch.tanh(tensor) or torch.nn.Tanh() layers, particularly visible in makemore_part4_backprop.ipynb and makemore_part5_cnn1.ipynb where tanh activations process hidden layer outputs in character-level language models.

ReLU – Sparse, Efficient, and Gradient-Robust

Mathematical Properties and Advantages

ReLU applies the element-wise operation ReLU(x) = max(0, x), outputting zero for negative inputs and preserving positive values unchanged. This piecewise linearity means gradients remain constant (1) for all positive activations, effectively eliminating the vanishing gradient problem for active neurons. The operation requires only a comparison, making it computationally cheaper than exponential activations like tanh. Additionally, ReLU induces sparse activation patterns where many neurons output exactly zero, reducing computational load and often improving generalization through implicit regularization.

Dead Neurons and Practical Considerations

Despite its advantages, ReLU suffers from the "dead neurons" problem: if a unit's weights push its pre-activation negative for all training inputs, the gradient becomes permanently zero, preventing further parameter updates. ReLU outputs are also not zero-centered, which can cause a bias shift in subsequent layers requiring careful learning rate tuning. While the nn-zero-to-hero notebooks focus on tanh for pedagogical clarity regarding smooth derivatives, replacing tanh with torch.relu in the provided architectures immediately demonstrates ReLU's impact on training dynamics and convergence speed.

Comparing tanh vs ReLU for Deep Learning

When selecting between these activations, consider these architectural implications:

  • tanh: Best suited for shallow networks, recurrent neural networks (RNNs), or scenarios requiring bounded, zero-centered outputs. The smooth, symmetric gradients work well when the vanishing gradient risk is manageable through moderate depth.

  • ReLU: Dominates modern convolutional neural networks (CNNs) and deep feed-forward architectures. The constant positive gradient allows training networks with hundreds of layers, though practitioners must monitor for dead neurons through proper initialization and learning rate selection.

In production PyTorch code, this choice manifests as nn.Tanh() versus nn.ReLU() layer definitions, or functional equivalents torch.tanh() versus torch.relu().

Practical Code Examples

The following implementation demonstrates identical two-layer architectures using tanh and ReLU activations, illustrating their behavioral differences during training:

import torch
import torch.nn as nn
import torch.nn.functional as F

# Dummy data

x = torch.randn(64, 10)          # batch of 64, 10-dimensional inputs

y = torch.randint(0, 2, (64,))   # binary targets

# Simple two-layer network with tanh

class NetTanh(nn.Module):
    def __init__(self):
        super().__init__()
        self.fc1 = nn.Linear(10, 32)
        self.fc2 = nn.Linear(32, 1)

    def forward(self, x):
        h = torch.tanh(self.fc1(x))      # ← tanh activation

        out = self.fc2(h)
        return out.squeeze()

# Same architecture with ReLU

class NetReLU(nn.Module):
    def __init__(self):
        super().__init__()
        self.fc1 = nn.Linear(10, 32)
        self.fc2 = nn.Linear(32, 1)

    def forward(self, x):
        h = F.relu(self.fc1(x))          # ← ReLU activation

        out = self.fc2(h)
        return out.squeeze()

# Training loop (shared)

def train(model, epochs=200, lr=0.01):
    opt = torch.optim.SGD(model.parameters(), lr=lr)
    loss_fn = nn.BCEWithLogitsLoss()
    for epoch in range(epochs):
        opt.zero_grad()
        logits = model(x)
        loss = loss_fn(logits, y.float())
        loss.backward()
        opt.step()
    return loss.item()

print("Final loss with tanh :", train(NetTanh()))
print("Final loss with ReLU :", train(NetReLU()))

Running this comparison typically shows the ReLU version converging faster and achieving lower final loss within the same epoch budget, demonstrating its gradient-friendly properties in practice.

Summary

  • Activation functions like tanh and ReLU introduce essential non-linearity that prevents deep networks from collapsing into linear transformations.
  • tanh provides smooth, zero-centered outputs bounded between -1 and 1, making it ideal for shallow networks and RNNs where output symmetry matters, though it risks vanishing gradients in deep architectures.
  • ReLU offers computationally efficient, sparse activations with robust gradient flow for positive inputs, enabling the training of very deep CNNs and modern architectures despite potential dead neuron issues.
  • The karpathy/nn-zero-to-hero repository demonstrates these concepts through handcrafted Value.tanh() implementations in micrograd_lecture_first_half_roughly.ipynb and practical PyTorch usage in the makemore lecture series.
  • Swapping torch.tanh for torch.relu in the repository's examples provides immediate empirical insight into how activation choice affects convergence speed and gradient propagation.

Frequently Asked Questions

What causes the vanishing gradient problem in tanh activation functions?

The vanishing gradient problem occurs because tanh saturates at -1 and 1 for large negative and positive inputs respectively. As the absolute value of the input increases, the derivative 1 - tanh²(x) approaches zero, causing error signals to diminish exponentially during backpropagation through deep layers. This phenomenon prevents early layers from receiving meaningful weight updates in deep networks, effectively halting learning.

Why do deep convolutional networks prefer ReLU over tanh?

Modern CNNs prefer ReLU because its gradient remains constant (1) for all positive activations, regardless of input magnitude. This prevents the vanishing gradient problem that plagues tanh in deep architectures, allowing networks with hundreds of layers to train effectively. Additionally, ReLU's computational simplicity—requiring only a threshold comparison rather than expensive exponential calculations—accelerates both forward and backward passes during training.

What are dead neurons in ReLU and how can they be avoided?

Dead neurons occur when a ReLU unit consistently receives negative pre-activations, causing it to output zero and receive zero gradients permanently. To prevent this, practitioners use techniques like Leaky ReLU (which allows small negative gradients), proper weight initialization schemes (He initialization), and careful learning rate selection. Monitoring the ratio of "active" neurons during training helps identify when dead neurons are impacting model capacity.

How does karpathy/nn-zero-to-hero demonstrate activation function mechanics?

The repository demonstrates activation mechanics through scalar-level automatic differentiation in micrograd_lecture_first_half_roughly.ipynb, where a Value class implements tanh() with explicit backward pass calculations. The notebooks visualize how non-linearities transform gradients during backpropagation, and the makemore series (makemore_part4_backprop.ipynb) shows these principles scaled to neural network layers using PyTorch's torch.tanh operations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →