# How micrograd's Automatic Differentiation Engine Works: A Deep Dive into Reverse-Mode Autograd

> Explore micrograd's reverse-mode autograd engine. Learn how its Value class builds a dynamic graph and backpropagates gradients for efficient deep learning.

- Repository: [Andrej/nn-zero-to-hero](https://github.com/karpathy/nn-zero-to-hero)
- Tags: deep-dive
- Published: 2026-05-23

---

**`micrograd` implements reverse-mode automatic differentiation using a single `Value` class that constructs a dynamic computational graph through operator overloading, then propagates gradients backward via topologically sorted nodes and locally defined `_backward` closures.**

The micrograd automatic differentiation engine, featured in Andrej Karpathy's `nn-zero-to-hero` educational series, demonstrates how to build a complete gradient computation system using nothing but standard Python. By examining the `Value` class defined in `lectures/micrograd/micrograd_lecture_first_half_roughly.ipynb`, you can see precisely how scalar values, mathematical operations, and gradient propagation merge into a functional deep learning primitive.

## The Value Class: Scalar Nodes with Gradient Tracking

At the heart of the micrograd automatic differentiation engine lies the `Value` class, defined in lines 72‑84 of the first lecture notebook. Each instance represents a scalar node in a computational graph and maintains five critical attributes:

- **`data`** – The actual scalar float value.
- **`grad`** – The accumulated gradient of the loss with respect to this node, initialized to `0`.
- **`_prev`** – A set of child `Value` objects that were inputs to this node's operation.
- **`_op`** – A string identifier for the operation that produced this node (e.g., `'+'`, `'*'`, `'tanh'`).
- **`_backward`** – A closure that defines how to propagate the upstream gradient to the node's children.

This design treats every mathematical operation as a node with dependencies, creating a directed acyclic graph (DAG) where gradients can flow from outputs back to inputs.

## Forward Pass: Building the Computational Graph

### Arithmetic Operations and Graph Construction

When you add or multiply `Value` objects, the overloaded operators (`__add__`, `__mul__`) do more than compute results. As shown in lines 86‑92 of `micrograd_lecture_first_half_roughly.ipynb`, the addition operation creates a new `Value` whose `_prev` contains the operands and whose `_op` is set to `'+'`:

```python
def __add__(self, other):
    other = other if isinstance(other, Value) else Value(other)
    out = Value(self.data + other.data, (self, other), '+')
    
    def _backward():
        self.grad += 1.0 * out.grad
        other.grad += 1.0 * out.grad
    out._backward = _backward
    
    return out

```

The `_backward` closure implements the local derivative of addition. Since `∂(a+b)/∂a = 1`, it distributes the upstream gradient (`out.grad`) equally to both inputs. Multiplication follows the same pattern but uses the multiplicative rule: `∂(a*b)/∂a = b` and `∂(a*b)/∂b = a`.

### Non-Linear Activations (tanh)

The engine supports non-linearities through methods like `tanh`, implemented in lines 105‑112. The method computes the hyperbolic tangent of the input value and stores its derivative `(1 - t**2)` in the `_backward` closure:

```python
def tanh(self):
    x = self.data
    t = (math.exp(2*x) - 1) / (math.exp(2*x) + 1)
    out = Value(t, (self,), 'tanh')
    
    def _backward():
        self.grad += (1 - t**2) * out.grad
    out._backward = _backward
    
    return out

```

This allows the chain rule to propagate through activation functions during the backward pass.

## Reverse-Mode Automatic Differentiation via Topological Sort

The `backward` method, spanning lines 122‑131 in the same notebook, executes reverse-mode automatic differentiation. This method performs two critical steps:

1. **Topological Ordering** – It builds a list of all nodes in the graph sorted such that every node appears before its dependencies. This ensures that when gradients flow backward, parent nodes always receive their gradients before children try to propagate further.

2. **Gradient Propagation** – It initializes the root node's gradient to `1.0` (since `∂L/∂L = 1`), then iterates through the topologically sorted list in reverse, invoking each node's `_backward` closure.

According to the source code in `karpathy/nn-zero-to-hero`, the algorithm looks like this:

```python
def backward(self):
    # Build topological ordering

    topo = []
    visited = set()
    def build(v):
        if v not in visited:
            visited.add(v)
            for child in v._prev:
                build(child)
            topo.append(v)
    build(self)
    
    # Accumulate gradients from output to inputs

    self.grad = 1.0
    for node in reversed(topo):
        node._backward()

```

Because each `_backward` closure knows the local Jacobian of its operation, calling them in reverse topological order automatically applies the chain rule, accumulating the correct gradients in each leaf node.

## Visualizing the Computational Graph

To debug and understand the flow of gradients, the second lecture notebook (`micrograd_lecture_second_half_roughly.ipynb`, lines 149‑169) provides `trace` and `draw_dot` utilities. These functions walk the `_prev` relationships to generate GraphViz diagrams:

```python
from graphviz import Digraph

def trace(root):
    nodes, edges = set(), set()
    def build(v):
        if v not in nodes:
            nodes.add(v)
            for child in v._prev:
                edges.add((child, v))
                build(child)
    build(root)
    return nodes, edges

def draw_dot(root):
    dot = Digraph(format='svg', graph_attr={'rankdir': 'LR'})
    nodes, edges = trace(root)
    for n in nodes:
        uid = str(id(n))
        dot.node(name=uid, label=f"{n.label} | data {n.data:.4f} | grad {n.grad:.4f}", shape='record')
        if n._op:
            dot.node(name=uid + n._op, label=n._op)
            dot.edge(uid + n._op, uid)
    for n1, n2 in edges:
        dot.edge(str(id(n1)), str(id(n2)) + n2._op)
    return dot

```

This visualization confirms that the automatic differentiation engine correctly built the intended computation structure before backpropagation begins.

## Practical Examples of micrograd in Action

### Example 1: Scalar Gradient Computation

```python

# Assuming Value class is defined from the notebook

a = Value(2.0, label='a')
b = Value(-3.0, label='b')
c = a * b               # c.data = -6.0

d = c + Value(10.0)     # d.data = 4.0

L = d * Value(-2.0)     # L.data = -8.0 (the loss)

L.backward()

print(f"∂L/∂a = {a.grad}")  # Returns 4.0 (chain rule: -2.0 * -3.0 * 1.0)

print(f"∂L/∂b = {b.grad}")  # Returns -4.0 (chain rule: -2.0 * 2.0 * 1.0)

```

*Implementation reference:* `lectures/micrograd/micrograd_lecture_first_half_roughly.ipynb`, lines 72‑131.

### Example 2: Single Neuron Forward and Backward Pass

```python

# Inputs

x1 = Value(2.0, label='x1')
w1 = Value(-3.0, label='w1')
b = Value(0.0, label='b')

# Forward pass

n = x1 * w1 + b         # Linear combination

o = n.tanh()            # Activation

# Backward pass

o.backward()

print(f"Weight gradient: {w1.grad}")  # ∂o/∂w1

print(f"Bias gradient: {b.grad}")    # ∂o/∂b

```

*Relevant implementation:* Lines 105‑131 for `tanh` and the `backward` method.

### Example 3: Generating Graph Visualization

```python

# After building computation graph with root node 'o'

dot = draw_dot(o)
dot.render('neuron_graph', format='svg')

```

*Graph utilities source:* `lectures/micrograd/micrograd_lecture_second_half_roughly.ipynb`, lines 149‑169.

## Summary

- **`Value` class** – The entire micrograd automatic differentiation engine centers on a single class that wraps scalars and tracks their relationships.
- **Dynamic graph construction** – Operator overloading (`__add__`, `__mul__`, `tanh`) builds a DAG on-the-fly during the forward pass, storing local derivatives in `_backward` closures.
- **Topological sort** – The `backward` method uses depth-first search to order nodes before reversing the list to ensure correct gradient flow from outputs to inputs.
- **Chain rule automation** – Each node’s `_backward` closure multiplies the incoming gradient by the local Jacobian, automatically implementing the chain rule without explicit matrix operations.
- **Visualization support** – Helper functions in the second lecture notebook allow inspection of the computational graph structure using GraphViz.

## Frequently Asked Questions

### What is reverse-mode automatic differentiation and why does micrograd use it?

Reverse-mode automatic differentiation computes the gradient of a scalar output with respect to all inputs in a single pass. It first evaluates the function forward to build the computational graph, then propagates derivatives backward from the output. Micrograd uses this approach because it is computationally efficient for scalar loss functions typical in neural network training, requiring only one backward sweep to obtain gradients for all parameters.

### How does the `Value` class apply the chain rule during backpropagation?

The `Value` class applies the chain rule through the `_backward` closures stored during the forward pass. When `backward()` iterates through the topologically sorted nodes in reverse, each node's `_backward` closure multiplies the upstream gradient (from `out.grad`) by the local partial derivative (e.g., `1.0` for addition, `other.data` for multiplication, or `1-t**2` for tanh). This product is then added to the child node's `grad` attribute, accumulating the total derivative via the chain rule.

### Can micrograd handle tensors or is it limited to scalars?

Micrograd is explicitly designed for scalar values only. The `Value` class operates on individual floats, and the `_backward` closures implement scalar derivatives. While you could theoretically wrap tensor operations, the engine as implemented in `micrograd_lecture_first_half_roughly.ipynb` does not support vectorized operations or matrix gradients—it is an educational tool demonstrating automatic differentiation fundamentals rather than a production tensor library.

### Where can I find the complete micrograd source code?

The complete implementation resides in the `karpathy/nn-zero-to-hero` repository. The core `Value` class and automatic differentiation logic are in `lectures/micrograd/micrograd_lecture_first_half_roughly.ipynb` (lines 72‑131), while visualization utilities appear in `lectures/micrograd/micrograd_lecture_second_half_roughly.ipynb` (lines 149‑169). The repository also includes a [`README.md`](https://github.com/karpathy/nn-zero-to-hero/blob/main/README.md) with links to video lectures explaining the mathematical foundations of the automatic differentiation engine.