How micrograd's Automatic Differentiation Engine Works: A Deep Dive into Reverse-Mode Autograd

micrograd implements reverse-mode automatic differentiation using a single Value class that constructs a dynamic computational graph through operator overloading, then propagates gradients backward via topologically sorted nodes and locally defined _backward closures.

The micrograd automatic differentiation engine, featured in Andrej Karpathy's nn-zero-to-hero educational series, demonstrates how to build a complete gradient computation system using nothing but standard Python. By examining the Value class defined in lectures/micrograd/micrograd_lecture_first_half_roughly.ipynb, you can see precisely how scalar values, mathematical operations, and gradient propagation merge into a functional deep learning primitive.

The Value Class: Scalar Nodes with Gradient Tracking

At the heart of the micrograd automatic differentiation engine lies the Value class, defined in lines 72‑84 of the first lecture notebook. Each instance represents a scalar node in a computational graph and maintains five critical attributes:

  • data – The actual scalar float value.
  • grad – The accumulated gradient of the loss with respect to this node, initialized to 0.
  • _prev – A set of child Value objects that were inputs to this node's operation.
  • _op – A string identifier for the operation that produced this node (e.g., '+', '*', 'tanh').
  • _backward – A closure that defines how to propagate the upstream gradient to the node's children.

This design treats every mathematical operation as a node with dependencies, creating a directed acyclic graph (DAG) where gradients can flow from outputs back to inputs.

Forward Pass: Building the Computational Graph

Arithmetic Operations and Graph Construction

When you add or multiply Value objects, the overloaded operators (__add__, __mul__) do more than compute results. As shown in lines 86‑92 of micrograd_lecture_first_half_roughly.ipynb, the addition operation creates a new Value whose _prev contains the operands and whose _op is set to '+':

def __add__(self, other):
    other = other if isinstance(other, Value) else Value(other)
    out = Value(self.data + other.data, (self, other), '+')
    
    def _backward():
        self.grad += 1.0 * out.grad
        other.grad += 1.0 * out.grad
    out._backward = _backward
    
    return out

The _backward closure implements the local derivative of addition. Since ∂(a+b)/∂a = 1, it distributes the upstream gradient (out.grad) equally to both inputs. Multiplication follows the same pattern but uses the multiplicative rule: ∂(a*b)/∂a = b and ∂(a*b)/∂b = a.

Non-Linear Activations (tanh)

The engine supports non-linearities through methods like tanh, implemented in lines 105‑112. The method computes the hyperbolic tangent of the input value and stores its derivative (1 - t**2) in the _backward closure:

def tanh(self):
    x = self.data
    t = (math.exp(2*x) - 1) / (math.exp(2*x) + 1)
    out = Value(t, (self,), 'tanh')
    
    def _backward():
        self.grad += (1 - t**2) * out.grad
    out._backward = _backward
    
    return out

This allows the chain rule to propagate through activation functions during the backward pass.

Reverse-Mode Automatic Differentiation via Topological Sort

The backward method, spanning lines 122‑131 in the same notebook, executes reverse-mode automatic differentiation. This method performs two critical steps:

  1. Topological Ordering – It builds a list of all nodes in the graph sorted such that every node appears before its dependencies. This ensures that when gradients flow backward, parent nodes always receive their gradients before children try to propagate further.

  2. Gradient Propagation – It initializes the root node's gradient to 1.0 (since ∂L/∂L = 1), then iterates through the topologically sorted list in reverse, invoking each node's _backward closure.

According to the source code in karpathy/nn-zero-to-hero, the algorithm looks like this:

def backward(self):
    # Build topological ordering

    topo = []
    visited = set()
    def build(v):
        if v not in visited:
            visited.add(v)
            for child in v._prev:
                build(child)
            topo.append(v)
    build(self)
    
    # Accumulate gradients from output to inputs

    self.grad = 1.0
    for node in reversed(topo):
        node._backward()

Because each _backward closure knows the local Jacobian of its operation, calling them in reverse topological order automatically applies the chain rule, accumulating the correct gradients in each leaf node.

Visualizing the Computational Graph

To debug and understand the flow of gradients, the second lecture notebook (micrograd_lecture_second_half_roughly.ipynb, lines 149‑169) provides trace and draw_dot utilities. These functions walk the _prev relationships to generate GraphViz diagrams:

from graphviz import Digraph

def trace(root):
    nodes, edges = set(), set()
    def build(v):
        if v not in nodes:
            nodes.add(v)
            for child in v._prev:
                edges.add((child, v))
                build(child)
    build(root)
    return nodes, edges

def draw_dot(root):
    dot = Digraph(format='svg', graph_attr={'rankdir': 'LR'})
    nodes, edges = trace(root)
    for n in nodes:
        uid = str(id(n))
        dot.node(name=uid, label=f"{n.label} | data {n.data:.4f} | grad {n.grad:.4f}", shape='record')
        if n._op:
            dot.node(name=uid + n._op, label=n._op)
            dot.edge(uid + n._op, uid)
    for n1, n2 in edges:
        dot.edge(str(id(n1)), str(id(n2)) + n2._op)
    return dot

This visualization confirms that the automatic differentiation engine correctly built the intended computation structure before backpropagation begins.

Practical Examples of micrograd in Action

Example 1: Scalar Gradient Computation


# Assuming Value class is defined from the notebook

a = Value(2.0, label='a')
b = Value(-3.0, label='b')
c = a * b               # c.data = -6.0

d = c + Value(10.0)     # d.data = 4.0

L = d * Value(-2.0)     # L.data = -8.0 (the loss)

L.backward()

print(f"∂L/∂a = {a.grad}")  # Returns 4.0 (chain rule: -2.0 * -3.0 * 1.0)

print(f"∂L/∂b = {b.grad}")  # Returns -4.0 (chain rule: -2.0 * 2.0 * 1.0)

Implementation reference: lectures/micrograd/micrograd_lecture_first_half_roughly.ipynb, lines 72‑131.

Example 2: Single Neuron Forward and Backward Pass


# Inputs

x1 = Value(2.0, label='x1')
w1 = Value(-3.0, label='w1')
b = Value(0.0, label='b')

# Forward pass

n = x1 * w1 + b         # Linear combination

o = n.tanh()            # Activation

# Backward pass

o.backward()

print(f"Weight gradient: {w1.grad}")  # ∂o/∂w1

print(f"Bias gradient: {b.grad}")    # ∂o/∂b

Relevant implementation: Lines 105‑131 for tanh and the backward method.

Example 3: Generating Graph Visualization


# After building computation graph with root node 'o'

dot = draw_dot(o)
dot.render('neuron_graph', format='svg')

Graph utilities source: lectures/micrograd/micrograd_lecture_second_half_roughly.ipynb, lines 149‑169.

Summary

  • Value class – The entire micrograd automatic differentiation engine centers on a single class that wraps scalars and tracks their relationships.
  • Dynamic graph construction – Operator overloading (__add__, __mul__, tanh) builds a DAG on-the-fly during the forward pass, storing local derivatives in _backward closures.
  • Topological sort – The backward method uses depth-first search to order nodes before reversing the list to ensure correct gradient flow from outputs to inputs.
  • Chain rule automation – Each node’s _backward closure multiplies the incoming gradient by the local Jacobian, automatically implementing the chain rule without explicit matrix operations.
  • Visualization support – Helper functions in the second lecture notebook allow inspection of the computational graph structure using GraphViz.

Frequently Asked Questions

What is reverse-mode automatic differentiation and why does micrograd use it?

Reverse-mode automatic differentiation computes the gradient of a scalar output with respect to all inputs in a single pass. It first evaluates the function forward to build the computational graph, then propagates derivatives backward from the output. Micrograd uses this approach because it is computationally efficient for scalar loss functions typical in neural network training, requiring only one backward sweep to obtain gradients for all parameters.

How does the Value class apply the chain rule during backpropagation?

The Value class applies the chain rule through the _backward closures stored during the forward pass. When backward() iterates through the topologically sorted nodes in reverse, each node's _backward closure multiplies the upstream gradient (from out.grad) by the local partial derivative (e.g., 1.0 for addition, other.data for multiplication, or 1-t**2 for tanh). This product is then added to the child node's grad attribute, accumulating the total derivative via the chain rule.

Can micrograd handle tensors or is it limited to scalars?

Micrograd is explicitly designed for scalar values only. The Value class operates on individual floats, and the _backward closures implement scalar derivatives. While you could theoretically wrap tensor operations, the engine as implemented in micrograd_lecture_first_half_roughly.ipynb does not support vectorized operations or matrix gradients—it is an educational tool demonstrating automatic differentiation fundamentals rather than a production tensor library.

Where can I find the complete micrograd source code?

The complete implementation resides in the karpathy/nn-zero-to-hero repository. The core Value class and automatic differentiation logic are in lectures/micrograd/micrograd_lecture_first_half_roughly.ipynb (lines 72‑131), while visualization utilities appear in lectures/micrograd/micrograd_lecture_second_half_roughly.ipynb (lines 149‑169). The repository also includes a README.md with links to video lectures explaining the mathematical foundations of the automatic differentiation engine.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →