What Is the Adam Optimizer and How Does It Differ From SGD?
Adam (Adaptive Moment Estimation) is a first-order gradient-based optimization algorithm that computes per-parameter adaptive learning rates using exponentially moving averages of past gradients and squared gradients, while SGD applies uniform learning rates without inherent per-parameter scaling.
The Adam optimizer has become the default choice for training modern deep neural networks, offering faster convergence and reduced hyperparameter tuning compared to classical stochastic gradient descent. In the nn-zero-to-hero educational repository by Andrej Karpathy, this algorithm is explicitly listed as a forthcoming topic in the curriculum roadmap. Understanding how Adam differs from SGD requires examining its dual moment estimation mechanism and adaptive step-size normalization.
How the Adam Optimizer Works: Moment Estimation
Adam maintains two moving averages per parameter to adapt the learning process dynamically. The first moment (m) estimates the mean of past gradients, functioning like momentum to smooth updates and accelerate convergence in consistent gradient directions. The second moment (v) estimates the uncentered variance of past gradients, scaling the learning rate inversely to the magnitude of recent gradients for each parameter.
The update rules contrast sharply with vanilla SGD:
- SGD:
θ ← θ – η·∇L(θ)(optionally with classical momentumv ← β·v + (1‑β)·∇L(θ)) - Adam:
m ← β₁·m + (1‑β₁)·∇L(θ)v ← β₂·v + (1‑β₂)·(∇L(θ))²θ ← θ – η·m̂ / (√v̂ + ε)
Here, m̂ and v̂ represent bias-corrected versions of the moments, calculated by dividing by (1 - β₁^t) and (1 - β₂^t) respectively to counteract initialization bias toward zero during early training steps.
Adam vs SGD: Key Technical Differences
Update Normalization: Adam automatically normalizes each parameter's update by its historical gradient variance, providing adaptive step sizes that shrink for parameters with large gradients and expand for those with small gradients. SGD lacks this per-coordinate scaling, requiring manual learning rate schedules or per-layer tuning to handle varying gradient magnitudes.
Momentum Handling: While SGD can incorporate momentum through a single velocity term, Adam embeds momentum naturally via the first moment estimate m with decay rate β₁ (typically 0.9), while simultaneously tracking second moments with β₂ (typically 0.999).
Memory Requirements: Adam stores two additional state tensors per parameter (m and v), consuming approximately twice the memory of SGD with momentum. This trade-off enables faster convergence and greater robustness to sparse or noisy gradients.
Robustness: The Adam optimizer demonstrates superior tolerance to noisy gradients and sparse data distributions, whereas SGD often requires careful learning rate annealing to prevent oscillation in high-curvature directions of the loss landscape.
Implementing the Adam Optimizer in PyTorch
The transition from SGD to Adam requires minimal code changes. The following examples demonstrate the implementation patterns compatible with the training loops shown in lectures/makemore/makemore_part5_cnn1.ipynb, which currently implements SGD-based training:
import torch
import torch.nn as nn
import torch.optim as optim
# Simple feed-forward network
model = nn.Sequential(
nn.Linear(784, 256),
nn.ReLU(),
nn.Linear(256, 10)
)
criterion = nn.CrossEntropyLoss()
inputs = torch.randn(64, 784)
targets = torch.randint(0, 10, (64,))
# SGD with momentum
sgd_opt = optim.SGD(model.parameters(), lr=0.01, momentum=0.9)
sgd_opt.zero_grad()
outputs = model(inputs)
loss = criterion(outputs, targets)
loss.backward()
sgd_opt.step()
# Adam optimizer - note the lower default learning rate
adam_opt = optim.Adam(model.parameters(), lr=0.001, betas=(0.9, 0.999))
adam_opt.zero_grad()
outputs = model(inputs)
loss = criterion(outputs, targets)
loss.backward()
adam_opt.step()
Key implementation details:
- Learning rate: Adam typically uses
lr=0.001compared to SGD's0.01or higher, as the adaptive scaling amplifies effective step sizes automatically. - Beta parameters: The default
betas=(0.9, 0.999)control the exponential decay rates for first and second moments respectively. - Epsilon: The
epsparameter (default1e-8) prevents division by zero when the second moment estimate is very small.
The Adam Optimizer in nn-zero-to-hero
According to the repository's README.md (specifically around line 40), the Adam optimizer is explicitly listed as a "notable todo" in the educational roadmap, positioned after batch normalization and residual connections. This placement reflects the curriculum's progression from foundational mechanics toward production-ready optimization.
The lectures/micrograd/micrograd_lecture_first_half_roughly.ipynb and lectures/micrograd/micrograd_lecture_second_half_roughly.ipynb notebooks establish the mathematical groundwork—implementing manual backpropagation and vanilla gradient descent—necessary to understand how Adam's moment estimates modify parameter updates. These files demonstrate the raw mechanics that Adam automates, while lectures/makemore/makemore_part5_cnn1.ipynb provides practical CNN training baselines where Adam can be substituted for SGD to observe convergence differences.
Summary
- Adam combines momentum (via first moment
m) and adaptive learning rates (via second momentv) to normalize parameter updates individually per coordinate. - The algorithm automatically scales steps based on historical gradient variance, eliminating the need for manual learning rate schedules required by SGD.
- Adam requires approximately twice the memory of SGD to store running estimates of gradient means and variances.
- In nn-zero-to-hero, Adam appears in the future curriculum roadmap, building upon the manual gradient descent implementations in the micrograd lecture series.
Frequently Asked Questions
Why does Adam use a lower default learning rate than SGD?
Adam's adaptive scaling through the second moment estimate (v) naturally amplifies step sizes for parameters with small gradients and dampens those with large gradients. Consequently, a base learning rate of 0.001 typically suffices, whereas SGD often requires 0.01 or higher to achieve comparable convergence speeds without the adaptive normalization.
Does Adam always outperform SGD?
No. While Adam generally converges faster during initial training phases and requires less hyperparameter tuning, SGD with momentum can achieve better final generalization on some computer vision benchmarks. Additionally, SGD consumes less memory—critical for training massive models with limited GPU resources—since it does not maintain separate first and second moment buffers for every parameter.
What are the recommended beta values for Adam?
The default PyTorch values of β₁ = 0.9 and β₂ = 0.999 work well for most deep learning applications. These coefficients control the exponential decay rates for the first and second moment estimates respectively, with β₂ closer to 1.0 ensuring stability for the second moment (squared gradients) which exhibits higher variance.
How does bias correction work in Adam?
Bias correction compensates for the initialization of moment estimates at zero. During early training steps, the uncorrected estimates m and v are biased toward zero. Adam divides these estimates by (1 - β₁^t) and (1 - β₂^t) respectively (where t is the step number) to obtain m̂ and v̂, ensuring accurate step sizes during the critical initial phase of training.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →