# RAdam vs Adam Optimizer: Key Differences for Stable Deep Learning Training

> Discover the key differences between RAdam and Adam optimizers for stable deep learning training. RAdam offers automatic warm-up and faster convergence.

- Repository: [labml.ai/annotated_deep_learning_paper_implementations](https://github.com/labmlai/annotated_deep_learning_paper_implementations)
- Tags: deep-dive
- Published: 2026-03-04

---

**RAdam (Rectified Adam) is a variance-aware refinement of the standard Adam optimizer that automatically stabilizes early training through an adaptive rectification term, eliminating the need for manual warm-up schedules while maintaining fast convergence.**

The **Adam** optimizer has become the default choice for training deep neural networks, but its adaptive learning rate can suffer from high variance during initial training steps. The **Rectified Adam (RAdam)** optimizer, as implemented in the `labmlai/annotated_deep_learning_paper_implementations` repository, addresses this limitation by introducing a rectification mechanism that controls variance based on the effective number of past gradients. Understanding the difference between RAdam and the standard Adam optimizer helps practitioners achieve more stable convergence without hand-crafted learning rate schedules.

## Core Algorithmic Differences

### Variance Rectification Mechanism

Standard **Adam** computes bias-corrected first-moment *m* and second-moment *v* estimates, applying the update rule with an adaptive learning rate scaled by √vₜ. During early training, this adaptive term exhibits high variance because the exponential moving average of squared gradients relies on few samples, potentially causing unstable updates or convergence to poor local minima.

**RAdam** retains the same *m* and *v* calculations but introduces a **rectification term** *rₜ* that rescales the update based on the *effective* number of past gradients, denoted as ρₜ (rho). According to the implementation in [`labml_nn/optimizers/radam.py`](https://github.com/labmlai/annotated_deep_learning_paper_implementations/blob/main/labml_nn/optimizers/radam.py), when ρₜ ≥ 5, the update is multiplied by √rₜ, where *rₜ* derives from the variance of the adaptive learning rate. This rectification dampens the step size when variance is high and gradually permits full Adam-like updates as training progresses.

### Warm-Up Behavior and Early Training Stability

Standard **Adam** frequently requires manual learning-rate warm-up—typically linearly increasing the learning rate over the first few thousand steps—to mitigate high variance in early adaptive learning rates. This warm-up proves particularly critical when training Transformer architectures.

**RAdam**'s rectification term effectively provides **automatic warm-up**. By dynamically scaling the learning rate based on the reliability of the second-moment estimate, RAdam stabilizes the optimizer from the first step without requiring hand-engineered warm-up schedules. This makes RAdam particularly effective for deep models where warm-up hyperparameters are difficult to tune.

### Fallback to SGD with Momentum

A distinctive feature of **RAdam** absent in standard **Adam** is its ability to **degenerate to SGD with momentum** during the earliest training steps. When ρₜ < 5, the rectification term becomes undefined (None), and if the `degenerate_to_sgd=True` flag is set, RAdam falls back to a simple momentum update rather than applying an unreliable adaptive learning rate.

Standard **Adam** possesses no such fallback mechanism—it always applies the adaptive update, even when the variance estimate is based on insufficient data.

## Implementation Details in LabML

### File Structure and Inheritance

In the `labmlai/annotated_deep_learning_paper_implementations` codebase, the optimizers follow a clear inheritance hierarchy demonstrating the architectural evolution:

- **[`labml_nn/optimizers/adam.py`](https://github.com/labmlai/annotated_deep_learning_paper_implementations/blob/main/labml_nn/optimizers/adam.py)**: Contains the reference implementation of standard Adam with bias-corrected moment estimates and optional optimized updates.
- **[`labml_nn/optimizers/amsgrad.py`](https://github.com/labmlai/annotated_deep_learning_paper_implementations/blob/main/labml_nn/optimizers/amsgrad.py)**: Provides the base AMSGrad functionality that RAdam extends.
- **[`labml_nn/optimizers/radam.py`](https://github.com/labmlai/annotated_deep_learning_paper_implementations/blob/main/labml_nn/optimizers/radam.py)**: Implements RAdam by inheriting from the AMSGrad base class and overriding the `step_param` logic to inject the rectification calculation.

### The Rectification Term Calculation

The critical difference in the source code lies in the `calc_rectification_term` method within [`labml_nn/optimizers/radam.py`](https://github.com/labmlai/annotated_deep_learning_paper_implementations/blob/main/labml_nn/optimizers/radam.py) (source lines 2003‑2019). This method computes the effective number of past gradients ρₜ and derives the rectification factor *rₜ*. The method returns the scaling factor √rₜ when tractable (ρₜ ≥ 5), otherwise returning `None` to trigger the SGD fallback path.

Both optimizers share the same core hyperparameters—`lr`, `betas`, `eps`, and `weight_decay`—but RAdam adds the `degenerate_to_sgd` boolean and `optimized_update` flag for computational efficiency.

## Practical Usage and Code Examples

Instantiating standard **Adam** follows the familiar PyTorch pattern:

```python
from labml_nn.optimizers import Adam
optimizer = Adam(model.parameters(),
                lr=1e-3,
                betas=(0.9, 0.999),
                eps=1e-8,
                weight_decay=WeightDecay())   # optional weight decay

```

For **RAdam**, the API is nearly identical, with the addition of the fallback option:

```python
from labml_nn.optimizers import RAdam
optimizer = RAdam(model.parameters(),
                  lr=1e-3,
                  betas=(0.9, 0.999),
                  eps=1e-8,
                  weight_decay=WeightDecay(),
                  degenerate_to_sgd=True)   # fall back to SGD when rectification is undefined

```

Both optimizers integrate seamlessly into standard PyTorch training loops: call `optimizer.zero_grad()`, compute your loss, execute `loss.backward()`, and step the optimizer with `optimizer.step()`.

## Summary

- **RAdam = Adam + variance rectification**: It uses identical first and second moment estimates but adds a rectification term *rₜ* based on the effective number of past gradients ρₜ.
- **Automatic stabilization**: RAdam's √rₜ scaling mitigates high variance during early training (ρₜ < 5), eliminating the manual warm-up schedules often required by standard Adam.
- **SGD fallback**: When variance estimates are unreliable (ρₜ < 5), RAdam can degenerate to SGD with momentum via the `degenerate_to_sgd` parameter, while Adam always applies adaptive updates regardless of estimate reliability.
- **Implementation architecture**: RAdam extends the AMSGrad base class in [`labml_nn/optimizers/radam.py`](https://github.com/labmlai/annotated_deep_learning_paper_implementations/blob/main/labml_nn/optimizers/radam.py), overriding `step_param` and implementing `calc_rectification_term` to inject variance-aware logic.

## Frequently Asked Questions

### Does RAdam completely replace the need for learning rate warm-up?

In many cases, yes. The rectification term in RAdam automatically reduces the effective learning rate when the variance of the adaptive term is high during early training steps. This provides an implicit warm-up effect that often eliminates the need for manual linear warm-up schedules commonly used with standard Adam, particularly when training Transformers or deep convolutional networks.

### When should I use the `degenerate_to_sgd` option in RAdam?

Enable `degenerate_to_sgd=True` when you want maximum stability during the very first training steps (typically fewer than 5 iterations, where ρₜ < 5). In this regime, the adaptive learning rate variance is undefined or unreliable, so falling back to standard SGD with momentum prevents erratic updates. Once sufficient gradients accumulate (ρₜ ≥ 5), RAdam automatically transitions to full adaptive updates.

### Is RAdam slower or more memory-intensive than standard Adam?

No. RAdam incurs negligible computational overhead compared to standard Adam. Both optimizers maintain the same two momentum buffers (first and second moments). The additional calculation of the rectification term *rₜ* involves simple scalar operations on ρₜ, making the performance and memory footprint virtually identical to Adam in the `labml_nn` implementation.

### Can I switch from Adam to RAdam mid-training without issues?

While technically possible, switching optimizers mid-training is generally not recommended as it disrupts the momentum buffers and variance estimates. If migrating to RAdam, it is best to start training from scratch or from a checkpoint with reset optimizer states. The `degenerate_to_sgd` feature is specifically designed to handle early training stability, which would be bypassed if you switch after the warm-up phase.