RAdam vs Adam Optimizer: Key Differences for Stable Deep Learning Training

RAdam (Rectified Adam) is a variance-aware refinement of the standard Adam optimizer that automatically stabilizes early training through an adaptive rectification term, eliminating the need for manual warm-up schedules while maintaining fast convergence.

The Adam optimizer has become the default choice for training deep neural networks, but its adaptive learning rate can suffer from high variance during initial training steps. The Rectified Adam (RAdam) optimizer, as implemented in the labmlai/annotated_deep_learning_paper_implementations repository, addresses this limitation by introducing a rectification mechanism that controls variance based on the effective number of past gradients. Understanding the difference between RAdam and the standard Adam optimizer helps practitioners achieve more stable convergence without hand-crafted learning rate schedules.

Core Algorithmic Differences

Variance Rectification Mechanism

Standard Adam computes bias-corrected first-moment m and second-moment v estimates, applying the update rule with an adaptive learning rate scaled by √vₜ. During early training, this adaptive term exhibits high variance because the exponential moving average of squared gradients relies on few samples, potentially causing unstable updates or convergence to poor local minima.

RAdam retains the same m and v calculations but introduces a rectification term rₜ that rescales the update based on the effective number of past gradients, denoted as ρₜ (rho). According to the implementation in labml_nn/optimizers/radam.py, when ρₜ ≥ 5, the update is multiplied by √rₜ, where rₜ derives from the variance of the adaptive learning rate. This rectification dampens the step size when variance is high and gradually permits full Adam-like updates as training progresses.

Warm-Up Behavior and Early Training Stability

Standard Adam frequently requires manual learning-rate warm-up—typically linearly increasing the learning rate over the first few thousand steps—to mitigate high variance in early adaptive learning rates. This warm-up proves particularly critical when training Transformer architectures.

RAdam's rectification term effectively provides automatic warm-up. By dynamically scaling the learning rate based on the reliability of the second-moment estimate, RAdam stabilizes the optimizer from the first step without requiring hand-engineered warm-up schedules. This makes RAdam particularly effective for deep models where warm-up hyperparameters are difficult to tune.

Fallback to SGD with Momentum

A distinctive feature of RAdam absent in standard Adam is its ability to degenerate to SGD with momentum during the earliest training steps. When ρₜ < 5, the rectification term becomes undefined (None), and if the degenerate_to_sgd=True flag is set, RAdam falls back to a simple momentum update rather than applying an unreliable adaptive learning rate.

Standard Adam possesses no such fallback mechanism—it always applies the adaptive update, even when the variance estimate is based on insufficient data.

Implementation Details in LabML

File Structure and Inheritance

In the labmlai/annotated_deep_learning_paper_implementations codebase, the optimizers follow a clear inheritance hierarchy demonstrating the architectural evolution:

The Rectification Term Calculation

The critical difference in the source code lies in the calc_rectification_term method within labml_nn/optimizers/radam.py (source lines 2003‑2019). This method computes the effective number of past gradients ρₜ and derives the rectification factor rₜ. The method returns the scaling factor √rₜ when tractable (ρₜ ≥ 5), otherwise returning None to trigger the SGD fallback path.

Both optimizers share the same core hyperparameters—lr, betas, eps, and weight_decay—but RAdam adds the degenerate_to_sgd boolean and optimized_update flag for computational efficiency.

Practical Usage and Code Examples

Instantiating standard Adam follows the familiar PyTorch pattern:

from labml_nn.optimizers import Adam
optimizer = Adam(model.parameters(),
                lr=1e-3,
                betas=(0.9, 0.999),
                eps=1e-8,
                weight_decay=WeightDecay())   # optional weight decay

For RAdam, the API is nearly identical, with the addition of the fallback option:

from labml_nn.optimizers import RAdam
optimizer = RAdam(model.parameters(),
                  lr=1e-3,
                  betas=(0.9, 0.999),
                  eps=1e-8,
                  weight_decay=WeightDecay(),
                  degenerate_to_sgd=True)   # fall back to SGD when rectification is undefined

Both optimizers integrate seamlessly into standard PyTorch training loops: call optimizer.zero_grad(), compute your loss, execute loss.backward(), and step the optimizer with optimizer.step().

Summary

  • RAdam = Adam + variance rectification: It uses identical first and second moment estimates but adds a rectification term rₜ based on the effective number of past gradients ρₜ.
  • Automatic stabilization: RAdam's √rₜ scaling mitigates high variance during early training (ρₜ < 5), eliminating the manual warm-up schedules often required by standard Adam.
  • SGD fallback: When variance estimates are unreliable (ρₜ < 5), RAdam can degenerate to SGD with momentum via the degenerate_to_sgd parameter, while Adam always applies adaptive updates regardless of estimate reliability.
  • Implementation architecture: RAdam extends the AMSGrad base class in labml_nn/optimizers/radam.py, overriding step_param and implementing calc_rectification_term to inject variance-aware logic.

Frequently Asked Questions

Does RAdam completely replace the need for learning rate warm-up?

In many cases, yes. The rectification term in RAdam automatically reduces the effective learning rate when the variance of the adaptive term is high during early training steps. This provides an implicit warm-up effect that often eliminates the need for manual linear warm-up schedules commonly used with standard Adam, particularly when training Transformers or deep convolutional networks.

When should I use the degenerate_to_sgd option in RAdam?

Enable degenerate_to_sgd=True when you want maximum stability during the very first training steps (typically fewer than 5 iterations, where ρₜ < 5). In this regime, the adaptive learning rate variance is undefined or unreliable, so falling back to standard SGD with momentum prevents erratic updates. Once sufficient gradients accumulate (ρₜ ≥ 5), RAdam automatically transitions to full adaptive updates.

Is RAdam slower or more memory-intensive than standard Adam?

No. RAdam incurs negligible computational overhead compared to standard Adam. Both optimizers maintain the same two momentum buffers (first and second moments). The additional calculation of the rectification term rₜ involves simple scalar operations on ρₜ, making the performance and memory footprint virtually identical to Adam in the labml_nn implementation.

Can I switch from Adam to RAdam mid-training without issues?

While technically possible, switching optimizers mid-training is generally not recommended as it disrupts the momentum buffers and variance estimates. If migrating to RAdam, it is best to start training from scratch or from a checkpoint with reset optimizer states. The degenerate_to_sgd feature is specifically designed to handle early training stability, which would be bypassed if you switch after the warm-up phase.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →