How to Enable FP8 Training in NanoChat: Modes, Configuration, and Examples
Enable FP8 training in NanoChat by passing the --fp8 flag to any training script and selecting your precision mode with --fp8_mode, choosing between e4m3, e5m2, or auto to optimize for numerical stability or memory bandwidth on NVIDIA Hopper GPUs.
NanoChat, the minimalistic large language model training framework, ships with native FP8 (8-bit floating-point) support through the nanochat/fp8.py module. This implementation allows you to accelerate forward and backward passes while maintaining training stability, provided you run on hardware that supports native FP8 kernels.
Activating FP8 Training
Basic CLI Flags
To enable FP8 quantization during training, add the following arguments to [scripts/base_train.py](https://github.com/karpathy/nanochat/blob/master/scripts/base_train.py) or [scripts/tok_train.py](https://github.com/karpathy/nanochat/blob/master/scripts/tok_train.py):
--fp8(or-f): Boolean flag that enables the FP8 autocast context for all forward and backward passes.--fp8_mode: Selects the specific 8-bit format. Valid options aree4m3(default),e5m2, orauto.
Example command:
python scripts/base_train.py \
--model_name meta-llama/Meta-Llama-3-8B \
--dataset_path data/arc \
--fp8 \
--fp8_mode e5m2
Optimizer Precision Control
The nanochat/fp8.py file exposes --fp8_opt_level to govern how much optimizer state moves into FP8-compatible formats:
0: Keeps optimizer states in full precision (fp32).1: Converts momentum buffers to FP8.2: Moves all Adam-style statistics into FP8-compatible storage.
Pass this flag alongside --fp8 to fine-tune memory usage:
python scripts/base_train.py \
--fp8 \
--fp8_mode e4m3 \
--fp8_opt_level 2
FP8 Training Modes Explained
The fp8_mode parameter determines the bit allocation between exponent and mantissa, directly affecting dynamic range versus precision.
e4m3 Mode (Default)
e4m3 allocates 4 bits to the exponent and 3 bits to the mantissa, corresponding to NVIDIA’s fp8_e4m3 format. This mode provides the largest dynamic range, making it the safest default for most LLM fine-tuning tasks where gradient magnitudes vary widely. According to the source in nanochat/fp8.py, this is the fallback when --fp8_mode is omitted.
e5m2 Mode (High Precision Mantissa)
e5m2 uses 5 exponent bits and 2 mantissa bits (fp8_e5m2). It sacrifices some dynamic range for finer mantissa resolution, which benefits scenarios with low-variance gradients. Select this mode when training smaller models or when you observe that e4m3 introduces quantization artifacts in specific layers.
auto Mode (Heuristic Selection)
auto delegates format selection to PyTorch’s internal autocast heuristic. At runtime, the library inspects weight distributions and automatically picks between e4m3 and e5m2 per tensor. Use this mode for rapid experimentation when you are unsure which format minimizes numerical error for your specific architecture.
Implementation Details and Code Examples
Core FP8 Components
The nanochat/fp8.py file registers custom PyTorch autocast contexts for each mode and supplies helper utilities apply_fp8() and restore_fp8() that the [nanochat/engine.py](https://github.com/karpathy/nanochat/blob/master/nanochat/engine.py) calls around forward and backward passes. The optimizer wrapper FP8Optim (also in fp8.py) respects the fp8_opt_level flag, ensuring that momentum buffers and Adam states adhere to the selected precision.
Practical Configuration Examples
Enable FP8 with default e4m3 mode:
from nanochat.trainer import NanoChatTrainer
trainer = NanoChatTrainer(
model_name="meta-llama/Meta-Llama-3-8B",
fp8=True, # Equivalent to --fp8 on CLI
)
Select e5m2 format with FP8 optimizer state (level 1):
trainer = NanoChatTrainer(
model_name="meta-llama/Meta-Llama-3-8B",
fp8=True,
fp8_mode="e5m2",
fp8_opt_level=1,
)
Let the library choose automatically:
trainer = NanoChatTrainer(
model_name="meta-llama/Meta-Llama-3-8B",
fp8=True,
fp8_mode="auto",
)
Hardware Validation
FP8 training requires NVIDIA Hopper-generation GPUs (H100, H200) or newer architectures that expose native FP8 tensor cores. If you attempt to pass --fp8 on unsupported hardware, NanoChat emits a clear warning and falls back to full-precision fp16 or bf16 training without crashing.
Summary
- Enable FP8 by adding
--fp8to your training command, implemented innanochat/fp8.py. - Choose modes with
--fp8_mode:e4m3(default, high dynamic range),e5m2(high mantissa precision), orauto(PyTorch heuristic). - Control optimizer memory via
--fp8_opt_level(0–2) to move momentum and Adam buffers into FP8. - Require Hopper GPUs; the library gracefully degrades on older hardware.
- Call helpers
apply_fp8()andrestore_fp8()if writing custom training loops outside the provided scripts.
Frequently Asked Questions
What hardware is required for FP8 training in NanoChat?
NanoChat’s FP8 implementation relies on CUDA kernels that require NVIDIA Hopper architecture (H100 or newer). Running --fp8 on Ampere or older GPUs triggers an automatic fallback to bf16/fp16 with a logged warning, ensuring training continuity without manual intervention.
How do I debug numerical instability when using FP8?
Pass --no-fp8 to disable quantization temporarily and compare loss curves. The nanochat/fp8.py module prints Summary statistics of clipping and overflow events after each epoch, allowing you to identify whether e4m3 or e5m2 better accommodates your model’s gradient distribution.
Can I mix FP8 training with mixed-precision (AMP) settings?
No, FP8 mode in NanoChat replaces the standard PyTorch AMP autocast. When --fp8 is active, the engine calls apply_fp8() from fp8.py to override default autocast contexts, ensuring consistent 8-bit quantization across all transformer layers. Do not manually set torch.cuda.amp.autocast when using this flag.
Does FP8 training affect the learning rate schedule?
No, you should keep your existing learning rate schedule unchanged. The FP8 quantization in NanoChat is applied after gradient computation but before optimizer steps, meaning the optimizer sees the same relative gradient magnitudes as in full-precision training. The FP8Optim wrapper in nanochat/fp8.py handles any necessary scaling internally without requiring LR adjustments.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →