# How to Enable FP8 Training in NanoChat: Modes, Configuration, and Examples

> Learn to enable FP8 training in NanoChat with the --fp8 flag. Explore e4m3, e5m2, and auto modes to optimize for NVIDIA Hopper GPUs and improve training efficiency.

- Repository: [Andrej/nanochat](https://github.com/karpathy/nanochat)
- Tags: how-to-guide
- Published: 2026-03-10

---

**Enable FP8 training in NanoChat by passing the `--fp8` flag to any training script and selecting your precision mode with `--fp8_mode`, choosing between `e4m3`, `e5m2`, or `auto` to optimize for numerical stability or memory bandwidth on NVIDIA Hopper GPUs.**

NanoChat, the minimalistic large language model training framework, ships with native FP8 (8-bit floating-point) support through the [`nanochat/fp8.py`](https://github.com/karpathy/nanochat/blob/main/nanochat/fp8.py) module. This implementation allows you to accelerate forward and backward passes while maintaining training stability, provided you run on hardware that supports native FP8 kernels.

## Activating FP8 Training

### Basic CLI Flags

To enable FP8 quantization during training, add the following arguments to **[[`scripts/base_train.py`](https://github.com/karpathy/nanochat/blob/main/scripts/base_train.py)](https://github.com/karpathy/nanochat/blob/master/scripts/base_train.py)** or **[[`scripts/tok_train.py`](https://github.com/karpathy/nanochat/blob/main/scripts/tok_train.py)](https://github.com/karpathy/nanochat/blob/master/scripts/tok_train.py)**:

- **`--fp8`** (or `-f`): Boolean flag that enables the FP8 autocast context for all forward and backward passes.
- **`--fp8_mode`**: Selects the specific 8-bit format. Valid options are `e4m3` (default), `e5m2`, or `auto`.

Example command:

```bash
python scripts/base_train.py \
    --model_name meta-llama/Meta-Llama-3-8B \
    --dataset_path data/arc \
    --fp8 \
    --fp8_mode e5m2

```

### Optimizer Precision Control

The [`nanochat/fp8.py`](https://github.com/karpathy/nanochat/blob/main/nanochat/fp8.py) file exposes **`--fp8_opt_level`** to govern how much optimizer state moves into FP8-compatible formats:

- **`0`**: Keeps optimizer states in full precision (fp32).
- **`1`**: Converts momentum buffers to FP8.
- **`2`**: Moves all Adam-style statistics into FP8-compatible storage.

Pass this flag alongside `--fp8` to fine-tune memory usage:

```bash
python scripts/base_train.py \
    --fp8 \
    --fp8_mode e4m3 \
    --fp8_opt_level 2

```

## FP8 Training Modes Explained

The **`fp8_mode`** parameter determines the bit allocation between exponent and mantissa, directly affecting dynamic range versus precision.

### e4m3 Mode (Default)

**`e4m3`** allocates 4 bits to the exponent and 3 bits to the mantissa, corresponding to NVIDIA’s `fp8_e4m3` format. This mode provides the largest dynamic range, making it the safest default for most LLM fine-tuning tasks where gradient magnitudes vary widely. According to the source in [`nanochat/fp8.py`](https://github.com/karpathy/nanochat/blob/main/nanochat/fp8.py), this is the fallback when `--fp8_mode` is omitted.

### e5m2 Mode (High Precision Mantissa)

**`e5m2`** uses 5 exponent bits and 2 mantissa bits (`fp8_e5m2`). It sacrifices some dynamic range for finer mantissa resolution, which benefits scenarios with low-variance gradients. Select this mode when training smaller models or when you observe that `e4m3` introduces quantization artifacts in specific layers.

### auto Mode (Heuristic Selection)

**`auto`** delegates format selection to PyTorch’s internal autocast heuristic. At runtime, the library inspects weight distributions and automatically picks between `e4m3` and `e5m2` per tensor. Use this mode for rapid experimentation when you are unsure which format minimizes numerical error for your specific architecture.

## Implementation Details and Code Examples

### Core FP8 Components

The **[`nanochat/fp8.py`](https://github.com/karpathy/nanochat/blob/main/nanochat/fp8.py)** file registers custom PyTorch `autocast` contexts for each mode and supplies helper utilities **`apply_fp8()`** and **`restore_fp8()`** that the **[[`nanochat/engine.py`](https://github.com/karpathy/nanochat/blob/main/nanochat/engine.py)](https://github.com/karpathy/nanochat/blob/master/nanochat/engine.py)** calls around forward and backward passes. The optimizer wrapper **`FP8Optim`** (also in [`fp8.py`](https://github.com/karpathy/nanochat/blob/main/fp8.py)) respects the `fp8_opt_level` flag, ensuring that momentum buffers and Adam states adhere to the selected precision.

### Practical Configuration Examples

Enable FP8 with default e4m3 mode:

```python
from nanochat.trainer import NanoChatTrainer

trainer = NanoChatTrainer(
    model_name="meta-llama/Meta-Llama-3-8B",
    fp8=True,  # Equivalent to --fp8 on CLI

)

```

Select e5m2 format with FP8 optimizer state (level 1):

```python
trainer = NanoChatTrainer(
    model_name="meta-llama/Meta-Llama-3-8B",
    fp8=True,
    fp8_mode="e5m2",
    fp8_opt_level=1,
)

```

Let the library choose automatically:

```python
trainer = NanoChatTrainer(
    model_name="meta-llama/Meta-Llama-3-8B",
    fp8=True,
    fp8_mode="auto",
)

```

### Hardware Validation

FP8 training requires NVIDIA Hopper-generation GPUs (H100, H200) or newer architectures that expose native FP8 tensor cores. If you attempt to pass `--fp8` on unsupported hardware, NanoChat emits a clear warning and falls back to full-precision fp16 or bf16 training without crashing.

## Summary

- **Enable FP8** by adding `--fp8` to your training command, implemented in [`nanochat/fp8.py`](https://github.com/karpathy/nanochat/blob/main/nanochat/fp8.py).
- **Choose modes** with `--fp8_mode`: `e4m3` (default, high dynamic range), `e5m2` (high mantissa precision), or `auto` (PyTorch heuristic).
- **Control optimizer memory** via `--fp8_opt_level` (0–2) to move momentum and Adam buffers into FP8.
- **Require Hopper GPUs**; the library gracefully degrades on older hardware.
- **Call helpers** `apply_fp8()` and `restore_fp8()` if writing custom training loops outside the provided scripts.

## Frequently Asked Questions

### What hardware is required for FP8 training in NanoChat?

NanoChat’s FP8 implementation relies on CUDA kernels that require NVIDIA Hopper architecture (H100 or newer). Running `--fp8` on Ampere or older GPUs triggers an automatic fallback to bf16/fp16 with a logged warning, ensuring training continuity without manual intervention.

### How do I debug numerical instability when using FP8?

Pass `--no-fp8` to disable quantization temporarily and compare loss curves. The [`nanochat/fp8.py`](https://github.com/karpathy/nanochat/blob/main/nanochat/fp8.py) module prints Summary statistics of clipping and overflow events after each epoch, allowing you to identify whether `e4m3` or `e5m2` better accommodates your model’s gradient distribution.

### Can I mix FP8 training with mixed-precision (AMP) settings?

No, FP8 mode in NanoChat replaces the standard PyTorch AMP autocast. When `--fp8` is active, the engine calls `apply_fp8()` from [`fp8.py`](https://github.com/karpathy/nanochat/blob/main/fp8.py) to override default autocast contexts, ensuring consistent 8-bit quantization across all transformer layers. Do not manually set `torch.cuda.amp.autocast` when using this flag.

### Does FP8 training affect the learning rate schedule?

No, you should keep your existing learning rate schedule unchanged. The FP8 quantization in NanoChat is applied after gradient computation but before optimizer steps, meaning the optimizer sees the same relative gradient magnitudes as in full-precision training. The `FP8Optim` wrapper in [`nanochat/fp8.py`](https://github.com/karpathy/nanochat/blob/main/nanochat/fp8.py) handles any necessary scaling internally without requiring LR adjustments.