# How to Train a LLM on a Single GPU Node with NanoChat

> Train a Transformer LLM on a single GPU node with NanoChat. Learn to install, configure, and run training scripts for efficient LLM development with compute-optimal scaling.

- Repository: [Andrej/nanochat](https://github.com/karpathy/nanochat)
- Tags: tutorial
- Published: 2026-03-10

---

**You can train a Transformer-based LLM on a single GPU by installing NanoChat via `pip install -e .`, configuring the `--depth` and `--device-batch-size` arguments to match your VRAM capacity, and executing `python -m scripts.base_train` to automatically apply compute-optimal scaling laws for batch size, learning rate, and training horizon.**

NanoChat is a minimal, production-ready framework developed by Andrej Karpathy that enables pre-training GPT-style language models on hardware ranging from single consumer GPUs to multi-node clusters. By deriving optimal hyperparameters from scaling laws implemented in [`scripts/base_train.py`](https://github.com/karpathy/nanochat/blob/main/scripts/base_train.py), NanoChat lets you **train a LLM on a single GPU node** without manually tuning model dimensions, learning rates, or batch sizes.

## Prerequisites and Environment Setup

### Installing NanoChat from Source

The framework requires a Python environment with dependencies listed in [`pyproject.toml`](https://github.com/karpathy/nanochat/blob/main/pyproject.toml). Install the package in editable mode to ensure the `nanochat` module is available for training scripts.

```bash
pip install -e .

```

### Configuring Compute Precision with NANOCHAT_DTYPE

Before launching training, set the `NANOCHAT_DTYPE` environment variable to control the compute dtype. On modern GPUs (SM 80+ such as H100), the framework defaults to `bfloat16`, but you can explicitly force `float32` for debugging or `float16` to enable `GradScaler` functionality.

```bash
export NANOCHAT_DTYPE=bfloat16  # Options: float32, float16, bfloat16

```

## Configuring Single-GPU Training Parameters

### Setting Model Depth with `--depth`

In [`nanochat/gpt.py`](https://github.com/karpathy/nanochat/blob/main/nanochat/gpt.py), the Transformer architecture derives its model dimension from the `--depth` flag using compute-optimal aspect ratios. This is the primary hyperparameter you control; all other dimensions (width, number of heads, and embedding size) scale automatically.

- Use `--depth=12` for GPT-1 scale models
- Use `--depth=24` for GPT-2 scale models  
- Use `--depth=26` or higher for larger experiments

### Adjusting Batch Size for VRAM Limits

The default `--device-batch-size=32` targets 80GB H100 GPUs. For smaller single-GPU nodes, reduce this value until the script fits in memory.

```bash

# For 24GB or 48GB GPUs

python -m scripts.base_train --depth=12 --device-batch-size=8

# For 80GB GPUs (default)

python -m scripts.base_train --depth=24 --device-batch-size=32

```

The script automatically calculates the optimal **total batch size** from scaling laws and implements a gradient accumulation loop to maintain compute optimality regardless of your per-GPU limit.

## The Training Pipeline Internals

### Scaling-Law Auto-Configuration in base_train.py

When you execute [`scripts/base_train.py`](https://github.com/karpathy/nanochat/blob/main/scripts/base_train.py), the script performs several automated steps before training begins:

1. **Meta-device initialization**: The model is built on a meta device to calculate parameter counts without allocating full GPU memory
2. **Weight initialization**: Parameters are initialized following GPT-2 schemes adapted for the calculated dimensions
3. **Training horizon calculation**: The script derives the optimal number of iterations from `--target-param-data-ratio` (default), or accepts explicit overrides via `--num-iterations` or `--target-flops`
4. **Gradient accumulation setup**: The training loop automatically accumulates gradients to reach the compute-optimal total batch size derived from Chinchilla scaling laws

### Optimizer Setup and Learning Rate Scaling

The `GPT.setup_optimizer` method in [`nanochat/gpt.py`](https://github.com/karpathy/nanochat/blob/main/nanochat/gpt.py) implements a sophisticated parameter grouping strategy using the **Muon + AdamW** optimizer combination defined in [`nanochat/optim.py`](https://github.com/karpathy/nanochat/blob/main/nanochat/optim.py). Learning rates automatically scale by the square root of the model dimension ratio (√(dim/768)), and the script applies batch-size scaling factors to maintain training stability across different hardware configurations.

You can override specific learning rates using flags like `--embedding-lr` or `--matrix-lr`, though the defaults are compute-optimal for the detected architecture.

## Executing the Training Run

For a standard single-GPU training run, use the Python module execution syntax. The `--run` flag configures Weights & Biases logging (use `--run=dummy` to disable).

```bash

# Train a 12-layer model on a single GPU

python -m scripts.base_train \
    --depth=12 \
    --run="single_gpu_d12" \
    --device-batch-size=8

```

If you are on a multi-GPU node but wish to restrict training to a single device, use `torchrun` with explicit process limiting:

```bash
torchrun --nproc_per_node=1 -m scripts.base_train \
    --depth=24 \
    --run="single_gpu_d24" \
    --device-batch-size=4

```

### Enabling FP8 on H100 GPUs

For H100 or newer Hopper architecture GPUs, add the `--fp8` flag to enable FP8 mixed precision training, which significantly increases throughput without impacting convergence.

```bash
python -m scripts.base_train --depth=24 --fp8 --device-batch-size=32

```

## Checkpointing and Evaluation

The [`nanochat/checkpoint_manager.py`](https://github.com/karpathy/nanochat/blob/main/nanochat/checkpoint_manager.py) module handles persistent storage of model weights, optimizer states, and dataloader position. By default, checkpoints save to `base_checkpoints/d<depth>/` at the end of training, or at intervals specified by `--save-every`.

During training, [`scripts/base_train.py`](https://github.com/karpathy/nanochat/blob/main/scripts/base_train.py) periodically evaluates **validation bits-per-byte (BPB)**, the **CORE metric**, and generates sample text completions. Control evaluation frequency using:

- `--eval-every`: Validation loss and BPB calculation
- `--core-metric-every`: CORE benchmark evaluation  
- `--sample-every`: Text generation sampling

All metrics are logged automatically to Weights & Biuses unless you specify `--run=dummy`.

## Summary

- **Install** NanoChat using `pip install -e .` and optionally set `NANOCHAT_DTYPE` for precision control
- **Configure** your single-GPU run using `--depth` for model size and `--device-batch-size` to fit VRAM constraints
- **Execute** training via `python -m scripts.base_train` or `torchrun --nproc_per_node=1` for explicit single-GPU targeting
- **Leverage** automatic scaling-law calculations in [`scripts/base_train.py`](https://github.com/karpathy/nanochat/blob/main/scripts/base_train.py) for batch size, learning rate, and training horizon
- **Optimize** using the Muon + AdamW optimizer in [`nanochat/optim.py`](https://github.com/karpathy/nanochat/blob/main/nanochat/optim.py) with automatic parameter grouping via `GPT.setup_optimizer`
- **Monitor** progress through built-in validation BPB, CORE metrics, and sample generation logged to Weights & Biases
- **Checkpoint** models using [`nanochat/checkpoint_manager.py`](https://github.com/karpathy/nanochat/blob/main/nanochat/checkpoint_manager.py) with configurable save intervals via `--save-every`

## Frequently Asked Questions

### What is the minimum GPU memory required to train a LLM with NanoChat?

You can train smaller models (depth 12) on GPUs with 24GB-48GB VRAM by setting `--device-batch-size` to 4 or 8. The framework automatically handles gradient accumulation to maintain the compute-optimal total batch size, so you only need to ensure the model parameters and optimizer states fit in memory. For reference, the default `--device-batch-size=32` assumes an 80GB H100.

### How does NanoChat calculate the optimal training duration?

In [`scripts/base_train.py`](https://github.com/karpathy/nanochat/blob/main/scripts/base_train.py), the script calculates training iterations using scaling-law formulas based on the `--target-param-data-ratio` parameter by default, ensuring Chinchilla-optimal training. You can override this with `--num-iterations` for a fixed step count or `--target-flops` for a specific compute budget, allowing precise control over training horizon without manual calculation.

### Can I use multiple GPUs even if I started with single-GPU training?

Yes, the same code supports distributed training without modification. Simply launch with `torchrun --nproc_per_node=N` where N matches your GPU count. The `GPT.setup_optimizer` function in [`nanochat/gpt.py`](https://github.com/karpathy/nanochat/blob/main/nanochat/gpt.py) automatically handles distributed-aware learning rate scaling and parameter sharding, ensuring seamless scaling from single-GPU experiments to multi-node clusters.

### Why does NanoChat use both Muon and AdamW optimizers?

According to [`nanochat/optim.py`](https://github.com/karpathy/nanochat/blob/main/nanochat/optim.py), the framework applies **Muon** optimizer to matrix-shaped parameters (projections, embeddings) and **AdamW** to remaining parameters like layer norms and biases. This mixed strategy exploits the orthogonalization properties of Muon for weight matrices while maintaining the stability of AdamW for smaller parameter groups, as implemented in the `GPT.setup_optimizer` method.