How to Train a LLM on a Single GPU Node with NanoChat

You can train a Transformer-based LLM on a single GPU by installing NanoChat via pip install -e ., configuring the --depth and --device-batch-size arguments to match your VRAM capacity, and executing python -m scripts.base_train to automatically apply compute-optimal scaling laws for batch size, learning rate, and training horizon.

NanoChat is a minimal, production-ready framework developed by Andrej Karpathy that enables pre-training GPT-style language models on hardware ranging from single consumer GPUs to multi-node clusters. By deriving optimal hyperparameters from scaling laws implemented in scripts/base_train.py, NanoChat lets you train a LLM on a single GPU node without manually tuning model dimensions, learning rates, or batch sizes.

Prerequisites and Environment Setup

Installing NanoChat from Source

The framework requires a Python environment with dependencies listed in pyproject.toml. Install the package in editable mode to ensure the nanochat module is available for training scripts.

pip install -e .

Configuring Compute Precision with NANOCHAT_DTYPE

Before launching training, set the NANOCHAT_DTYPE environment variable to control the compute dtype. On modern GPUs (SM 80+ such as H100), the framework defaults to bfloat16, but you can explicitly force float32 for debugging or float16 to enable GradScaler functionality.

export NANOCHAT_DTYPE=bfloat16  # Options: float32, float16, bfloat16

Configuring Single-GPU Training Parameters

Setting Model Depth with --depth

In nanochat/gpt.py, the Transformer architecture derives its model dimension from the --depth flag using compute-optimal aspect ratios. This is the primary hyperparameter you control; all other dimensions (width, number of heads, and embedding size) scale automatically.

  • Use --depth=12 for GPT-1 scale models
  • Use --depth=24 for GPT-2 scale models
  • Use --depth=26 or higher for larger experiments

Adjusting Batch Size for VRAM Limits

The default --device-batch-size=32 targets 80GB H100 GPUs. For smaller single-GPU nodes, reduce this value until the script fits in memory.


# For 24GB or 48GB GPUs

python -m scripts.base_train --depth=12 --device-batch-size=8

# For 80GB GPUs (default)

python -m scripts.base_train --depth=24 --device-batch-size=32

The script automatically calculates the optimal total batch size from scaling laws and implements a gradient accumulation loop to maintain compute optimality regardless of your per-GPU limit.

The Training Pipeline Internals

Scaling-Law Auto-Configuration in base_train.py

When you execute scripts/base_train.py, the script performs several automated steps before training begins:

  1. Meta-device initialization: The model is built on a meta device to calculate parameter counts without allocating full GPU memory
  2. Weight initialization: Parameters are initialized following GPT-2 schemes adapted for the calculated dimensions
  3. Training horizon calculation: The script derives the optimal number of iterations from --target-param-data-ratio (default), or accepts explicit overrides via --num-iterations or --target-flops
  4. Gradient accumulation setup: The training loop automatically accumulates gradients to reach the compute-optimal total batch size derived from Chinchilla scaling laws

Optimizer Setup and Learning Rate Scaling

The GPT.setup_optimizer method in nanochat/gpt.py implements a sophisticated parameter grouping strategy using the Muon + AdamW optimizer combination defined in nanochat/optim.py. Learning rates automatically scale by the square root of the model dimension ratio (√(dim/768)), and the script applies batch-size scaling factors to maintain training stability across different hardware configurations.

You can override specific learning rates using flags like --embedding-lr or --matrix-lr, though the defaults are compute-optimal for the detected architecture.

Executing the Training Run

For a standard single-GPU training run, use the Python module execution syntax. The --run flag configures Weights & Biases logging (use --run=dummy to disable).


# Train a 12-layer model on a single GPU

python -m scripts.base_train \
    --depth=12 \
    --run="single_gpu_d12" \
    --device-batch-size=8

If you are on a multi-GPU node but wish to restrict training to a single device, use torchrun with explicit process limiting:

torchrun --nproc_per_node=1 -m scripts.base_train \
    --depth=24 \
    --run="single_gpu_d24" \
    --device-batch-size=4

Enabling FP8 on H100 GPUs

For H100 or newer Hopper architecture GPUs, add the --fp8 flag to enable FP8 mixed precision training, which significantly increases throughput without impacting convergence.

python -m scripts.base_train --depth=24 --fp8 --device-batch-size=32

Checkpointing and Evaluation

The nanochat/checkpoint_manager.py module handles persistent storage of model weights, optimizer states, and dataloader position. By default, checkpoints save to base_checkpoints/d<depth>/ at the end of training, or at intervals specified by --save-every.

During training, scripts/base_train.py periodically evaluates validation bits-per-byte (BPB), the CORE metric, and generates sample text completions. Control evaluation frequency using:

  • --eval-every: Validation loss and BPB calculation
  • --core-metric-every: CORE benchmark evaluation
  • --sample-every: Text generation sampling

All metrics are logged automatically to Weights & Biuses unless you specify --run=dummy.

Summary

  • Install NanoChat using pip install -e . and optionally set NANOCHAT_DTYPE for precision control
  • Configure your single-GPU run using --depth for model size and --device-batch-size to fit VRAM constraints
  • Execute training via python -m scripts.base_train or torchrun --nproc_per_node=1 for explicit single-GPU targeting
  • Leverage automatic scaling-law calculations in scripts/base_train.py for batch size, learning rate, and training horizon
  • Optimize using the Muon + AdamW optimizer in nanochat/optim.py with automatic parameter grouping via GPT.setup_optimizer
  • Monitor progress through built-in validation BPB, CORE metrics, and sample generation logged to Weights & Biases
  • Checkpoint models using nanochat/checkpoint_manager.py with configurable save intervals via --save-every

Frequently Asked Questions

What is the minimum GPU memory required to train a LLM with NanoChat?

You can train smaller models (depth 12) on GPUs with 24GB-48GB VRAM by setting --device-batch-size to 4 or 8. The framework automatically handles gradient accumulation to maintain the compute-optimal total batch size, so you only need to ensure the model parameters and optimizer states fit in memory. For reference, the default --device-batch-size=32 assumes an 80GB H100.

How does NanoChat calculate the optimal training duration?

In scripts/base_train.py, the script calculates training iterations using scaling-law formulas based on the --target-param-data-ratio parameter by default, ensuring Chinchilla-optimal training. You can override this with --num-iterations for a fixed step count or --target-flops for a specific compute budget, allowing precise control over training horizon without manual calculation.

Can I use multiple GPUs even if I started with single-GPU training?

Yes, the same code supports distributed training without modification. Simply launch with torchrun --nproc_per_node=N where N matches your GPU count. The GPT.setup_optimizer function in nanochat/gpt.py automatically handles distributed-aware learning rate scaling and parameter sharding, ensuring seamless scaling from single-GPU experiments to multi-node clusters.

Why does NanoChat use both Muon and AdamW optimizers?

According to nanochat/optim.py, the framework applies Muon optimizer to matrix-shaped parameters (projections, embeddings) and AdamW to remaining parameters like layer norms and biases. This mixed strategy exploits the orthogonalization properties of Muon for weight matrices while maintaining the stability of AdamW for smaller parameter groups, as implemented in the GPT.setup_optimizer method.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →