How to Train a LLM on a Single GPU Node with NanoChat
You can train a Transformer-based LLM on a single GPU by installing NanoChat via pip install -e ., configuring the --depth and --device-batch-size arguments to match your VRAM capacity, and executing python -m scripts.base_train to automatically apply compute-optimal scaling laws for batch size, learning rate, and training horizon.
NanoChat is a minimal, production-ready framework developed by Andrej Karpathy that enables pre-training GPT-style language models on hardware ranging from single consumer GPUs to multi-node clusters. By deriving optimal hyperparameters from scaling laws implemented in scripts/base_train.py, NanoChat lets you train a LLM on a single GPU node without manually tuning model dimensions, learning rates, or batch sizes.
Prerequisites and Environment Setup
Installing NanoChat from Source
The framework requires a Python environment with dependencies listed in pyproject.toml. Install the package in editable mode to ensure the nanochat module is available for training scripts.
pip install -e .
Configuring Compute Precision with NANOCHAT_DTYPE
Before launching training, set the NANOCHAT_DTYPE environment variable to control the compute dtype. On modern GPUs (SM 80+ such as H100), the framework defaults to bfloat16, but you can explicitly force float32 for debugging or float16 to enable GradScaler functionality.
export NANOCHAT_DTYPE=bfloat16 # Options: float32, float16, bfloat16
Configuring Single-GPU Training Parameters
Setting Model Depth with --depth
In nanochat/gpt.py, the Transformer architecture derives its model dimension from the --depth flag using compute-optimal aspect ratios. This is the primary hyperparameter you control; all other dimensions (width, number of heads, and embedding size) scale automatically.
- Use
--depth=12for GPT-1 scale models - Use
--depth=24for GPT-2 scale models - Use
--depth=26or higher for larger experiments
Adjusting Batch Size for VRAM Limits
The default --device-batch-size=32 targets 80GB H100 GPUs. For smaller single-GPU nodes, reduce this value until the script fits in memory.
# For 24GB or 48GB GPUs
python -m scripts.base_train --depth=12 --device-batch-size=8
# For 80GB GPUs (default)
python -m scripts.base_train --depth=24 --device-batch-size=32
The script automatically calculates the optimal total batch size from scaling laws and implements a gradient accumulation loop to maintain compute optimality regardless of your per-GPU limit.
The Training Pipeline Internals
Scaling-Law Auto-Configuration in base_train.py
When you execute scripts/base_train.py, the script performs several automated steps before training begins:
- Meta-device initialization: The model is built on a meta device to calculate parameter counts without allocating full GPU memory
- Weight initialization: Parameters are initialized following GPT-2 schemes adapted for the calculated dimensions
- Training horizon calculation: The script derives the optimal number of iterations from
--target-param-data-ratio(default), or accepts explicit overrides via--num-iterationsor--target-flops - Gradient accumulation setup: The training loop automatically accumulates gradients to reach the compute-optimal total batch size derived from Chinchilla scaling laws
Optimizer Setup and Learning Rate Scaling
The GPT.setup_optimizer method in nanochat/gpt.py implements a sophisticated parameter grouping strategy using the Muon + AdamW optimizer combination defined in nanochat/optim.py. Learning rates automatically scale by the square root of the model dimension ratio (√(dim/768)), and the script applies batch-size scaling factors to maintain training stability across different hardware configurations.
You can override specific learning rates using flags like --embedding-lr or --matrix-lr, though the defaults are compute-optimal for the detected architecture.
Executing the Training Run
For a standard single-GPU training run, use the Python module execution syntax. The --run flag configures Weights & Biases logging (use --run=dummy to disable).
# Train a 12-layer model on a single GPU
python -m scripts.base_train \
--depth=12 \
--run="single_gpu_d12" \
--device-batch-size=8
If you are on a multi-GPU node but wish to restrict training to a single device, use torchrun with explicit process limiting:
torchrun --nproc_per_node=1 -m scripts.base_train \
--depth=24 \
--run="single_gpu_d24" \
--device-batch-size=4
Enabling FP8 on H100 GPUs
For H100 or newer Hopper architecture GPUs, add the --fp8 flag to enable FP8 mixed precision training, which significantly increases throughput without impacting convergence.
python -m scripts.base_train --depth=24 --fp8 --device-batch-size=32
Checkpointing and Evaluation
The nanochat/checkpoint_manager.py module handles persistent storage of model weights, optimizer states, and dataloader position. By default, checkpoints save to base_checkpoints/d<depth>/ at the end of training, or at intervals specified by --save-every.
During training, scripts/base_train.py periodically evaluates validation bits-per-byte (BPB), the CORE metric, and generates sample text completions. Control evaluation frequency using:
--eval-every: Validation loss and BPB calculation--core-metric-every: CORE benchmark evaluation--sample-every: Text generation sampling
All metrics are logged automatically to Weights & Biuses unless you specify --run=dummy.
Summary
- Install NanoChat using
pip install -e .and optionally setNANOCHAT_DTYPEfor precision control - Configure your single-GPU run using
--depthfor model size and--device-batch-sizeto fit VRAM constraints - Execute training via
python -m scripts.base_trainortorchrun --nproc_per_node=1for explicit single-GPU targeting - Leverage automatic scaling-law calculations in
scripts/base_train.pyfor batch size, learning rate, and training horizon - Optimize using the Muon + AdamW optimizer in
nanochat/optim.pywith automatic parameter grouping viaGPT.setup_optimizer - Monitor progress through built-in validation BPB, CORE metrics, and sample generation logged to Weights & Biases
- Checkpoint models using
nanochat/checkpoint_manager.pywith configurable save intervals via--save-every
Frequently Asked Questions
What is the minimum GPU memory required to train a LLM with NanoChat?
You can train smaller models (depth 12) on GPUs with 24GB-48GB VRAM by setting --device-batch-size to 4 or 8. The framework automatically handles gradient accumulation to maintain the compute-optimal total batch size, so you only need to ensure the model parameters and optimizer states fit in memory. For reference, the default --device-batch-size=32 assumes an 80GB H100.
How does NanoChat calculate the optimal training duration?
In scripts/base_train.py, the script calculates training iterations using scaling-law formulas based on the --target-param-data-ratio parameter by default, ensuring Chinchilla-optimal training. You can override this with --num-iterations for a fixed step count or --target-flops for a specific compute budget, allowing precise control over training horizon without manual calculation.
Can I use multiple GPUs even if I started with single-GPU training?
Yes, the same code supports distributed training without modification. Simply launch with torchrun --nproc_per_node=N where N matches your GPU count. The GPT.setup_optimizer function in nanochat/gpt.py automatically handles distributed-aware learning rate scaling and parameter sharding, ensuring seamless scaling from single-GPU experiments to multi-node clusters.
Why does NanoChat use both Muon and AdamW optimizers?
According to nanochat/optim.py, the framework applies Muon optimizer to matrix-shaped parameters (projections, embeddings) and AdamW to remaining parameters like layer norms and biases. This mixed strategy exploits the orthogonalization properties of Muon for weight matrices while maintaining the stability of AdamW for smaller parameter groups, as implemented in the GPT.setup_optimizer method.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →