Training, Validation, and Test Data Splits for Neural Networks: The 80/10/10 Guide
The karpathy/nn-zero-to-hero repository implements a standard 80/10/10 split for training, validation, and test data, where the training set fits model parameters, the validation set tunes hyperparameters, and the test set provides the final unbiased performance estimate.
Splitting your dataset correctly is fundamental to building neural networks that generalize beyond memorization. In the karpathy/nn-zero-to-hero course, Andrej Karpathy demonstrates exactly how to create training, validation, and test data splits for neural networks using reproducible randomization and precise indexing. This approach prevents data leakage and ensures that your model evaluation reflects true performance on unseen examples.
The Purpose of the Three-Way Split
Neural network development requires three distinct data partitions to balance learning, tuning, and evaluation:
- Training split (80%): Used to update weights and biases via backpropagation. The model directly optimizes on this data.
- Validation/Dev split (10%): Monitored during training to tune hyperparameters like learning rate and regularization strength. This split detects overfitting when training loss decreases but validation loss plateaus or increases.
- Test split (10%): Held in reserve until final evaluation. It provides an unbiased estimate of how the trained model performs on completely unseen data.
Implementing the 80/10/10 Split in nn-zero-to-hero
According to the source code in lectures/makemore/makemore_part2_mlp.ipynb, the dataset of names is first shuffled with a fixed random seed, then sliced at the 80th and 90th percentiles. This guarantees reproducibility across runs.
Shuffling and Dividing the Dataset
The implementation uses random.seed(42) to ensure identical splits every time the code executes.
import random
# Assume `words` is a list of strings (e.g., names from the dataset)
random.seed(42) # Fixed seed for reproducibility
random.shuffle(words)
# Calculate split indices: 80% train, 10% dev, 10% test
n1 = int(0.8 * len(words)) # First 80% → training
n2 = int(0.9 * len(words)) # Next 10% → validation/dev
# Create datasets using the build_dataset helper function
Xtr, Ytr = build_dataset(words[:n1]) # Training tensors
Xdev, Ydev = build_dataset(words[n1:n2]) # Validation tensors
Xte, Yte = build_dataset(words[n2:]) # Test tensors
print('train:', Xtr.shape, Ytr.shape)
print('dev :', Xdev.shape, Ydev.shape)
print('test :', Xte.shape, Yte.shape)
The build_dataset function converts the word lists into input tensors (X) and target tensors (Y) suitable for PyTorch training.
Training Workflow with Validation Monitoring
During model development, you train exclusively on Xtr/Ytr while periodically evaluating loss on Xdev/Ydev without updating gradients.
Training Loop with Validation Checks
This pattern from the MLP notebook demonstrates how to monitor validation loss each epoch to detect overfitting:
for epoch in range(num_epochs):
# --- Training step on training data ---
ix = torch.randint(0, Xtr.shape[0], (batch_size,))
emb = C[Xtr[ix]]
h = torch.tanh(emb.view(-1, hidden_dim) @ W1 + b1)
logits = h @ W2 + b2
loss = torch.nn.functional.cross_entropy(logits, Ytr[ix])
loss.backward()
# Update parameters
for p in parameters:
p.data -= learning_rate * p.grad
# --- Validation step on dev data ---
with torch.no_grad():
emb = C[Xdev]
h = torch.tanh(emb.view(-1, hidden_dim) @ W1 + b1)
logits = h @ W2 + b2
val_loss = torch.nn.functional.cross_entropy(logits, Ydev)
print(f'Epoch {epoch:02d} | train loss {loss.item():.4f} | val loss {val_loss.item():.4f}')
Final Test Set Evaluation
Only after hyperparameter tuning is complete should you evaluate on the test set. This code provides the unbiased final metric:
with torch.no_grad():
emb = C[Xte]
h = torch.tanh(emb.view(-1, hidden_dim) @ W1 + b1)
logits = h @ W2 + b2
test_loss = torch.nn.functional.cross_entropy(logits, Yte)
print('Final test loss:', test_loss.item())
Key Source Files
These files contain the reference implementations for data splits in the repository:
lectures/makemore/makemore_part2_mlp.ipynb: Contains the concrete 80/10/10 split implementation and training loop logic for the multilayer perceptron.lectures/makemore/makemore_part1_bigrams.ipynb: Demonstrates the same splitting pattern applied to the simpler bigram model.README.md: Introduces the concept of train/dev/test splits as a core machine learning practice in the course overview.
Summary
- Training data (80%) fits the model parameters through gradient descent.
- Validation data (10%) tunes hyperparameters and detects overfitting without contaminating the final evaluation.
- Test data (10%) provides the unbiased final performance metric and must only be used once at the end of development.
- The
random.seed(42)shuffle ensures your training, validation, and test data splits for neural networks remain reproducible across experiments.
Frequently Asked Questions
Why is the test set kept separate from the validation set?
The validation set guides hyperparameter tuning and model selection decisions during development. If you used the test set for this purpose, you would overfit the hyperparameters to that specific data, making your final performance estimate overly optimistic and invalid.
What happens if I don't shuffle the data before splitting?
Without shuffling, sequential ordering in the dataset (such as alphabetical name sorting) could concentrate similar examples into one split. This creates distribution mismatch between splits, causing poor generalization and unreliable validation metrics.
Can I use different split ratios like 70/15/15?
While the nn-zero-to-hero repository uses 80/10/10, other ratios like 70/15/15 are valid depending on dataset size. Smaller datasets often require larger validation/test portions to ensure statistically significant metrics, while massive datasets might use 98/1/1 splits.
Why does the code use random.seed(42) specifically?
The value 42 is arbitrary; any fixed integer ensures reproducibility. Setting the seed guarantees that random.shuffle(words) produces the same order every run, allowing you to compare model architectures and hyperparameters fairly without variance from different data splits.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →