# Training, Validation, and Test Data Splits for Neural Networks: The 80/10/10 Guide

> Understand training, validation, and test data splits for neural networks. Learn the 80/10/10 method from karpathy nn-zero-to-hero for optimal model performance.

- Repository: [Andrej/nn-zero-to-hero](https://github.com/karpathy/nn-zero-to-hero)
- Tags: deep-dive
- Published: 2026-05-23

---

**The karpathy/nn-zero-to-hero repository implements a standard 80/10/10 split for training, validation, and test data, where the training set fits model parameters, the validation set tunes hyperparameters, and the test set provides the final unbiased performance estimate.**

Splitting your dataset correctly is fundamental to building neural networks that generalize beyond memorization. In the **karpathy/nn-zero-to-hero** course, Andrej Karpathy demonstrates exactly how to create **training, validation, and test data splits for neural networks** using reproducible randomization and precise indexing. This approach prevents data leakage and ensures that your model evaluation reflects true performance on unseen examples.

## The Purpose of the Three-Way Split

Neural network development requires three distinct data partitions to balance learning, tuning, and evaluation:

- **Training split (80%)**: Used to update weights and biases via backpropagation. The model directly optimizes on this data.
- **Validation/Dev split (10%)**: Monitored during training to tune hyperparameters like learning rate and regularization strength. This split detects overfitting when training loss decreases but validation loss plateaus or increases.
- **Test split (10%)**: Held in reserve until final evaluation. It provides an unbiased estimate of how the trained model performs on completely unseen data.

## Implementing the 80/10/10 Split in nn-zero-to-hero

According to the source code in `lectures/makemore/makemore_part2_mlp.ipynb`, the dataset of names is first shuffled with a fixed random seed, then sliced at the 80th and 90th percentiles. This guarantees reproducibility across runs.

### Shuffling and Dividing the Dataset

The implementation uses `random.seed(42)` to ensure identical splits every time the code executes.

```python
import random

# Assume `words` is a list of strings (e.g., names from the dataset)

random.seed(42)  # Fixed seed for reproducibility

random.shuffle(words)

# Calculate split indices: 80% train, 10% dev, 10% test

n1 = int(0.8 * len(words))   # First 80% → training

n2 = int(0.9 * len(words))   # Next 10% → validation/dev

# Create datasets using the build_dataset helper function

Xtr, Ytr = build_dataset(words[:n1])      # Training tensors

Xdev, Ydev = build_dataset(words[n1:n2])  # Validation tensors

Xte, Yte = build_dataset(words[n2:])        # Test tensors

print('train:', Xtr.shape, Ytr.shape)
print('dev  :', Xdev.shape, Ydev.shape)
print('test :', Xte.shape, Yte.shape)

```

The `build_dataset` function converts the word lists into input tensors (`X`) and target tensors (`Y`) suitable for PyTorch training.

## Training Workflow with Validation Monitoring

During model development, you train exclusively on `Xtr`/`Ytr` while periodically evaluating loss on `Xdev`/`Ydev` without updating gradients.

### Training Loop with Validation Checks

This pattern from the MLP notebook demonstrates how to monitor validation loss each epoch to detect overfitting:

```python
for epoch in range(num_epochs):
    # --- Training step on training data ---

    ix = torch.randint(0, Xtr.shape[0], (batch_size,))
    emb = C[Xtr[ix]]
    h = torch.tanh(emb.view(-1, hidden_dim) @ W1 + b1)
    logits = h @ W2 + b2
    loss = torch.nn.functional.cross_entropy(logits, Ytr[ix])
    loss.backward()
    
    # Update parameters

    for p in parameters:
        p.data -= learning_rate * p.grad

    # --- Validation step on dev data ---

    with torch.no_grad():
        emb = C[Xdev]
        h = torch.tanh(emb.view(-1, hidden_dim) @ W1 + b1)
        logits = h @ W2 + b2
        val_loss = torch.nn.functional.cross_entropy(logits, Ydev)
    
    print(f'Epoch {epoch:02d} | train loss {loss.item():.4f} | val loss {val_loss.item():.4f}')

```

### Final Test Set Evaluation

Only after hyperparameter tuning is complete should you evaluate on the test set. This code provides the unbiased final metric:

```python
with torch.no_grad():
    emb = C[Xte]
    h = torch.tanh(emb.view(-1, hidden_dim) @ W1 + b1)
    logits = h @ W2 + b2
    test_loss = torch.nn.functional.cross_entropy(logits, Yte)

print('Final test loss:', test_loss.item())

```

## Key Source Files

These files contain the reference implementations for data splits in the repository:

- **`lectures/makemore/makemore_part2_mlp.ipynb`**: Contains the concrete 80/10/10 split implementation and training loop logic for the multilayer perceptron.
- **`lectures/makemore/makemore_part1_bigrams.ipynb`**: Demonstrates the same splitting pattern applied to the simpler bigram model.
- **[`README.md`](https://github.com/karpathy/nn-zero-to-hero/blob/main/README.md)**: Introduces the concept of train/dev/test splits as a core machine learning practice in the course overview.

## Summary

- **Training data (80%)** fits the model parameters through gradient descent.
- **Validation data (10%)** tunes hyperparameters and detects overfitting without contaminating the final evaluation.
- **Test data (10%)** provides the unbiased final performance metric and must only be used once at the end of development.
- The `random.seed(42)` shuffle ensures your **training, validation, and test data splits for neural networks** remain reproducible across experiments.

## Frequently Asked Questions

### Why is the test set kept separate from the validation set?

The **validation set** guides hyperparameter tuning and model selection decisions during development. If you used the **test set** for this purpose, you would overfit the hyperparameters to that specific data, making your final performance estimate overly optimistic and invalid.

### What happens if I don't shuffle the data before splitting?

Without shuffling, sequential ordering in the dataset (such as alphabetical name sorting) could concentrate similar examples into one split. This creates distribution mismatch between splits, causing poor generalization and unreliable validation metrics.

### Can I use different split ratios like 70/15/15?

While the nn-zero-to-hero repository uses **80/10/10**, other ratios like 70/15/15 are valid depending on dataset size. Smaller datasets often require larger validation/test portions to ensure statistically significant metrics, while massive datasets might use 98/1/1 splits.

### Why does the code use `random.seed(42)` specifically?

The value `42` is arbitrary; any fixed integer ensures reproducibility. Setting the seed guarantees that `random.shuffle(words)` produces the same order every run, allowing you to compare model architectures and hyperparameters fairly without variance from different data splits.