# Which Depth Delivers GPT-2 Capability in NanoChat: The 26-Layer Benchmark

> Discover the 26-layer benchmark for GPT-2 capability in nanochat. Explore the experimental findings detailing the optimal depth for advanced performance in the nanochat repository.

- Repository: [Andrej/nanochat](https://github.com/karpathy/nanochat)
- Tags: deep-dive
- Published: 2026-03-10

---

**GPT-2‑level capability in nanochat is achieved at a model depth of approximately 26 layers, with the [`dev/LOG.md`](https://github.com/karpathy/nanochat/blob/main/dev/LOG.md) experiments identifying an effective range spanning depths 24 through 26.**

NanoChat, Karpathy's streamlined implementation for training GPT‑style language models from scratch, relies on a configurable `depth` parameter to control architectural capacity. To match the performance characteristics of OpenAI's GPT‑2 small (124M parameter) model, practitioners must scale this parameter according to empirical findings documented in the project's experimental logs. Analysis of the training records in [`dev/LOG.md`](https://github.com/karpathy/nanochat/blob/main/dev/LOG.md) reveals that crossing the GPT‑2 capability threshold requires a specific minimum layer count that balances representational power with trainability.

## Understanding Model Depth in NanoChat

In the nanochat codebase, **depth** refers to the number of stacked transformer decoder blocks in the architecture. Each layer consists of a multi‑head self‑attention mechanism followed by a position‑wise feed‑forward network, mirroring the original GPT‑2 design. This parameter directly determines the model's total parameter count and its ability to capture long‑range dependencies in text. Increasing depth adds computational cost but is necessary to reach the capacity required for human‑like text generation and few‑shot learning.

## The GPT-2 Capability Threshold: Depth 26

According to the experimental logs recorded in [`dev/LOG.md`](https://github.com/karpathy/nanochat/blob/main/dev/LOG.md), **GPT‑2 capability corresponds to a depth of 26 layers**. The documentation explicitly notes that GPT‑2‑level performance falls within the **d24‑d26 range**, with depth 26 serving as the definitive benchmark for achieving full parity with the 124M parameter GPT‑2 small model. This conclusion derives from systematic experiments comparing validation loss and generation quality across progressively deeper architectures.

### The d24-d26 Performance Range

While depth 26 represents the target configuration, the **d24‑d26 range** indicates where models begin exhibiting GPT‑2‑like behaviors. Architectures at depth 24 or 25 may demonstrate strong language modeling performance but typically exhibit higher perplexity or reduced coherence compared to the full 26‑layer implementation. Depth 26 consistently delivers the parameter density and layerwise transformations necessary to replicate GPT‑2's emergent capabilities.

## Configuring the 26-Layer Architecture

To replicate GPT‑2 capability, set the `depth` parameter to `26` in your training configuration. This aligns the nanochat architecture with the 124M parameter count of GPT‑2 small.

```python

# model configuration in train.py or config file

model_config = {
    "depth": 26,      # GPT-2 capability threshold

    "dim": 768,       # Embedding dimension

    "n_head": 12,     # Number of attention heads

    "vocab_size": 50257,
}

```

When launching training from the command line, explicitly pass the depth argument to ensure the model scales to the required capacity:

```bash
python train.py --depth=26 --dim=768 --n_head=12 --batch_size=64 --max_iters=100000

```

## Computational Requirements at Depth 26

Training at **depth 26** yields approximately 124 million parameters, requiring significant GPU memory and compute resources. This configuration typically demands multi‑GPU setups or extended training times on high‑memory single GPUs to process batch sizes necessary for stable optimization. The investment, however, delivers the full generative performance and few‑shot task capabilities that define GPT‑2‑level models.

## Summary

- **Depth 26** corresponds to GPT‑2 capability in nanochat, matching the 124M parameter GPT‑2 small architecture.
- The **d24‑d26 range** represents the transition zone where models approach GPT‑2 performance, with depth 26 serving as the definitive benchmark.
- Configuration is implemented via the `depth` parameter in model configs, as documented in [`dev/LOG.md`](https://github.com/karpathy/nanochat/blob/main/dev/LOG.md).
- This architecture requires substantial computational resources but delivers full GPT‑2 parity in text generation quality.

## Frequently Asked Questions

### What is the exact depth needed for GPT-2 capability in nanochat?

**A depth of 26 layers is required to achieve full GPT‑2 capability** in nanochat. While depths 24 and 25 may approximate this performance based on the d24‑d26 range documented in [`dev/LOG.md`](https://github.com/karpathy/nanochat/blob/main/dev/LOG.md), depth 26 represents the definitive threshold where the model's capacity matches the 124M parameter GPT‑2 small model.

### Can I train a capable model with fewer than 26 layers?

Models trained at depth 24 or 25 can achieve strong language modeling results but typically fail to reach true GPT‑2 capability. These configurations often exhibit higher validation perplexity and less coherent long‑form generation compared to the standard 26‑layer architecture required for full parity.

### Where is the depth recommendation documented?

The empirical relationship between depth and GPT‑2 capability is documented in the repository's [`dev/LOG.md`](https://github.com/karpathy/nanochat/blob/main/dev/LOG.md) file, which logs training experiments showing that GPT‑2‑level performance emerges specifically within the depth 24–26 range, with depth 26 as the reference implementation.

### How does increasing depth to 26 affect training resources?

Scaling from depth 24 to depth 26 increases the parameter count by approximately 8% and adds corresponding memory and compute overhead per training step. This additional capacity is essential for capturing the complex patterns that define GPT‑2‑level generation quality, making the resource cost unavoidable for achieving target performance.