Tokens-to-Parameters Ratio in NanoChat’s Compute-Optimal Models

NanoChat’s compute-optimal models use a fixed tokens-to-parameters ratio of 10.5, meaning every trainable parameter is trained on approximately 10.5 tokens.

This ratio is hard-coded as the default value for the --target-param-data-ratio argument in the training scripts and serves as the mathematical foundation for calculating training budgets across the entire model family. According to the source code in karpathy/nanochat, the framework automatically derives the optimal number of training iterations by multiplying this ratio by the total parameter count.

What Is the Tokens-to-Parameters Ratio?

In neural network training, the tokens-to-parameters ratio (often denoted as D/N where D is tokens and N is parameters) determines how much data a model sees relative to its size. NanoChat adopts a compute-optimal scaling law where this ratio is held constant at 10.5 across different model depths.

This approach contrasts with the Chinchilla scaling laws (which recommend ~20 tokens per parameter) and represents a more aggressive training regime where models are trained on fewer tokens per parameter but with optimized architectures. The ratio ensures that as you increase model size by adjusting --depth, the training budget scales proportionally to maintain computational efficiency.

Where the Ratio Is Defined in the Code

CLI Argument Definition

The default ratio is defined in scripts/base_train.py at line 58, where the argument parser establishes the scaling constant:

parser.add_argument("--target-param-data-ratio", type=float, default=10.5,
    help="calculate num_iterations to maintain data:param ratio (Chinchilla=20, -1 = disable)")

A later comment in the same file (around line 260) explicitly documents the design philosophy: "The compute-optimal models satisfy the Tokens:Params ratio of --target-param-data-ratio". This confirms that the 10.5 value is the canonical ratio for NanoChat's compute-optimal training regime.

Scaling Parameters Helper

The actual parameter counting occurs in nanochat/optim.py within the get_scaling_params function. This utility calculates the total number of trainable parameters (N) which is then multiplied by the ratio to determine the token budget (D):

def get_scaling_params(model_meta):
    """
    Return a dict with number of parameters ("N") and tokens ("D") for the model.
    The scaling law used in nanochat is D = ratio * N where ratio is target_param_data_ratio.
    """
    # Compute N: total number of trainable parameters

    N = sum(p.numel() for p in model_meta.parameters() if p.requires_grad)
    # D is approximated as N * args.target_param_data_ratio later in training script

    return {"N": N, "D": None}

How the Ratio Drives Training Budgets

During initialization, the training script calculates the optimal token budget using the formula:

target_tokens = int(args.target_param_data_ratio * num_scaling_params)

Here, num_scaling_params represents the parameter count (N) returned by get_scaling_params. For example, a model with 200 million parameters would be allocated approximately 2.1 billion tokens (200,000,000 × 10.5). This automatic calculation ensures that researchers can switch between model sizes simply by changing the --depth flag while maintaining compute-optimal training conditions.

Practical Code Examples

Inspecting the Default Value

You can verify the default ratio directly from the command line:

python -m scripts.base_train --help | grep target-param-data-ratio

The output confirms the default value:


--target-param-data-ratio  FLOAT  default=10.5  calculate num_iterations to maintain data:param ratio (Chinchilla=20, -1 = disable)

Calculating Token Budget for a Model

The following snippet demonstrates how to compute the compute-optimal token budget for any NanoChat model configuration:

from nanochat.gpt import GPT, GPTConfig
from nanochat.optim import get_scaling_params

# Configure a depth-12 model (d12)

depth = 12
aspect_ratio = 64
head_dim = 128
base_dim = depth * aspect_ratio
model_dim = ((base_dim + head_dim - 1) // head_dim) * head_dim
num_heads = model_dim // head_dim

config = GPTConfig(
    sequence_len=2048,
    vocab_size=262,  # 256 + 6 for the tokenizer

    n_layer=depth,
    n_head=num_heads,
    n_kv_head=num_heads,
    n_embd=model_dim,
    window_pattern="L"
)

# Create meta-model and count parameters

model_meta = GPT(config)
stats = get_scaling_params(model_meta)
num_params = stats["N"]

# Apply NanoChat's compute-optimal ratio

tokens_per_param = 10.5
optimal_tokens = int(num_params * tokens_per_param)

print(f"Model parameters: {num_params:,}")
print(f"Compute-optimal tokens: {optimal_tokens:,}")
print(f"Ratio: {tokens_per_param} tokens per parameter")

Overriding the Ratio for Custom Training

While 10.5 is the default for compute-optimal training, you can adjust this value to match other scaling laws (such as Chinchilla's 20 tokens per parameter):


# Train with Chinchilla-style ratio

python -m scripts.base_train --depth 16 --target-param-data-ratio 20

# Disable automatic budget calculation (train for fixed steps)

python -m scripts.base_train --depth 12 --target-param-data-ratio -1

Summary

  • NanoChat's compute-optimal ratio is 10.5 tokens per parameter, defined as the default value for --target-param-data-ratio in scripts/base_train.py.
  • This ratio is used to automatically calculate training budgets via target_tokens = ratio × num_parameters, implemented in the training initialization logic.
  • The get_scaling_params function in nanochat/optim.py handles parameter counting (N), while the training script applies the ratio to determine the token budget (D).
  • Users can override the default ratio via CLI arguments to align with alternative scaling laws like Chinchilla (20:1) or disable automatic calculation entirely.

Frequently Asked Questions

How does NanoChat's 10.5 ratio compare to Chinchilla scaling laws?

NanoChat's default ratio of 10.5 tokens per parameter is approximately half of the Chinchilla recommendation of 20 tokens per parameter. This reflects a different compute-optimal regime where models are trained more aggressively on less data per parameter. The training scripts explicitly document this comparison in the help text for --target-param-data-ratio, allowing users to switch to Chinchilla-style scaling by setting the value to 20.

Can I train a NanoChat model without using the tokens-to-parameters ratio?

Yes. Setting --target-param-data-ratio to -1 disables the automatic training budget calculation, allowing you to specify a fixed number of training iterations manually. This is useful for experimental setups where you want to test different training durations independent of the compute-optimal scaling laws.

Where is the actual token count calculated during training?

The calculation occurs in scripts/base_train.py where the training script multiplies the parameter count (obtained from get_scaling_params in nanochat/optim.py) by the ratio value. The resulting target_tokens integer determines how many tokens the model will process during training, automatically adjusting the number of iterations based on batch size and sequence length.

Does the ratio apply to all model depths in NanoChat?

Yes. The compute-optimal ratio is designed to scale with model size. When you change the --depth flag (e.g., from 8 to 24 layers), the get_scaling_params function recalculates the total parameter count, and the training script automatically adjusts the token budget to maintain the constant 10.5 ratio. This ensures that every model in the family follows the same compute-optimal scaling principle.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →