What Information is Stored in the params.json File in Llama 2 Checkpoints?

The params.json file in Llama 2 checkpoints stores the complete model architecture configuration—including hidden dimensions, layer counts, attention head specifications, and runtime limits—required to reconstruct the transformer from raw binary weight tensors.

When working with the meta-llama/llama inference codebase, the params.json file serves as the critical bridge between binary checkpoint data and runnable model configuration. This JSON manifest, located alongside the *.pth weight files in the checkpoint directory, contains the architectural hyperparameters that the loading code in llama/generation.py uses to instantiate the ModelArgs dataclass and properly shape the transformer layers.

The Role of params.json in Llama 2 Checkpoint Loading

The params.json file functions as a machine-readable manifest that describes the neural network architecture independent of the actual weight values. According to the source code in llama/generation.py, the loading process explicitly reads this file to configure the model before weights are loaded:

with open(Path(ckpt_dir) / "params.json", "r") as f:
    params = json.loads(f.read())

This dictionary is then unpacked into the ModelArgs dataclass defined in llama/model.py, which determines the dimensions of every tensor in the checkpoint. Without this file, the binary .pth tensors cannot be correctly reshaped and assigned to the appropriate layers.

Complete List of Parameters Stored in params.json

The JSON keys in params.json map directly to fields in the ModelArgs dataclass. Each parameter controls a specific aspect of the transformer architecture:

  • dim: The hidden dimension size used throughout each transformer block (default: 4096).
  • n_layers: The total number of transformer layers in the stack (default: 32).
  • n_heads: The number of attention heads for queries (default: 32).
  • n_kv_heads: The number of key/value heads; defaults to n_heads if omitted, enabling Grouped-Query Attention (GQA) when set lower.
  • vocab_size: The size of the vocabulary; typically overridden at load time by the tokenizer's actual word count (default: -1).
  • multiple_of: Ensures hidden dimensions are multiples of this value, primarily for SwiGLU feed-forward network alignment (default: 256).
  • ffn_dim_multiplier: An optional multiplier applied to the feed-forward hidden dimension calculation.
  • norm_eps: The epsilon value used in RMSNorm layers for numerical stability (default: 1e-5).
  • max_seq_len: The maximum sequence length the model can process during inference (default: 2048).
  • max_batch_size: The maximum number of sequences processed simultaneously during inference (default: 32).

These values collectively define the tensor shapes for all weights, biases, and embeddings stored in the accompanying .pth files.

How params.json Maps to ModelArgs in llama/model.py

The architectural configuration flows directly from the JSON file into Python objects. In llama/model.py, the ModelArgs dataclass serves as the schema:

@dataclass
class ModelArgs:
    dim: int = 4096
    n_layers: int = 32
    n_heads: int = 32
    n_kv_heads: Optional[int] = None
    vocab_size: int = -1
    multiple_of: int = 256
    ffn_dim_multiplier: Optional[float] = None
    norm_eps: float = 1e-5
    max_seq_len: int = 2048
    max_batch_size: int = 32

When the checkpoint loader in llama/generation.py instantiates ModelArgs, it unpacks the JSON dictionary using the **params syntax:

model_args = ModelArgs(
    max_seq_len=max_seq_len,
    max_batch_size=max_batch_size,
    **params,  # Injects dim, n_layers, n_heads, etc. from params.json

)

This design ensures that the model architecture is strictly defined by the checkpoint metadata rather than hardcoded assumptions.

Practical Example: Loading a Llama 2 Checkpoint Using params.json

To demonstrate how params.json functions in practice, the following example shows the complete loading sequence used in the reference implementation:

from pathlib import Path
import json
import torch
from llama.generation import Llama
from llama.model import ModelArgs, Transformer
from llama.tokenizer import Tokenizer

def load_llama_checkpoint(ckpt_dir: str, tokenizer_path: str,
                         max_seq_len: int = 2048,
                         max_batch_size: int = 32):
    # Step 1: Load the architectural parameters from params.json

    with open(Path(ckpt_dir) / "params.json", "r") as f:
        params = json.loads(f.read())
    
    # Step 2: Build ModelArgs by merging JSON values with runtime constraints

    model_args = ModelArgs(
        max_seq_len=max_seq_len,
        max_batch_size=max_batch_size,
        **params,
    )
    
    # Step 3: Override vocab_size using the actual tokenizer

    tokenizer = Tokenizer(model_path=tokenizer_path)
    model_args.vocab_size = tokenizer.n_words
    
    # Step 4: Load binary weight tensors from .pth files

    checkpoints = sorted(Path(ckpt_dir).glob("*.pth"))
    ckpt_path = checkpoints[0]  # Single-GPU loading example

    state_dict = torch.load(ckpt_path, map_location="cpu")
    
    # Step 5: Instantiate the Transformer with architecture from params.json

    torch.set_default_tensor_type(torch.cuda.HalfTensor)
    model = Transformer(model_args)
    model.load_state_dict(state_dict, strict=False)
    
    return Llama(model, tokenizer)

# Usage example

llama = load_llama_checkpoint(
    ckpt_dir="llama-2-7b-chat/",
    tokenizer_path="tokenizer.model"
)

This workflow confirms that params.json is the authoritative source for architectural dimensions, while the .pth files supply only the numerical parameters.

Summary

  • The params.json file in Llama 2 checkpoints stores the complete model architecture configuration, including hidden dimensions, layer counts, and attention head specifications.
  • Located alongside the binary .pth weight files, this JSON manifest is parsed by llama/generation.py to instantiate the ModelArgs dataclass defined in llama/model.py.
  • Key parameters include dim, n_layers, n_heads, n_kv_heads, max_seq_len, and norm_eps, which collectively define the tensor shapes required to load the checkpoint correctly.
  • The separation of architecture metadata (params.json) from weight values (.pth files) enables flexible model loading and ensures reproducibility across different runtime environments.

Frequently Asked Questions

What happens if the params.json file is missing or corrupted?

If params.json is missing from the checkpoint directory, the loading code in llama/generation.py will raise a FileNotFoundError when attempting to open the path. Without this file, the code cannot instantiate ModelArgs with the correct architectural dimensions, making it impossible to reshape the flat weight tensors from the .pth files into the proper transformer layers. You must restore the original params.json from the official Meta Llama 2 distribution to load the checkpoint successfully.

Can I modify params.json to change the model architecture?

Modifying params.json to change values like dim, n_layers, or n_heads will cause a shape mismatch error when load_state_dict() attempts to map the binary weight tensors into the model. The .pth files contain tensors with fixed shapes that correspond to the original architecture defined in the unmodified params.json. If you need to alter the architecture, you must reinitialize the model with new dimensions and train from scratch, or use a checkpoint that matches your desired configuration.

How does params.json support Grouped-Query Attention (GQA)?

The n_kv_heads parameter in params.json enables Grouped-Query Attention by allowing the number of key/value heads to differ from the query heads (n_heads). When n_kv_heads is set lower than n_heads—for example, 8 key/value heads versus 32 query heads in Llama 2 70B—the model shares key and value representations across multiple query heads, reducing memory bandwidth during inference. If n_kv_heads is omitted from params.json, the code defaults to n_heads, implementing standard multi-head attention.

Does params.json contain the model weights?

No, params.json contains only metadata and hyperparameters, not the actual model weights. The numerical parameters—embedding matrices, attention weights, and feed-forward network biases—are stored separately in the binary *.pth (PyTorch) files. The JSON file provides the architectural blueprint that tells the loader how to reshape and assign those raw tensors into the correct layers of the Transformer model defined in llama/model.py. This separation allows the same weight files to be loaded with different runtime constraints by adjusting params.json or the runtime arguments.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →