# Converting GPT Checkpoints to Other Model Architectures: A Complete Guide to LLMs-from-scratch

> Convert GPT checkpoints to PyTorch with LLMs-from-scratch. Explore original GPT-2, LLaMA-style, and Qwen-3-style model architectures easily.

- Repository: [Sebastian Raschka/LLMs-from-scratch](https://github.com/rasbt/LLMs-from-scratch)
- Tags: how-to-guide
- Published: 2026-05-12

---

**The LLMs-from-scratch repository provides utilities to convert OpenAI GPT-2 TensorFlow checkpoints into three distinct PyTorch architectures: original GPT-2, LLaMA-style, and Qwen-3-style models.**

Converting GPT checkpoints to other model architectures allows researchers to leverage pre-trained weights while experimenting with modern architectural innovations. The `rasbt/LLMs-from-scratch` repository implements a clean pipeline that downloads TensorFlow checkpoints and maps them into PyTorch models with different attention mechanisms, normalization schemes, and projection layers. This article walks through the complete conversion workflow using the actual source implementations.

## Overview of the Conversion Pipeline

The conversion system supports three target architectures, each with dedicated loader functions:

| Target architecture | Core file | Loader function |
|---------------------|-----------|-----------------|
| **GPT-2 (original)** | [`pkg/llms_from_scratch/ch05.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/pkg/llms_from_scratch/ch05.py) | `load_weights_into_gpt` |
| **LLaMA-style** | [`pkg/llms_from_scratch/llama3.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/pkg/llms_from_scratch/llama3.py) | `load_weights_into_llama` |
| **Qwen-3 style** | [`pkg/llms_from_scratch/qwen3.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/pkg/llms_from_scratch/qwen3.py) | `load_weights_into_qwen` |

Every conversion follows the same three-stage pattern. First, download the TensorFlow checkpoint files (`checkpoint`, `model.ckpt.*`, [`encoder.json`](https://github.com/rasbt/LLMs-from-scratch/blob/main/encoder.json)). Second, parse the checkpoint into a pure-Python dictionary using `load_gpt2_params_from_tf_ckpt`. Third, map the dictionary onto the target PyTorch model using the appropriate `load_weights_into_*` helper. All loaders share a common `assign` helper that validates tensor shape compatibility before copying values.

## Downloading and Parsing GPT-2 Checkpoints

The entry point for any conversion is `download_and_load_gpt2` in **[`pkg/llms_from_scratch/ch05.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/pkg/llms_from_scratch/ch05.py)**. This function pulls files from the official OpenAI bucket, discovers the latest checkpoint with `tf.train.latest_checkpoint`, and converts TensorFlow variables into NumPy arrays stored in a nested dictionary.

```python
from llms_from_scratch.ch05 import download_and_load_gpt2

settings, raw_params = download_and_load_gpt2(
    model_size="124M",
    models_dir="~/models/gpt2"
)

```

Under the hood, the function enumerates TF variables with `tf.train.list_variables` and recursively inserts them into a Python dict. For example, `"model/h.0/ln_1/g"` becomes `params["blocks"][0]["ln_1"]["g"]`. The implementation spans lines 33-54 in **[`ch05.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch05.py)** (see [source](https://github.com/rasbt/LLMs-from-scratch/blob/main/pkg/llms_from_scratch/ch05.py#L33-L54)).

## Loading Checkpoints into GPT-2 Architecture

To instantiate the original GPT-2 architecture using the `GPT2Model` class, use the `load_weights_into_gpt` function. This is the baseline conversion that maintains the exact original architecture.

```python
from llms_from_scratch.ch05 import GPT2Model, load_weights_into_gpt

gpt = GPT2Model(settings)
load_weights_into_gpt(gpt, raw_params)

```

The `load_weights_into_gpt` function is defined at lines 27-86 in **[`ch05.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch05.py)** ([source](https://github.com/rasbt/LLMs-from-scratch/blob/main/pkg/llms_from_scratch/ch05.py#L27-L86)). It maps the raw parameter dictionary onto the PyTorch model's embedding, transformer blocks, and output layers.

## Converting Checkpoints to LLaMA-Style Models

The repository includes a compact LLaMA-compatible implementation (`Llama3Model`). Converting GPT-2 checkpoints to this architecture requires remapping the attention and feed-forward layers to match LLaMA's conventions.

```python
from llms_from_scratch.llama3 import Llama3Model, load_weights_into_llama

llama_cfg = {
    "vocab_size": settings["n_vocab"],
    "context_length": settings["n_ctx"],
    "emb_dim": settings["n_embd"],
    "n_heads": settings["n_head"],
    "n_layers": settings["n_layer"],
    "hidden_dim": settings["n_embd"] * 4,
    "head_dim": None,
    "qk_norm": False,
    "n_kv_groups": 1,
    "rope_base": 10000.0,
    "dtype": torch.float32,
}
llama = Llama3Model(llama_cfg)

param_cfg = {
    "n_layers": llama_cfg["n_layers"],
    "hidden_dim": llama_cfg["hidden_dim"]
}

load_weights_into_llama(llama, param_cfg, raw_params)

```

The `load_weights_into_llama` function iterates over each layer, copying query/key/value matrices, output projections, layer-normalization weights, and the three feed-forward matrices (`gate_proj`, `up_proj`, `down_proj`). This loader is implemented at lines 67-124 in **[`pkg/llms_from_scratch/llama3.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/pkg/llms_from_scratch/llama3.py)** ([source](https://github.com/rasbt/LLMs-from-scratch/blob/main/pkg/llms_from_scratch/llama3.py#L67-L124)).

## Converting Checkpoints to Qwen-3 Architecture

Qwen-3 introduces **grouped-query attention** and optional RMSNorm scaling. The conversion process follows the same pattern as LLaMA but includes additional steps for Q- and K-norm scales.

```python
from llms_from_scratch.qwen3 import Qwen3Model, load_weights_into_qwen

qwen_cfg = {
    "vocab_size": settings["n_vocab"],
    "context_length": settings["n_ctx"],
    "emb_dim": settings["n_embd"],
    "n_heads": settings["n_head"],
    "n_layers": settings["n_layer"],
    "hidden_dim": settings["n_embd"] * 3,
    "head_dim": 128,
    "qk_norm": True,
    "n_kv_groups": 8,
    "rope_base": 1_000_000.0,
    "dtype": torch.bfloat16,
}
qwen = Qwen3Model(qwen_cfg)

param_cfg = {
    "n_layers": qwen_cfg["n_layers"],
    "hidden_dim": qwen_cfg["hidden_dim"]
}

load_weights_into_qwen(qwen, param_cfg, raw_params)

```

The `load_weights_into_qwen` function handles the specific grouped-query attention mapping where `n_kv_groups` differs from `n_heads`. See the implementation at lines 51-115 in **[`pkg/llms_from_scratch/qwen3.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/pkg/llms_from_scratch/qwen3.py)** ([source](https://github.com/rasbt/LLMs-from-scratch/blob/main/pkg/llms_from_scratch/qwen3.py#L51-L115)).

## The Assign Helper: Validating Tensor Compatibility

All three loaders rely on a shared `assign` helper function defined at lines 1-8 in **[`ch05.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch05.py)** ([source](https://github.com/rasbt/LLMs-from-scratch/blob/main/pkg/llms_from_scratch/ch05.py#L1-L8)). This utility ensures source and destination tensors have identical shapes before copying:

```python
def assign(left, right, tensor_name="unknown"):
    if left.shape != right.shape:
        raise ValueError(f"Shape mismatch in tensor '{tensor_name}'. "
                         f"Left: {left.shape}, Right: {right.shape}")
    with torch.no_grad():
        if isinstance(right, torch.Tensor):
            left.copy_(right)
        else:
            left.copy_(torch.as_tensor(right, dtype=left.dtype, device=left.device))
    return left

```

This validation step prevents silent shape mismatches during weight transfer and is reused across [`llama3.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/llama3.py) and [`qwen3.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/qwen3.py).

## Complete Conversion Workflow Example

Below is a unified script demonstrating the full pipeline for converting GPT-2 checkpoints to both LLaMA and Qwen-3 architectures:

```python
import torch
from llms_from_scratch.ch05 import download_and_load_gpt2
from llms_from_scratch.llama3 import Llama3Model, load_weights_into_llama
from llms_from_scratch.qwen3 import Qwen3Model, load_weights_into_qwen

# 1. Download raw checkpoint

settings, raw_params = download_and_load_gpt2(
    model_size="124M",
    models_dir="./models"
)

# 2. Convert to LLaMA-style

llama_cfg = {
    "vocab_size": settings["n_vocab"],
    "context_length": settings["n_ctx"],
    "emb_dim": settings["n_embd"],
    "n_heads": settings["n_head"],
    "n_layers": settings["n_layer"],
    "hidden_dim": settings["n_embd"] * 4,
    "head_dim": None,
    "qk_norm": False,
    "n_kv_groups": 1,
    "rope_base": 10_000.0,
    "dtype": torch.float32,
}
llama = Llama3Model(llama_cfg)
load_weights_into_llama(
    llama,
    {"n_layers": llama_cfg["n_layers"], "hidden_dim": llama_cfg["hidden_dim"]},
    raw_params
)

# 3. Convert to Qwen-3-style

qwen_cfg = {
    "vocab_size": settings["n_vocab"],
    "context_length": settings["n_ctx"],
    "emb_dim": settings["n_embd"],
    "n_heads": settings["n_head"],
    "n_layers": settings["n_layer"],
    "hidden_dim": settings["n_embd"] * 3,
    "head_dim": 128,
    "qk_norm": True,
    "n_kv_groups": 8,
    "rope_base": 1_000_000.0,
    "dtype": torch.bfloat16,
}
qwen = Qwen3Model(qwen_cfg)
load_weights_into_qwen(
    qwen,
    {"n_layers": qwen_cfg["n_layers"], "hidden_dim": qwen_cfg["hidden_dim"]},
    raw_params
)

# 4. Verify forward pass

dummy_input = torch.randint(0, llama_cfg["vocab_size"], (2, 16), dtype=torch.long)
logits = llama(dummy_input)
print("Output shape:", logits.shape)

```

## Summary

- The **LLMs-from-scratch** repository provides a unified pipeline for converting OpenAI GPT-2 TensorFlow checkpoints into modern PyTorch architectures.
- Three target architectures are supported: original **GPT-2** (`load_weights_into_gpt`), **LLaMA-style** (`load_weights_into_llama`), and **Qwen-3** (`load_weights_into_qwen`).
- The conversion process extracts TensorFlow checkpoints into Python dictionaries via `download_and_load_gpt2` in [`ch05.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/ch05.py), then maps parameters using architecture-specific loaders.
- All loaders use the `assign` helper to validate tensor shape compatibility before copying weights.
- Configuration dictionaries must specify architectural hyperparameters including `n_layers`, `hidden_dim`, `n_kv_groups`, and `qk_norm` to match the target model structure.

## Frequently Asked Questions

### Can I convert GPT-2 checkpoints to LLaMA architecture using this library?

Yes. The `load_weights_into_llama` function in [`pkg/llms_from_scratch/llama3.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/pkg/llms_from_scratch/llama3.py) specifically handles this conversion. You must provide a configuration dictionary mapping GPT-2 hyperparameters (vocab size, embedding dimension, number of heads) to LLaMA's expected structure, including the feed-forward hidden dimension (typically 4x the embedding dimension) and RoPE base frequency.

### What is the `assign` helper function used for in checkpoint conversion?

The `assign` function validates that source and destination tensors have identical shapes before copying data. Defined in [`pkg/llms_from_scratch/ch05.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/pkg/llms_from_scratch/ch05.py), it raises a `ValueError` if shapes mismatch and handles type conversion between NumPy arrays and PyTorch tensors. This prevents silent errors when mapping between different tensor layouts across architectures.

### How does the Qwen-3 conversion differ from LLaMA conversion?

Qwen-3 conversion supports **grouped-query attention** via the `n_kv_groups` parameter and optional **QK normalization** through the `qk_norm` flag. The `load_weights_into_qwen` function in [`qwen3.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/qwen3.py) handles these additional scaling factors and uses different default values for RoPE base frequency (1,000,000 vs. 10,000) and hidden dimension ratios (3x vs. 4x the embedding size).

### Where are the source files located for these conversion utilities?

The core conversion logic resides in three main files: [`pkg/llms_from_scratch/ch05.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/pkg/llms_from_scratch/ch05.py) contains the GPT-2 loader and checkpoint download utilities, [`pkg/llms_from_scratch/llama3.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/pkg/llms_from_scratch/llama3.py) implements the LLaMA conversion, and [`pkg/llms_from_scratch/qwen3.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/pkg/llms_from_scratch/qwen3.py) handles Qwen-3 mapping. End-to-end tests verifying conversion accuracy against HuggingFace implementations are located in `pkg/llms_from_scratch/tests/`.