# How to Convert a GPT Architecture to Llama 3.2: RoPE, Scaling, and KV-Cache Implementation

> Convert GPT to Llama 3.2 with our LLMs-from-scratch repository. Swap embeddings, adjust scaling, and update KV-cache, preserving all pretrained weights for seamless integration.

- Repository: [Sebastian Raschka/LLMs-from-scratch](https://github.com/rasbt/LLMs-from-scratch)
- Tags: deep-dive
- Published: 2026-05-12

---

**The LLMs-from-scratch repository provides a plug-and-play conversion that transforms a GPT-style Transformer into Llama 3.2 by swapping absolute positional embeddings for Rotary Positional Embeddings (RoPE), adjusting the attention scaling factor to account for head dimensions, and replacing the KV-cache implementation while preserving all pretrained weights.**

The `rasbt/LLMs-from-scratch` repository demonstrates a systematic approach to converting a classic GPT architecture into a modern Llama 3.2 model. Since both architectures share identical underlying components—including multi-head attention, feed-forward layers, and residual connections—this conversion requires only surgical modifications to the positional encoding, attention scaling mathematics, and cache handling layers to achieve full compatibility.

## Core Architectural Differences Between GPT and Llama 3.2

### Rotary Positional Embeddings (RoPE)

GPT models rely on **absolute sinusoidal positional embeddings** added to input token embeddings at the base layer. Llama 3.2 replaces this with **Rotary Positional Embeddings (RoPE)**, which apply rotation matrices directly to the query and key vectors within each attention head. In [`llama3.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/llama3.py), the `apply_rotary_emb` helper function computes these rotations, providing rotation-invariant positional bias and improved extrapolation to longer sequences than the original absolute encoding scheme.

### Scaled Dot-Product Attention with Head Dimension Scaling

While GPT scales attention scores by `1/√d_k` (where `d_k` is the key dimension), Llama 3.2 introduces an additional scaling factor that accounts for the attention head dimension (`head_dim`). The `scaled_dot_product_attention` function in [`llama3.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/llama3.py) implements this modified scaling arithmetic, reproducing the exact attention weight calculations specified in the Llama architecture and ensuring numerical stability across different sequence lengths.

### Batched KV-Cache Layout for Autoregressive Generation

Both architectures utilize key-value caching for efficient text generation, but Llama 3.2 employs a **batched KV-cache layout** that differs from GPT's implementation. The repository maintains two distinct cache classes: `GPT2KVCache` in [`kv_cache/gpt2.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/kv_cache/gpt2.py) for the original GPT-style cache, and `Llama3KVCache` in [`kv_cache/llama3.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/kv_cache/llama3.py) for the Llama-compatible batched structure. During conversion, the model swaps these classes to ensure optimal memory layout and retrieval speed during token generation.

## Implementing the GPT to Llama 3.2 Conversion

The conversion process centers on the `convert_to_llama3()` function provided in [`llama3.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/llama3.py). This utility performs a weight-preserving transformation that reconfigures the model's forward pass logic without destroying learned representations.

Load a pretrained GPT model and convert it using the repository's conversion utility:

```python
from llms_from_scratch import gpt2
from llms_from_scratch.llama3 import convert_to_llama3

# Load pre-trained GPT-2 weights

gpt = gpt2.GPT2Model.from_pretrained("gpt2_small")

# Convert architecture to Llama 3.2

llama = convert_to_llama3(gpt)

```

Both models share an identical generation API, allowing immediate inference with the converted architecture:

```python

# Generate text using the converted Llama 3.2 model

prompt = "In a distant future, AI"
tokens = llama.tokenizer.encode(prompt)
output_ids = llama.generate(tokens, max_new_tokens=30)

print(llama.tokenizer.decode(output_ids))

```

Verify that the conversion correctly swapped the positional embedding and KV-cache implementations:

```python

# Inspect architectural component changes

print("GPT positional layer:", type(gpt.positional_embedding))
print("Llama positional layer:", type(llama.positional_embedding))

# <class 'llms_from_scratch.gpt2.AbsolutePosEmbedding'>

# <class 'llms_from_scratch.llama3.RotaryPosEmbedding'>

print("GPT KV cache:", type(gpt.kv_cache))
print("Llama KV cache:", type(llama.kv_cache))

# <class 'llms_from_scratch.kv_cache.gpt2.GPT2KVCache'>

# <class 'llms_from_scratch.kv_cache.llama3.Llama3KVCache'>

```

## Source Code Organization and Testing

The repository separates concerns across modular files to facilitate testing and verification:

- **[`llama3.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/llama3.py)** – Contains the core conversion logic, `convert_to_llama3()` function, RoPE implementation via `apply_rotary_emb()`, and the modified `scaled_dot_product_attention()` with head-specific scaling.
- **[`kv_cache/llama3.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/kv_cache/llama3.py)** – Implements the `Llama3KVCache` class with batched tensor layouts optimized for Llama's generation pattern.
- **[`kv_cache/gpt2.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/kv_cache/gpt2.py)** – Houses the reference `GPT2KVCache` implementation used by original GPT models.
- **[`tests/test_llama3.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/tests/test_llama3.py)** – Unit tests validating RoPE calculations, attention scaling, and cache behavior against expected Llama 3.2 outputs.
- **[`tests/test_gpt_to_llama.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/tests/test_gpt_to_llama.py)** – Integration tests ensuring that post-conversion models maintain weight consistency and produce valid token sequences.

## Summary

Converting a GPT architecture to Llama 3.2 requires precise modifications to three specific components while preserving the core Transformer weights:

- **Rotary Positional Embeddings** replace absolute sinusoidal encodings to enable better length extrapolation and relative position encoding through rotation matrices applied in [`llama3.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/llama3.py).
- **Modified Attention Scaling** adjusts the dot-product scaling factor to incorporate the head dimension (`head_dim`), altering how attention scores are computed before the softmax operation.
- **KV-Cache Restructuring** switches from the GPT cache layout to a batched KV-cache implementation (`Llama3KVCache`) that matches Llama 3.2's memory optimization patterns.

The `convert_to_llama3()` function executes these swaps automatically, enabling immediate inference with Llama-style behavior without retraining.

## Frequently Asked Questions

### What is the main difference between GPT and Llama 3.2 positional embeddings?

GPT models use absolute positional embeddings added to token embeddings at the input layer, while Llama 3.2 utilizes Rotary Positional Embeddings (RoPE) that are applied directly to query and key vectors within attention heads. This change, implemented via `apply_rotary_emb` in [`llama3.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/llama3.py), allows Llama to better generalize to sequence lengths beyond training data by encoding relative positions through rotation matrices.

### How does the attention scaling differ between GPT and Llama 3.2?

Standard GPT architectures scale attention scores by the inverse square root of the key dimension (`1/√d_k`). Llama 3.2 introduces an additional scaling term that accounts for the head dimension (`head_dim`), implemented in the `scaled_dot_product_attention` function to match the specific numerical properties of the Llama architecture.

### Why does the KV-cache implementation need to change during conversion?

Llama 3.2 uses a batched KV-cache layout that optimizes memory access patterns for its specific attention mechanism and generation requirements. The repository provides distinct cache classes—`GPT2KVCache` in [`kv_cache/gpt2.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/kv_cache/gpt2.py) and `Llama3KVCache` in [`kv_cache/llama3.py`](https://github.com/rasbt/LLMs-from-scratch/blob/main/kv_cache/llama3.py)—to ensure efficient autoregressive generation compatible with each architecture's expectations.

### Can I convert any GPT-2 model to Llama 3.2 using this method?

The conversion works for any GPT-2-style model that shares the same base hyperparameters (embedding dimensions, number of layers, attention heads) as the target Llama 3.2 configuration. The `convert_to_llama3()` function preserves all weight tensors where shapes match, making it suitable for pretrained checkpoints that align with the repository's expected tensor layouts.