How to Convert a GPT Architecture to Llama 3.2: RoPE, Scaling, and KV-Cache Implementation

The LLMs-from-scratch repository provides a plug-and-play conversion that transforms a GPT-style Transformer into Llama 3.2 by swapping absolute positional embeddings for Rotary Positional Embeddings (RoPE), adjusting the attention scaling factor to account for head dimensions, and replacing the KV-cache implementation while preserving all pretrained weights.

The rasbt/LLMs-from-scratch repository demonstrates a systematic approach to converting a classic GPT architecture into a modern Llama 3.2 model. Since both architectures share identical underlying components—including multi-head attention, feed-forward layers, and residual connections—this conversion requires only surgical modifications to the positional encoding, attention scaling mathematics, and cache handling layers to achieve full compatibility.

Core Architectural Differences Between GPT and Llama 3.2

Rotary Positional Embeddings (RoPE)

GPT models rely on absolute sinusoidal positional embeddings added to input token embeddings at the base layer. Llama 3.2 replaces this with Rotary Positional Embeddings (RoPE), which apply rotation matrices directly to the query and key vectors within each attention head. In llama3.py, the apply_rotary_emb helper function computes these rotations, providing rotation-invariant positional bias and improved extrapolation to longer sequences than the original absolute encoding scheme.

Scaled Dot-Product Attention with Head Dimension Scaling

While GPT scales attention scores by 1/√d_k (where d_k is the key dimension), Llama 3.2 introduces an additional scaling factor that accounts for the attention head dimension (head_dim). The scaled_dot_product_attention function in llama3.py implements this modified scaling arithmetic, reproducing the exact attention weight calculations specified in the Llama architecture and ensuring numerical stability across different sequence lengths.

Batched KV-Cache Layout for Autoregressive Generation

Both architectures utilize key-value caching for efficient text generation, but Llama 3.2 employs a batched KV-cache layout that differs from GPT's implementation. The repository maintains two distinct cache classes: GPT2KVCache in kv_cache/gpt2.py for the original GPT-style cache, and Llama3KVCache in kv_cache/llama3.py for the Llama-compatible batched structure. During conversion, the model swaps these classes to ensure optimal memory layout and retrieval speed during token generation.

Implementing the GPT to Llama 3.2 Conversion

The conversion process centers on the convert_to_llama3() function provided in llama3.py. This utility performs a weight-preserving transformation that reconfigures the model's forward pass logic without destroying learned representations.

Load a pretrained GPT model and convert it using the repository's conversion utility:

from llms_from_scratch import gpt2
from llms_from_scratch.llama3 import convert_to_llama3

# Load pre-trained GPT-2 weights

gpt = gpt2.GPT2Model.from_pretrained("gpt2_small")

# Convert architecture to Llama 3.2

llama = convert_to_llama3(gpt)

Both models share an identical generation API, allowing immediate inference with the converted architecture:


# Generate text using the converted Llama 3.2 model

prompt = "In a distant future, AI"
tokens = llama.tokenizer.encode(prompt)
output_ids = llama.generate(tokens, max_new_tokens=30)

print(llama.tokenizer.decode(output_ids))

Verify that the conversion correctly swapped the positional embedding and KV-cache implementations:


# Inspect architectural component changes

print("GPT positional layer:", type(gpt.positional_embedding))
print("Llama positional layer:", type(llama.positional_embedding))

# <class 'llms_from_scratch.gpt2.AbsolutePosEmbedding'>

# <class 'llms_from_scratch.llama3.RotaryPosEmbedding'>

print("GPT KV cache:", type(gpt.kv_cache))
print("Llama KV cache:", type(llama.kv_cache))

# <class 'llms_from_scratch.kv_cache.gpt2.GPT2KVCache'>

# <class 'llms_from_scratch.kv_cache.llama3.Llama3KVCache'>

Source Code Organization and Testing

The repository separates concerns across modular files to facilitate testing and verification:

  • llama3.py – Contains the core conversion logic, convert_to_llama3() function, RoPE implementation via apply_rotary_emb(), and the modified scaled_dot_product_attention() with head-specific scaling.
  • kv_cache/llama3.py – Implements the Llama3KVCache class with batched tensor layouts optimized for Llama's generation pattern.
  • kv_cache/gpt2.py – Houses the reference GPT2KVCache implementation used by original GPT models.
  • tests/test_llama3.py – Unit tests validating RoPE calculations, attention scaling, and cache behavior against expected Llama 3.2 outputs.
  • tests/test_gpt_to_llama.py – Integration tests ensuring that post-conversion models maintain weight consistency and produce valid token sequences.

Summary

Converting a GPT architecture to Llama 3.2 requires precise modifications to three specific components while preserving the core Transformer weights:

  • Rotary Positional Embeddings replace absolute sinusoidal encodings to enable better length extrapolation and relative position encoding through rotation matrices applied in llama3.py.
  • Modified Attention Scaling adjusts the dot-product scaling factor to incorporate the head dimension (head_dim), altering how attention scores are computed before the softmax operation.
  • KV-Cache Restructuring switches from the GPT cache layout to a batched KV-cache implementation (Llama3KVCache) that matches Llama 3.2's memory optimization patterns.

The convert_to_llama3() function executes these swaps automatically, enabling immediate inference with Llama-style behavior without retraining.

Frequently Asked Questions

What is the main difference between GPT and Llama 3.2 positional embeddings?

GPT models use absolute positional embeddings added to token embeddings at the input layer, while Llama 3.2 utilizes Rotary Positional Embeddings (RoPE) that are applied directly to query and key vectors within attention heads. This change, implemented via apply_rotary_emb in llama3.py, allows Llama to better generalize to sequence lengths beyond training data by encoding relative positions through rotation matrices.

How does the attention scaling differ between GPT and Llama 3.2?

Standard GPT architectures scale attention scores by the inverse square root of the key dimension (1/√d_k). Llama 3.2 introduces an additional scaling term that accounts for the head dimension (head_dim), implemented in the scaled_dot_product_attention function to match the specific numerical properties of the Llama architecture.

Why does the KV-cache implementation need to change during conversion?

Llama 3.2 uses a batched KV-cache layout that optimizes memory access patterns for its specific attention mechanism and generation requirements. The repository provides distinct cache classes—GPT2KVCache in kv_cache/gpt2.py and Llama3KVCache in kv_cache/llama3.py—to ensure efficient autoregressive generation compatible with each architecture's expectations.

Can I convert any GPT-2 model to Llama 3.2 using this method?

The conversion works for any GPT-2-style model that shares the same base hyperparameters (embedding dimensions, number of layers, attention heads) as the target Llama 3.2 configuration. The convert_to_llama3() function preserves all weight tensors where shapes match, making it suitable for pretrained checkpoints that align with the repository's expected tensor layouts.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →