# How to Convert HuggingFace Checkpoints to Megatron Format Using MegaDLMs

> Convert HuggingFace checkpoints to Megatron format with MegaDLMs. Learn how to reshape tensors and permute QKV for seamless integration using our dedicated conversion utility.

- Repository: [Jinjie Ni/megadlms](https://github.com/jinjieni/megadlms)
- Tags: how-to-guide
- Published: 2026-03-04

---

**MegaDLMs provides a dedicated conversion utility in [`tools/weights_conversion/hf_to_megatron_te.py`](https://github.com/jinjieni/megadlms/blob/main/tools/weights_conversion/hf_to_megatron_te.py) that rewrites HuggingFace model weights into the checkpoint format expected by Megatron-LM, handling tensor reshaping, QKV permutation, and architecture-specific mappings.**

The MegaDLMs repository (jinjieni/megadlms) ships with a complete toolchain to convert HuggingFace checkpoints to Megatron format. This conversion is essential when migrating pretrained models from the HuggingFace ecosystem into Megatron-LM for continued training or large-scale inference.

## Understanding the Conversion Pipeline

The conversion process follows a five-stage pipeline implemented in [`tools/weights_conversion/hf_to_megatron_te.py`](https://github.com/jinjieni/megadlms/blob/main/tools/weights_conversion/hf_to_megatron_te.py). First, the script loads the original HuggingFace checkpoint using `AutoModelForCausalLM.from_pretrained`. Next, it dispatches to a model-specific conversion routine (such as `llama_to_megatron` or `mistral_to_megatron`) that re-maps tensor names and reshapes weight matrices.

The routine then applies QKV permutation using `permute_qkv` from [`utils/permute_qkv.py`](https://github.com/jinjieni/megadlms/blob/main/utils/permute_qkv.py) to match Megatron's attention layout. After reshaping the weights, the script constructs a Megatron-compatible argument namespace containing architectural parameters. Finally, it creates the Megatron checkpoint directory structure and writes the converted tensors to disk.

## Supported Model Architectures

MegaDLMs supports converting several popular transformer architectures from HuggingFace to Megatron format:

- **LLaMA / LLaMA-2 / CodeLlama** – Use `--model llama`, `llama2`, or `codellama`
- **Mistral** – Use `--model mistral`
- **Qwen / Qwen-2.5-Math** – Use `--model qwen` or `qwen_2_5_math`
- **Falcon** – Use `--model falcon` (implementation present, pending full activation)
- **Gemma** – Future support planned

Each architecture uses specialized mapping functions to handle unique features such as grouped-query attention (GQA) in Mistral and Qwen, or bias terms in Qwen models.

## Step-by-Step Conversion Process

### 1. Load the HuggingFace Checkpoint

The script begins by invoking `AutoModelForCausalLM.from_pretrained` to fetch the model weights. This returns a `state_dict` containing all raw tensors from the HuggingFace checkpoint. You can specify a local cache directory with `--cache-dir` to avoid re-downloading.

### 2. Select the Model-Specific Converter

Based on the `--model` argument, the script dispatches to the appropriate conversion function. For example, `--model llama2` triggers `llama_to_megatron`, while `--model mistral` triggers `mistral_to_megatron`. These functions handle architecture-specific tensor name remapping and weight concatenation.

### 3. Permute QKV Matrices

Megatron-LM expects query, key, and value weight matrices in a specific layout that differs from HuggingFace. The conversion routines call `permute_qkv` from [`tools/weights_conversion/utils/permute_qkv.py`](https://github.com/jinjieni/megadlms/blob/main/tools/weights_conversion/utils/permute_qkv.py) to rearrange the tensor dimensions. This permutation is critical for ensuring the converted model produces identical outputs to the original.

### 4. Create Megatron-Compatible Arguments

After reshaping the weights, the script constructs a model-argument namespace containing architectural parameters such as `num_layers`, `hidden_size`, `num_attention_heads`, and `num_query_groups`. These values are derived from static configuration tables (`llama_s2layer`, `qwen_s2heads`, etc.) embedded in the conversion script.

### 5. Write the Checkpoint

Finally, the script creates the Megatron checkpoint directory structure (`release/mp_rank_00`), stores the converted tensors under the `"model"` key, and writes a metadata file ([`latest_checkpointed_iteration.txt`](https://github.com/jinjieni/megadlms/blob/main/latest_checkpointed_iteration.txt)). The checkpoint is saved as a `.pt` file using `torch.save`.

## Verification and Correctness Testing

After conversion, use the [`verify_correctness_dlm.py`](https://github.com/jinjieni/megadlms/blob/main/verify_correctness_dlm.py) utility to ensure numerical fidelity between the original HuggingFace model and the converted Megatron checkpoint. This script loads both models, runs forward passes on identical inputs, and compares logits and loss values.

```bash
python tools/weights_conversion/utils/verify_correctness_dlm.py \
    --load /path/to/megatron/checkpoint \
    --cache-dir /path/to/hf/cache \
    --activation-sample-path /tmp/activation_logs \
    --huggingface-device cuda:0

```

The script reports the maximum absolute error between implementations. Values below `1e-5` typically indicate a successful conversion with no numerical drift.

## Practical Code Examples

### Converting LLaMA-2-7B from HuggingFace

```bash
python tools/weights_conversion/hf_to_megatron_te.py llama2 \
    --size 7 \
    --out /tmp/llama2_7b_megatron \
    --cache-dir /tmp/hf_cache \
    --model-path meta-llama/Llama-2-7b-hf

```

This command downloads the LLaMA-2-7B checkpoint (or uses the cached version), converts it using `llama_to_megatron`, and saves the result to `/tmp/llama2_7b_megatron/release/mp_rank_00/model_optim_rng.pt`.

### Converting Qwen-2.5-Math-7B with Grouped-Query Attention

```bash
python tools/weights_conversion/hf_to_megatron_te.py qwen_2_5_math \
    --size 7 \
    --out /tmp/qwen2_5_math_7b_megatron \
    --cache-dir /tmp/hf_cache \
    --model-path Qwen/Qwen2.5-Math-7B

```

The conversion routine automatically configures **grouped-query attention** parameters based on the `qwen_s2kvheads` mapping table, setting `num_query_groups` and enabling bias handling (`add_qkv_bias=True`).

### Converting Mistral-7B

```bash
python tools/weights_conversion/hf_to_megatron_te.py mistral \
    --size 7 \
    --out /tmp/mistral_7b_megatron \
    --cache-dir /tmp/hf_cache \
    --model-path mistralai/Mistral-7B-v0.1

```

Mistral uses the same `permute_qkv` logic as LLaMA but applies architecture-specific parameters for sliding window attention and grouped-query heads.

## Key Files and Their Roles

| File | Role |
|------|------|
| [`tools/weights_conversion/hf_to_megatron_te.py`](https://github.com/jinjieni/megadlms/blob/main/tools/weights_conversion/hf_to_megatron_te.py) | Main conversion entry point; parses arguments, loads HF weights, dispatches to model-specific converters, writes Megatron checkpoint. |
| [`tools/weights_conversion/utils/permute_qkv.py`](https://github.com/jinjieni/megadlms/blob/main/tools/weights_conversion/utils/permute_qkv.py) | Implements low-level QKV permutation required by Megatron's attention layout. |
| [`tools/weights_conversion/utils/verify_correctness_dlm.py`](https://github.com/jinjieni/megadlms/blob/main/tools/weights_conversion/utils/verify_correctness_dlm.py) | Validates numerical correctness by comparing logits between original HF and converted Megatron models. |
| [`tools/weights_conversion/utils/merge_llama.py`](https://github.com/jinjieni/megadlms/blob/main/tools/weights_conversion/utils/merge_llama.py) | Helper for merging Meta-provided LLaMA weights when not using HuggingFace directly. |

## Summary

- **MegaDLMs** provides a complete toolchain to convert HuggingFace checkpoints to Megatron format via [`tools/weights_conversion/hf_to_megatron_te.py`](https://github.com/jinjieni/megadlms/blob/main/tools/weights_conversion/hf_to_megatron_te.py).
- The conversion process handles **tensor reshaping**, **QKV permutation**, and **architecture-specific mappings** for LLaMA, Mistral, Qwen, and other models.
- **Grouped-query attention** and bias terms are automatically configured based on static mapping tables within the conversion script.
- Always verify converted checkpoints using [`verify_correctness_dlm.py`](https://github.com/jinjieni/megadlms/blob/main/verify_correctness_dlm.py) to ensure **numerical fidelity** between HuggingFace and Megatron implementations.

## Frequently Asked Questions

### What model architectures are supported for converting HuggingFace checkpoints to Megatron format?

MegaDLMs supports **LLaMA**, **LLaMA-2**, **CodeLlama**, **Mistral**, **Qwen** (including Qwen-2.5-Math), **Falcon** (implementation present, pending full activation), and **Gemma** (future support). Each architecture uses specialized conversion routines to handle unique attention mechanisms and tensor layouts.

### How does the conversion handle grouped-query attention?

The conversion script automatically configures **grouped-query attention** (GQA) by referencing static mapping tables such as `qwen_s2kvheads` and `mistral_s2heads`. When converting models like Qwen-2.5-Math or Mistral, the script sets the `num_query_groups` parameter and adjusts tensor concatenation to match Megatron's GQA expectations.

### Can I convert HuggingFace checkpoints without internet access?

Yes. The conversion script can operate offline by specifying a local `--cache-dir` containing previously downloaded HuggingFace weights, or by using `--model-path` pointing to a local directory containing [`config.json`](https://github.com/jinjieni/megadlms/blob/main/config.json) and model weights (`.bin` or `.safetensors` files). The `AutoModelForCausalLM.from_pretrained` call will load from the local path without requiring network access.

### How do I verify the conversion was successful?

Use the [`tools/weights_conversion/utils/verify_correctness_dlm.py`](https://github.com/jinjieni/megadlms/blob/main/tools/weights_conversion/utils/verify_correctness_dlm.py) utility to perform numerical validation. This script loads both the original HuggingFace model and the converted Megatron checkpoint, runs identical forward passes, and compares the resulting logits and loss values. Maximum absolute errors below `1e-5` indicate a successful conversion with no numerical drift.