How to Convert HuggingFace Checkpoints to Megatron Format Using MegaDLMs

MegaDLMs provides a dedicated conversion utility in tools/weights_conversion/hf_to_megatron_te.py that rewrites HuggingFace model weights into the checkpoint format expected by Megatron-LM, handling tensor reshaping, QKV permutation, and architecture-specific mappings.

The MegaDLMs repository (jinjieni/megadlms) ships with a complete toolchain to convert HuggingFace checkpoints to Megatron format. This conversion is essential when migrating pretrained models from the HuggingFace ecosystem into Megatron-LM for continued training or large-scale inference.

Understanding the Conversion Pipeline

The conversion process follows a five-stage pipeline implemented in tools/weights_conversion/hf_to_megatron_te.py. First, the script loads the original HuggingFace checkpoint using AutoModelForCausalLM.from_pretrained. Next, it dispatches to a model-specific conversion routine (such as llama_to_megatron or mistral_to_megatron) that re-maps tensor names and reshapes weight matrices.

The routine then applies QKV permutation using permute_qkv from utils/permute_qkv.py to match Megatron's attention layout. After reshaping the weights, the script constructs a Megatron-compatible argument namespace containing architectural parameters. Finally, it creates the Megatron checkpoint directory structure and writes the converted tensors to disk.

Supported Model Architectures

MegaDLMs supports converting several popular transformer architectures from HuggingFace to Megatron format:

  • LLaMA / LLaMA-2 / CodeLlama – Use --model llama, llama2, or codellama
  • Mistral – Use --model mistral
  • Qwen / Qwen-2.5-Math – Use --model qwen or qwen_2_5_math
  • Falcon – Use --model falcon (implementation present, pending full activation)
  • Gemma – Future support planned

Each architecture uses specialized mapping functions to handle unique features such as grouped-query attention (GQA) in Mistral and Qwen, or bias terms in Qwen models.

Step-by-Step Conversion Process

1. Load the HuggingFace Checkpoint

The script begins by invoking AutoModelForCausalLM.from_pretrained to fetch the model weights. This returns a state_dict containing all raw tensors from the HuggingFace checkpoint. You can specify a local cache directory with --cache-dir to avoid re-downloading.

2. Select the Model-Specific Converter

Based on the --model argument, the script dispatches to the appropriate conversion function. For example, --model llama2 triggers llama_to_megatron, while --model mistral triggers mistral_to_megatron. These functions handle architecture-specific tensor name remapping and weight concatenation.

3. Permute QKV Matrices

Megatron-LM expects query, key, and value weight matrices in a specific layout that differs from HuggingFace. The conversion routines call permute_qkv from tools/weights_conversion/utils/permute_qkv.py to rearrange the tensor dimensions. This permutation is critical for ensuring the converted model produces identical outputs to the original.

4. Create Megatron-Compatible Arguments

After reshaping the weights, the script constructs a model-argument namespace containing architectural parameters such as num_layers, hidden_size, num_attention_heads, and num_query_groups. These values are derived from static configuration tables (llama_s2layer, qwen_s2heads, etc.) embedded in the conversion script.

5. Write the Checkpoint

Finally, the script creates the Megatron checkpoint directory structure (release/mp_rank_00), stores the converted tensors under the "model" key, and writes a metadata file (latest_checkpointed_iteration.txt). The checkpoint is saved as a .pt file using torch.save.

Verification and Correctness Testing

After conversion, use the verify_correctness_dlm.py utility to ensure numerical fidelity between the original HuggingFace model and the converted Megatron checkpoint. This script loads both models, runs forward passes on identical inputs, and compares logits and loss values.

python tools/weights_conversion/utils/verify_correctness_dlm.py \
    --load /path/to/megatron/checkpoint \
    --cache-dir /path/to/hf/cache \
    --activation-sample-path /tmp/activation_logs \
    --huggingface-device cuda:0

The script reports the maximum absolute error between implementations. Values below 1e-5 typically indicate a successful conversion with no numerical drift.

Practical Code Examples

Converting LLaMA-2-7B from HuggingFace

python tools/weights_conversion/hf_to_megatron_te.py llama2 \
    --size 7 \
    --out /tmp/llama2_7b_megatron \
    --cache-dir /tmp/hf_cache \
    --model-path meta-llama/Llama-2-7b-hf

This command downloads the LLaMA-2-7B checkpoint (or uses the cached version), converts it using llama_to_megatron, and saves the result to /tmp/llama2_7b_megatron/release/mp_rank_00/model_optim_rng.pt.

Converting Qwen-2.5-Math-7B with Grouped-Query Attention

python tools/weights_conversion/hf_to_megatron_te.py qwen_2_5_math \
    --size 7 \
    --out /tmp/qwen2_5_math_7b_megatron \
    --cache-dir /tmp/hf_cache \
    --model-path Qwen/Qwen2.5-Math-7B

The conversion routine automatically configures grouped-query attention parameters based on the qwen_s2kvheads mapping table, setting num_query_groups and enabling bias handling (add_qkv_bias=True).

Converting Mistral-7B

python tools/weights_conversion/hf_to_megatron_te.py mistral \
    --size 7 \
    --out /tmp/mistral_7b_megatron \
    --cache-dir /tmp/hf_cache \
    --model-path mistralai/Mistral-7B-v0.1

Mistral uses the same permute_qkv logic as LLaMA but applies architecture-specific parameters for sliding window attention and grouped-query heads.

Key Files and Their Roles

File Role
tools/weights_conversion/hf_to_megatron_te.py Main conversion entry point; parses arguments, loads HF weights, dispatches to model-specific converters, writes Megatron checkpoint.
tools/weights_conversion/utils/permute_qkv.py Implements low-level QKV permutation required by Megatron's attention layout.
tools/weights_conversion/utils/verify_correctness_dlm.py Validates numerical correctness by comparing logits between original HF and converted Megatron models.
tools/weights_conversion/utils/merge_llama.py Helper for merging Meta-provided LLaMA weights when not using HuggingFace directly.

Summary

  • MegaDLMs provides a complete toolchain to convert HuggingFace checkpoints to Megatron format via tools/weights_conversion/hf_to_megatron_te.py.
  • The conversion process handles tensor reshaping, QKV permutation, and architecture-specific mappings for LLaMA, Mistral, Qwen, and other models.
  • Grouped-query attention and bias terms are automatically configured based on static mapping tables within the conversion script.
  • Always verify converted checkpoints using verify_correctness_dlm.py to ensure numerical fidelity between HuggingFace and Megatron implementations.

Frequently Asked Questions

What model architectures are supported for converting HuggingFace checkpoints to Megatron format?

MegaDLMs supports LLaMA, LLaMA-2, CodeLlama, Mistral, Qwen (including Qwen-2.5-Math), Falcon (implementation present, pending full activation), and Gemma (future support). Each architecture uses specialized conversion routines to handle unique attention mechanisms and tensor layouts.

How does the conversion handle grouped-query attention?

The conversion script automatically configures grouped-query attention (GQA) by referencing static mapping tables such as qwen_s2kvheads and mistral_s2heads. When converting models like Qwen-2.5-Math or Mistral, the script sets the num_query_groups parameter and adjusts tensor concatenation to match Megatron's GQA expectations.

Can I convert HuggingFace checkpoints without internet access?

Yes. The conversion script can operate offline by specifying a local --cache-dir containing previously downloaded HuggingFace weights, or by using --model-path pointing to a local directory containing config.json and model weights (.bin or .safetensors files). The AutoModelForCausalLM.from_pretrained call will load from the local path without requiring network access.

How do I verify the conversion was successful?

Use the tools/weights_conversion/utils/verify_correctness_dlm.py utility to perform numerical validation. This script loads both the original HuggingFace model and the converted Megatron checkpoint, runs identical forward passes, and compares the resulting logits and loss values. Maximum absolute errors below 1e-5 indicate a successful conversion with no numerical drift.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →