How to Convert and Consolidate Eagle Model Checkpoints for Deployment
To convert and consolidate Eagle model checkpoints for deployment, load the sharded checkpoint directory using EagleModel.from_pretrained() to automatically merge shards, then export a single consolidated file using torch.save() or save_pretrained() for production inference.
Eagle, developed by NVlabs, generates sharded checkpoints during distributed training that must be consolidated into a single deployable artifact. Understanding how to convert and consolidate Eagle model checkpoints is essential for production inference pipelines. This guide walks through the exact process using the repository's native utilities and HuggingFace compatibility layers.
Understanding Eagle Checkpoint Structure
During training, Eagle's scripts—such as Eagle/scripts/pretrain-eagle-x5-vicuna-7b.sh—save checkpoints in subdirectories named checkpoint-<step> within the --output_dir path. According to the source code, each checkpoint folder contains:
pytorch_model.bin– The model'sstate_dict(potentially sharded if model parallelism is used)optimizer.pt– Optimizer state (not required for inference)trainer_state.json– Training metadata including learning rate schedule and global stepargs.json– The training arguments used to generate the checkpoint
For deployment, you must merge these sharded components into a single consolidated file that can be loaded efficiently without the training infrastructure.
Step-by-Step Conversion Process
Locate the Checkpoint Directory
First, identify the specific checkpoint you want to deploy. Typically, this is the final checkpoint or the one with the best validation metrics.
# Example: Locate the latest checkpoint after pre-training
CHECKPOINT_ROOT=./checkpoints # Same path passed to --output_dir in training scripts
LATEST=$(python -c "import glob, os; \
ckpts = sorted(glob.glob(os.path.join('$CHECKPOINT_ROOT', 'checkpoint-*')), \
key=lambda p: int(p.split('-')[-1])); print(ckpts[-1])")
echo "Using checkpoint: $LATEST"
The training scripts in Eagle/scripts/pretrain-eagle-x5-vicuna-7b.sh demonstrate how the --output_dir parameter dictates where these folders are created.
Load and Merge Sharded Weights
Eagle's model classes inherit from HuggingFace's PreTrainedModel, enabling automatic shard merging via from_pretrained(). If your training used the unsloth gradient-checkpointing patch, apply it before loading.
import os
import torch
from eagle import EagleModel # Entry point defined in Eagle/eagle/__init__.py
from eaglevl.patch.unsloth_checkpoint import patch_unsloth_gradient_checkpointing
# Optional: Apply unsloth patch if checkpoint was trained with unsloth gradient checkpointing
# See Embodied/eaglevl/patch/unsloth_checkpoint.py for implementation details
patch_unsloth_gradient_checkpointing()
# Load model - automatically merges sharded tensors from pytorch_model.bin shards
checkpoint_dir = os.path.abspath(LATEST) # e.g., ./checkpoints/checkpoint-3000
model = EagleModel.from_pretrained(
checkpoint_dir,
torch_dtype=torch.float16,
device_map="auto"
)
model.eval()
print(f"Model loaded - Total parameters: {sum(p.numel() for p in model.parameters()):,}")
The EagleModel.from_pretrained() method (implemented via the HuggingFace integration in Embodied/eaglevl/model/moon_vit/modeling_vit.py) handles the reconstruction of sharded tensors automatically, including those saved with DeepSpeed ZeRO-3 or model-parallel strategies.
Export the Consolidated Checkpoint
Once loaded, export to a single file or HuggingFace-compatible directory structure.
export_dir = "./eagle_deploy"
os.makedirs(export_dir, exist_ok=True)
# Option A: Single torch file for custom inference engines
consolidated_path = os.path.join(export_dir, "eagle_consolidated.pth")
torch.save(model.state_dict(), consolidated_path)
print(f"Saved consolidated weights to: {consolidated_path}")
# Option B: HuggingFace format (recommended for compatibility)
# Creates config.json and model.safetensors/pytorch_model.bin
model.save_pretrained(export_dir)
print(f"Saved HF-compatible checkpoint to: {export_dir}")
Consolidating eliminates the need to manage multiple shard files in production and ensures compatibility with inference engines that expect monolithic weight files.
Optional: Export to ONNX for Optimized Inference
For ultra-fast CPU or GPU inference, convert the consolidated checkpoint to ONNX format.
dummy_input = torch.randn(1, 3, 224, 224).to(next(model.parameters()).device)
onnx_path = os.path.join(export_dir, "eagle_model.onnx")
torch.onnx.export(
model,
dummy_input,
onnx_path,
opset_version=17,
input_names=["pixel_values"],
output_names=["logits"],
dynamic_axes={"pixel_values": {0: "batch"}, "logits": {0: "batch"}}
)
print(f"ONNX model exported to: {onnx_path}")
The resulting ONNX file can be deployed using ONNX Runtime, TensorRT, or other high-performance inference backends.
Summary
- Eagle generates sharded checkpoints during training in
checkpoint-<step>folders containingpytorch_model.bin,optimizer.pt, and metadata files. - Use
EagleModel.from_pretrained()to automatically merge sharded weights, applying thepatch_unsloth_gradient_checkpointing()utility first if the model was trained with unsloth optimizations. - Export via
torch.save()for a single consolidated.pthfile, or usesave_pretrained()for HuggingFace-compatible directory structures. - Consolidated checkpoints reduce deployment complexity and enable direct loading by production inference servers without training dependencies.
Frequently Asked Questions
Why does Eagle save sharded checkpoints instead of a single file?
Eagle leverages HuggingFace Trainer and distributed training strategies (such as DeepSpeed ZeRO-3) that shard model weights across GPUs to reduce memory pressure. Each pytorch_model.bin in the checkpoint directory may represent a portion of the full state dict, requiring reassembly during the conversion process described above.
Do I need the optimizer.pt file when converting checkpoints for deployment?
No. The optimizer.pt file contains optimizer states (momentum buffers, Adam moments) necessary for resuming training but irrelevant for inference. When converting and consolidating Eagle model checkpoints for deployment, you only need pytorch_model.bin (the weights) and config.json (the architecture definition).
Can I convert checkpoints trained with DeepSpeed ZeRO-3 using this method?
Yes. The EagleModel.from_pretrained() method automatically handles DeepSpeed ZeRO-3 sharded checkpoints. When loading the checkpoint directory, the method reconstructs the full state dict from the sharded pytorch_model.bin files, provided you have sufficient CPU or GPU memory to hold the consolidated model.
What is the difference between torch.save and save_pretrained for Eagle models?
torch.save(model.state_dict(), file) creates a single raw PyTorch pickle file containing only weights, suitable for custom inference pipelines. model.save_pretrained(directory) creates a standardized HuggingFace directory with config.json and properly formatted weight files (.bin or .safetensors), enabling seamless reloading via EagleModel.from_pretrained() in downstream applications.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →