How to Convert and Consolidate Eagle Model Checkpoints for Deployment

To convert and consolidate Eagle model checkpoints for deployment, load the sharded checkpoint directory using EagleModel.from_pretrained() to automatically merge shards, then export a single consolidated file using torch.save() or save_pretrained() for production inference.

Eagle, developed by NVlabs, generates sharded checkpoints during distributed training that must be consolidated into a single deployable artifact. Understanding how to convert and consolidate Eagle model checkpoints is essential for production inference pipelines. This guide walks through the exact process using the repository's native utilities and HuggingFace compatibility layers.

Understanding Eagle Checkpoint Structure

During training, Eagle's scripts—such as Eagle/scripts/pretrain-eagle-x5-vicuna-7b.sh—save checkpoints in subdirectories named checkpoint-<step> within the --output_dir path. According to the source code, each checkpoint folder contains:

  • pytorch_model.bin – The model's state_dict (potentially sharded if model parallelism is used)
  • optimizer.pt – Optimizer state (not required for inference)
  • trainer_state.json – Training metadata including learning rate schedule and global step
  • args.json – The training arguments used to generate the checkpoint

For deployment, you must merge these sharded components into a single consolidated file that can be loaded efficiently without the training infrastructure.

Step-by-Step Conversion Process

Locate the Checkpoint Directory

First, identify the specific checkpoint you want to deploy. Typically, this is the final checkpoint or the one with the best validation metrics.


# Example: Locate the latest checkpoint after pre-training

CHECKPOINT_ROOT=./checkpoints  # Same path passed to --output_dir in training scripts

LATEST=$(python -c "import glob, os; \
    ckpts = sorted(glob.glob(os.path.join('$CHECKPOINT_ROOT', 'checkpoint-*')), \
    key=lambda p: int(p.split('-')[-1])); print(ckpts[-1])")
echo "Using checkpoint: $LATEST"

The training scripts in Eagle/scripts/pretrain-eagle-x5-vicuna-7b.sh demonstrate how the --output_dir parameter dictates where these folders are created.

Load and Merge Sharded Weights

Eagle's model classes inherit from HuggingFace's PreTrainedModel, enabling automatic shard merging via from_pretrained(). If your training used the unsloth gradient-checkpointing patch, apply it before loading.

import os
import torch
from eagle import EagleModel  # Entry point defined in Eagle/eagle/__init__.py

from eaglevl.patch.unsloth_checkpoint import patch_unsloth_gradient_checkpointing

# Optional: Apply unsloth patch if checkpoint was trained with unsloth gradient checkpointing

# See Embodied/eaglevl/patch/unsloth_checkpoint.py for implementation details

patch_unsloth_gradient_checkpointing()

# Load model - automatically merges sharded tensors from pytorch_model.bin shards

checkpoint_dir = os.path.abspath(LATEST)  # e.g., ./checkpoints/checkpoint-3000

model = EagleModel.from_pretrained(
    checkpoint_dir,
    torch_dtype=torch.float16,
    device_map="auto"
)

model.eval()
print(f"Model loaded - Total parameters: {sum(p.numel() for p in model.parameters()):,}")

The EagleModel.from_pretrained() method (implemented via the HuggingFace integration in Embodied/eaglevl/model/moon_vit/modeling_vit.py) handles the reconstruction of sharded tensors automatically, including those saved with DeepSpeed ZeRO-3 or model-parallel strategies.

Export the Consolidated Checkpoint

Once loaded, export to a single file or HuggingFace-compatible directory structure.

export_dir = "./eagle_deploy"
os.makedirs(export_dir, exist_ok=True)

# Option A: Single torch file for custom inference engines

consolidated_path = os.path.join(export_dir, "eagle_consolidated.pth")
torch.save(model.state_dict(), consolidated_path)
print(f"Saved consolidated weights to: {consolidated_path}")

# Option B: HuggingFace format (recommended for compatibility)

# Creates config.json and model.safetensors/pytorch_model.bin

model.save_pretrained(export_dir)
print(f"Saved HF-compatible checkpoint to: {export_dir}")

Consolidating eliminates the need to manage multiple shard files in production and ensures compatibility with inference engines that expect monolithic weight files.

Optional: Export to ONNX for Optimized Inference

For ultra-fast CPU or GPU inference, convert the consolidated checkpoint to ONNX format.

dummy_input = torch.randn(1, 3, 224, 224).to(next(model.parameters()).device)
onnx_path = os.path.join(export_dir, "eagle_model.onnx")

torch.onnx.export(
    model,
    dummy_input,
    onnx_path,
    opset_version=17,
    input_names=["pixel_values"],
    output_names=["logits"],
    dynamic_axes={"pixel_values": {0: "batch"}, "logits": {0: "batch"}}
)
print(f"ONNX model exported to: {onnx_path}")

The resulting ONNX file can be deployed using ONNX Runtime, TensorRT, or other high-performance inference backends.

Summary

  • Eagle generates sharded checkpoints during training in checkpoint-<step> folders containing pytorch_model.bin, optimizer.pt, and metadata files.
  • Use EagleModel.from_pretrained() to automatically merge sharded weights, applying the patch_unsloth_gradient_checkpointing() utility first if the model was trained with unsloth optimizations.
  • Export via torch.save() for a single consolidated .pth file, or use save_pretrained() for HuggingFace-compatible directory structures.
  • Consolidated checkpoints reduce deployment complexity and enable direct loading by production inference servers without training dependencies.

Frequently Asked Questions

Why does Eagle save sharded checkpoints instead of a single file?

Eagle leverages HuggingFace Trainer and distributed training strategies (such as DeepSpeed ZeRO-3) that shard model weights across GPUs to reduce memory pressure. Each pytorch_model.bin in the checkpoint directory may represent a portion of the full state dict, requiring reassembly during the conversion process described above.

Do I need the optimizer.pt file when converting checkpoints for deployment?

No. The optimizer.pt file contains optimizer states (momentum buffers, Adam moments) necessary for resuming training but irrelevant for inference. When converting and consolidating Eagle model checkpoints for deployment, you only need pytorch_model.bin (the weights) and config.json (the architecture definition).

Can I convert checkpoints trained with DeepSpeed ZeRO-3 using this method?

Yes. The EagleModel.from_pretrained() method automatically handles DeepSpeed ZeRO-3 sharded checkpoints. When loading the checkpoint directory, the method reconstructs the full state dict from the sharded pytorch_model.bin files, provided you have sufficient CPU or GPU memory to hold the consolidated model.

What is the difference between torch.save and save_pretrained for Eagle models?

torch.save(model.state_dict(), file) creates a single raw PyTorch pickle file containing only weights, suitable for custom inference pipelines. model.save_pretrained(directory) creates a standardized HuggingFace directory with config.json and properly formatted weight files (.bin or .safetensors), enabling seamless reloading via EagleModel.from_pretrained() in downstream applications.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →