# How to Convert and Consolidate Eagle Model Checkpoints for Deployment

> Easily convert and consolidate Eagle model checkpoints for deployment. Load sharded checkpoints, merge them automatically, and export a single file for production inference.

- Repository: [NVIDIA Research Projects/Eagle](https://github.com/NVlabs/Eagle)
- Tags: how-to-guide
- Published: 2026-06-28

---

**To convert and consolidate Eagle model checkpoints for deployment, load the sharded checkpoint directory using `EagleModel.from_pretrained()` to automatically merge shards, then export a single consolidated file using `torch.save()` or `save_pretrained()` for production inference.**

Eagle, developed by NVlabs, generates sharded checkpoints during distributed training that must be consolidated into a single deployable artifact. Understanding how to convert and consolidate Eagle model checkpoints is essential for production inference pipelines. This guide walks through the exact process using the repository's native utilities and HuggingFace compatibility layers.

## Understanding Eagle Checkpoint Structure

During training, Eagle's scripts—such as [`Eagle/scripts/pretrain-eagle-x5-vicuna-7b.sh`](https://github.com/NVlabs/Eagle/blob/main/Eagle/scripts/pretrain-eagle-x5-vicuna-7b.sh)—save checkpoints in subdirectories named `checkpoint-<step>` within the `--output_dir` path. According to the source code, each checkpoint folder contains:

- `pytorch_model.bin` – The model's `state_dict` (potentially sharded if model parallelism is used)
- `optimizer.pt` – Optimizer state (not required for inference)
- [`trainer_state.json`](https://github.com/NVlabs/Eagle/blob/main/trainer_state.json) – Training metadata including learning rate schedule and global step
- [`args.json`](https://github.com/NVlabs/Eagle/blob/main/args.json) – The training arguments used to generate the checkpoint

For deployment, you must merge these sharded components into a single consolidated file that can be loaded efficiently without the training infrastructure.

## Step-by-Step Conversion Process

### Locate the Checkpoint Directory

First, identify the specific checkpoint you want to deploy. Typically, this is the final checkpoint or the one with the best validation metrics.

```bash

# Example: Locate the latest checkpoint after pre-training

CHECKPOINT_ROOT=./checkpoints  # Same path passed to --output_dir in training scripts

LATEST=$(python -c "import glob, os; \
    ckpts = sorted(glob.glob(os.path.join('$CHECKPOINT_ROOT', 'checkpoint-*')), \
    key=lambda p: int(p.split('-')[-1])); print(ckpts[-1])")
echo "Using checkpoint: $LATEST"

```

The training scripts in [`Eagle/scripts/pretrain-eagle-x5-vicuna-7b.sh`](https://github.com/NVlabs/Eagle/blob/main/Eagle/scripts/pretrain-eagle-x5-vicuna-7b.sh) demonstrate how the `--output_dir` parameter dictates where these folders are created.

### Load and Merge Sharded Weights

Eagle's model classes inherit from HuggingFace's `PreTrainedModel`, enabling automatic shard merging via `from_pretrained()`. If your training used the unsloth gradient-checkpointing patch, apply it before loading.

```python
import os
import torch
from eagle import EagleModel  # Entry point defined in Eagle/eagle/__init__.py

from eaglevl.patch.unsloth_checkpoint import patch_unsloth_gradient_checkpointing

# Optional: Apply unsloth patch if checkpoint was trained with unsloth gradient checkpointing

# See Embodied/eaglevl/patch/unsloth_checkpoint.py for implementation details

patch_unsloth_gradient_checkpointing()

# Load model - automatically merges sharded tensors from pytorch_model.bin shards

checkpoint_dir = os.path.abspath(LATEST)  # e.g., ./checkpoints/checkpoint-3000

model = EagleModel.from_pretrained(
    checkpoint_dir,
    torch_dtype=torch.float16,
    device_map="auto"
)

model.eval()
print(f"Model loaded - Total parameters: {sum(p.numel() for p in model.parameters()):,}")

```

The `EagleModel.from_pretrained()` method (implemented via the HuggingFace integration in [`Embodied/eaglevl/model/moon_vit/modeling_vit.py`](https://github.com/NVlabs/Eagle/blob/main/Embodied/eaglevl/model/moon_vit/modeling_vit.py)) handles the reconstruction of sharded tensors automatically, including those saved with DeepSpeed ZeRO-3 or model-parallel strategies.

### Export the Consolidated Checkpoint

Once loaded, export to a single file or HuggingFace-compatible directory structure.

```python
export_dir = "./eagle_deploy"
os.makedirs(export_dir, exist_ok=True)

# Option A: Single torch file for custom inference engines

consolidated_path = os.path.join(export_dir, "eagle_consolidated.pth")
torch.save(model.state_dict(), consolidated_path)
print(f"Saved consolidated weights to: {consolidated_path}")

# Option B: HuggingFace format (recommended for compatibility)

# Creates config.json and model.safetensors/pytorch_model.bin

model.save_pretrained(export_dir)
print(f"Saved HF-compatible checkpoint to: {export_dir}")

```

Consolidating eliminates the need to manage multiple shard files in production and ensures compatibility with inference engines that expect monolithic weight files.

## Optional: Export to ONNX for Optimized Inference

For ultra-fast CPU or GPU inference, convert the consolidated checkpoint to ONNX format.

```python
dummy_input = torch.randn(1, 3, 224, 224).to(next(model.parameters()).device)
onnx_path = os.path.join(export_dir, "eagle_model.onnx")

torch.onnx.export(
    model,
    dummy_input,
    onnx_path,
    opset_version=17,
    input_names=["pixel_values"],
    output_names=["logits"],
    dynamic_axes={"pixel_values": {0: "batch"}, "logits": {0: "batch"}}
)
print(f"ONNX model exported to: {onnx_path}")

```

The resulting ONNX file can be deployed using ONNX Runtime, TensorRT, or other high-performance inference backends.

## Summary

- **Eagle generates sharded checkpoints** during training in `checkpoint-<step>` folders containing `pytorch_model.bin`, `optimizer.pt`, and metadata files.
- **Use `EagleModel.from_pretrained()`** to automatically merge sharded weights, applying the `patch_unsloth_gradient_checkpointing()` utility first if the model was trained with unsloth optimizations.
- **Export via `torch.save()`** for a single consolidated `.pth` file, or use `save_pretrained()` for HuggingFace-compatible directory structures.
- **Consolidated checkpoints** reduce deployment complexity and enable direct loading by production inference servers without training dependencies.

## Frequently Asked Questions

### Why does Eagle save sharded checkpoints instead of a single file?

Eagle leverages HuggingFace Trainer and distributed training strategies (such as DeepSpeed ZeRO-3) that shard model weights across GPUs to reduce memory pressure. Each `pytorch_model.bin` in the checkpoint directory may represent a portion of the full state dict, requiring reassembly during the conversion process described above.

### Do I need the optimizer.pt file when converting checkpoints for deployment?

No. The `optimizer.pt` file contains optimizer states (momentum buffers, Adam moments) necessary for resuming training but irrelevant for inference. When converting and consolidating Eagle model checkpoints for deployment, you only need `pytorch_model.bin` (the weights) and [`config.json`](https://github.com/NVlabs/Eagle/blob/main/config.json) (the architecture definition).

### Can I convert checkpoints trained with DeepSpeed ZeRO-3 using this method?

Yes. The `EagleModel.from_pretrained()` method automatically handles DeepSpeed ZeRO-3 sharded checkpoints. When loading the checkpoint directory, the method reconstructs the full state dict from the sharded `pytorch_model.bin` files, provided you have sufficient CPU or GPU memory to hold the consolidated model.

### What is the difference between torch.save and save_pretrained for Eagle models?

`torch.save(model.state_dict(), file)` creates a single raw PyTorch pickle file containing only weights, suitable for custom inference pipelines. `model.save_pretrained(directory)` creates a standardized HuggingFace directory with [`config.json`](https://github.com/NVlabs/Eagle/blob/main/config.json) and properly formatted weight files (`.bin` or `.safetensors`), enabling seamless reloading via `EagleModel.from_pretrained()` in downstream applications.