How to Integrate Eagle with TensorRT for Optimized Inference

You can integrate Eagle with TensorRT by exporting the vision encoder to ONNX using the export_vision_onnx.py utility, building a serialized TensorRT engine with explicit batch profiles, and replacing the PyTorch vision forward pass with TensorRT execution at inference time.

The NVlabs/Eagle repository provides a streamlined deployment pipeline for accelerating vision-language inference with NVIDIA TensorRT. By splitting the computation between a TensorRT-optimized vision encoder and the native language model, you can significantly reduce latency while maintaining accuracy. This guide walks through the exact file paths and code patterns needed to convert Eagle’s ViT-based vision component into a high-performance TensorRT engine.

Prerequisites

Before integrating Eagle with TensorRT, ensure your environment meets these requirements:

  • NVIDIA GPU with driver version ≥ 525
  • TensorRT ≥ 8.6 Python package (pip install nvidia-tensorrt)
  • PyTorch ≥ 2.0 with bfloat16 support
  • Eagle source checkout with dependencies installed (pip install -e .)

Step 1: Export the Vision Encoder to ONNX

Eagle’s inference pipeline separates the vision encoder (ViT-based) from the language model. The repository ships Eagle2_5/deployment/export_vision_onnx.py to automate the ONNX export of the vision component.

Understanding the FeatureExtractorWrapper

In export_vision_onnx.py, the script loads a pretrained Eagle 2.5 model and wraps the vision portion in a FeatureExtractorWrapper class (lines 87-127). This wrapper ensures the model outputs match the expected tensor shapes for downstream processing. The actual export occurs at lines 61-71, where the script calls torch.onnx.export to produce a vision-only ONNX file.

Running the Export Command

Execute the following to generate the ONNX model:

python -m Eagle2_5.deployment.export_vision_onnx \
    --config_path ./Eagle2_5/model_configs/eagle2_5_8b.yaml \
    --model_path ./checkpoints/eagle2_5_8b \
    --onnx_file eagle_vision.onnx

This creates eagle_vision.onnx in your current directory, containing the frozen vision encoder graph.

Step 2: Build the TensorRT Engine

The same export_vision_onnx.py script contains the generate_trt_engine function that converts the ONNX file to an optimized TensorRT engine.

Configuring Builder Profiles

The function configures the TensorRT builder with explicit-batch support. Key configurations include:

  • Builder optimization (lines 46-53): Enables FP16 precision and sets memory limits
  • Profile setup (lines 66-73): Defines min, opt, and max batch size profiles for dynamic batching
  • Serialization: Calls builder.build_serialized_network to create the .plan file

Generating the Engine

Build the optimized engine with your desired batch size range:

python -m Eagle2_5.deployment.export_vision_onnx \
    --onnx_file eagle_vision.onnx \
    --trt_engine eagle_vision.plan \
    --minBS 1 --optBS 8 --maxBS 32

The resulting eagle_vision.plan contains the serialized TensorRT engine optimized for your GPU.

Step 3: Integrate TensorRT into Eagle Inference

With the engine built, you can load it at runtime and bypass the PyTorch vision forward pass.

Loading and Executing the TensorRT Engine

Use the TensorRT Python API to deserialize the engine and create an execution context:

import tensorrt as trt
import numpy as np
import torch

TRT_LOGGER = trt.Logger(trt.Logger.VERBOSE)

def load_trt_engine(plan_path: str):
    with open(plan_path, "rb") as f:
        runtime = trt.Runtime(TRT_LOGGER)
        return runtime.deserialize_cuda_engine(f.read())

engine = load_trt_engine("eagle_vision.plan")
context = engine.create_execution_context()

# Prepare input (match the image_size from config, typically 448)

device = torch.cuda.current_device()
batch = 1
img_size = 448
dummy_imgs = torch.randn(batch, 3, img_size, img_size,
                         device=device, dtype=torch.bfloat16)

# TensorRT expects FP16 numpy buffers

input_np = dummy_imgs.to(torch.float16).contiguous().cpu().numpy()
output_shape = (batch, engine.get_binding_shape(1)[1])
output_np = np.empty(output_shape, dtype=np.float16)

# Set up CUDA bindings

import tensorrt.cuda as trt_cuda
bindings = [
    int(trt_cuda.cupy.asarray(input_np).data),
    int(trt_cuda.cupy.asarray(output_np).data)
]

# Configure batch dimension and execute

context.set_binding_shape(0, input_np.shape)
context.execute_v2(bindings)

# Convert back to torch tensor

vision_feats = torch.from_numpy(output_np).to(device).to(torch.bfloat16)

Fusing with the Language Model

Inject the TensorRT-computed features into the full Eagle model:

from Eagle import Eagle2_5_VLForConditionalGeneration

# Load the complete model (weights shared with vision encoder)

eagle = Eagle2_5_VLForConditionalGeneration.from_pretrained(
    "./checkpoints/eagle2_5_8b",
    trust_remote_code=True,
    torch_dtype=torch.bfloat16,
).to(device).eval()

# Disable the original PyTorch vision head and inject TRT features

eagle.vision_model = None
eagle.vision_feats = vision_feats

# Generate text using the accelerated features

outputs = eagle.generate(
    pixel_values=None,
    vision_feats=vision_feats,
    max_length=64,
)
print(outputs)

Optional: Accelerating the Language Model

While vision encoder acceleration provides the largest speedup, you can also optimize the language model using Eagle2_5/deployment/export_llm_engine.py. This script converts the LLM checkpoint into TensorRT-LLM format for end-to-end GPU acceleration.

Summary

  • Export the vision encoder using Eagle2_5/deployment/export_vision_onnx.py, which utilizes FeatureExtractorWrapper (lines 87-127) to handle the torch.onnx.export call (lines 61-71).
  • Build the TensorRT engine with explicit batch profiles (min/opt/max) and FP16 optimization via the generate_trt_engine function (builder configuration at lines 46-53).
  • Execute by loading the .plan file with trt.Runtime, creating an execution context, and replacing the PyTorch vision forward pass with context.execute_v2().
  • Integrate by setting eagle.vision_model = None and passing pre-computed vision_feats to the generate() method.

Frequently Asked Questions

Can I use TensorRT with the full Eagle model or just the vision encoder?

According to the NVlabs/Eagle source code, you can accelerate both components. The vision encoder (Eagle2_5/deployment/export_vision_onnx.py) is the primary target for TensorRT optimization due to its compute intensity. For the language model, use export_llm_engine.py to generate TensorRT-LLM checkpoints, though most users find vision-only acceleration sufficient for latency reduction.

What batch sizes should I specify when building the TensorRT engine?

Configure the min, opt, and max batch sizes based on your expected traffic patterns. The example uses --minBS 1 --optBS 8 --maxBS 32, where the optimal batch size (8) represents your most common inference volume. These values feed into the builder profile setup (lines 66-73 in export_vision_onnx.py) that defines the memory allocation strategy.

How do I handle the bfloat16 to float16 conversion for TensorRT?

Eagle models typically use bfloat16 for weights, but TensorRT engines often run in FP16 for maximum performance. Convert your input tensors using .to(torch.float16) before creating the numpy buffer, then convert the output back to bfloat16 with .to(torch.bfloat16) before feeding features into the language model. This precision conversion occurs in the inference bindings setup shown in the code examples above.

Where can I find the image size configuration for Eagle models?

The default image size is defined in Eagle/eagle/constants.py within the vision_config settings, typically set to 448 pixels for Eagle 2.5 models. When preparing input tensors for TensorRT, ensure your dummy data matches this image_size (e.g., torch.randn(batch, 3, 448, 448)) to maintain alignment with the exported ONNX graph.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →