How to Integrate Eagle with TensorRT for Optimized Inference
You can integrate Eagle with TensorRT by exporting the vision encoder to ONNX using the export_vision_onnx.py utility, building a serialized TensorRT engine with explicit batch profiles, and replacing the PyTorch vision forward pass with TensorRT execution at inference time.
The NVlabs/Eagle repository provides a streamlined deployment pipeline for accelerating vision-language inference with NVIDIA TensorRT. By splitting the computation between a TensorRT-optimized vision encoder and the native language model, you can significantly reduce latency while maintaining accuracy. This guide walks through the exact file paths and code patterns needed to convert Eagle’s ViT-based vision component into a high-performance TensorRT engine.
Prerequisites
Before integrating Eagle with TensorRT, ensure your environment meets these requirements:
- NVIDIA GPU with driver version ≥ 525
- TensorRT ≥ 8.6 Python package (
pip install nvidia-tensorrt) - PyTorch ≥ 2.0 with bfloat16 support
- Eagle source checkout with dependencies installed (
pip install -e .)
Step 1: Export the Vision Encoder to ONNX
Eagle’s inference pipeline separates the vision encoder (ViT-based) from the language model. The repository ships Eagle2_5/deployment/export_vision_onnx.py to automate the ONNX export of the vision component.
Understanding the FeatureExtractorWrapper
In export_vision_onnx.py, the script loads a pretrained Eagle 2.5 model and wraps the vision portion in a FeatureExtractorWrapper class (lines 87-127). This wrapper ensures the model outputs match the expected tensor shapes for downstream processing. The actual export occurs at lines 61-71, where the script calls torch.onnx.export to produce a vision-only ONNX file.
Running the Export Command
Execute the following to generate the ONNX model:
python -m Eagle2_5.deployment.export_vision_onnx \
--config_path ./Eagle2_5/model_configs/eagle2_5_8b.yaml \
--model_path ./checkpoints/eagle2_5_8b \
--onnx_file eagle_vision.onnx
This creates eagle_vision.onnx in your current directory, containing the frozen vision encoder graph.
Step 2: Build the TensorRT Engine
The same export_vision_onnx.py script contains the generate_trt_engine function that converts the ONNX file to an optimized TensorRT engine.
Configuring Builder Profiles
The function configures the TensorRT builder with explicit-batch support. Key configurations include:
- Builder optimization (lines 46-53): Enables FP16 precision and sets memory limits
- Profile setup (lines 66-73): Defines min, opt, and max batch size profiles for dynamic batching
- Serialization: Calls
builder.build_serialized_networkto create the.planfile
Generating the Engine
Build the optimized engine with your desired batch size range:
python -m Eagle2_5.deployment.export_vision_onnx \
--onnx_file eagle_vision.onnx \
--trt_engine eagle_vision.plan \
--minBS 1 --optBS 8 --maxBS 32
The resulting eagle_vision.plan contains the serialized TensorRT engine optimized for your GPU.
Step 3: Integrate TensorRT into Eagle Inference
With the engine built, you can load it at runtime and bypass the PyTorch vision forward pass.
Loading and Executing the TensorRT Engine
Use the TensorRT Python API to deserialize the engine and create an execution context:
import tensorrt as trt
import numpy as np
import torch
TRT_LOGGER = trt.Logger(trt.Logger.VERBOSE)
def load_trt_engine(plan_path: str):
with open(plan_path, "rb") as f:
runtime = trt.Runtime(TRT_LOGGER)
return runtime.deserialize_cuda_engine(f.read())
engine = load_trt_engine("eagle_vision.plan")
context = engine.create_execution_context()
# Prepare input (match the image_size from config, typically 448)
device = torch.cuda.current_device()
batch = 1
img_size = 448
dummy_imgs = torch.randn(batch, 3, img_size, img_size,
device=device, dtype=torch.bfloat16)
# TensorRT expects FP16 numpy buffers
input_np = dummy_imgs.to(torch.float16).contiguous().cpu().numpy()
output_shape = (batch, engine.get_binding_shape(1)[1])
output_np = np.empty(output_shape, dtype=np.float16)
# Set up CUDA bindings
import tensorrt.cuda as trt_cuda
bindings = [
int(trt_cuda.cupy.asarray(input_np).data),
int(trt_cuda.cupy.asarray(output_np).data)
]
# Configure batch dimension and execute
context.set_binding_shape(0, input_np.shape)
context.execute_v2(bindings)
# Convert back to torch tensor
vision_feats = torch.from_numpy(output_np).to(device).to(torch.bfloat16)
Fusing with the Language Model
Inject the TensorRT-computed features into the full Eagle model:
from Eagle import Eagle2_5_VLForConditionalGeneration
# Load the complete model (weights shared with vision encoder)
eagle = Eagle2_5_VLForConditionalGeneration.from_pretrained(
"./checkpoints/eagle2_5_8b",
trust_remote_code=True,
torch_dtype=torch.bfloat16,
).to(device).eval()
# Disable the original PyTorch vision head and inject TRT features
eagle.vision_model = None
eagle.vision_feats = vision_feats
# Generate text using the accelerated features
outputs = eagle.generate(
pixel_values=None,
vision_feats=vision_feats,
max_length=64,
)
print(outputs)
Optional: Accelerating the Language Model
While vision encoder acceleration provides the largest speedup, you can also optimize the language model using Eagle2_5/deployment/export_llm_engine.py. This script converts the LLM checkpoint into TensorRT-LLM format for end-to-end GPU acceleration.
Summary
- Export the vision encoder using
Eagle2_5/deployment/export_vision_onnx.py, which utilizesFeatureExtractorWrapper(lines 87-127) to handle thetorch.onnx.exportcall (lines 61-71). - Build the TensorRT engine with explicit batch profiles (min/opt/max) and FP16 optimization via the
generate_trt_enginefunction (builder configuration at lines 46-53). - Execute by loading the
.planfile withtrt.Runtime, creating an execution context, and replacing the PyTorch vision forward pass withcontext.execute_v2(). - Integrate by setting
eagle.vision_model = Noneand passing pre-computedvision_featsto thegenerate()method.
Frequently Asked Questions
Can I use TensorRT with the full Eagle model or just the vision encoder?
According to the NVlabs/Eagle source code, you can accelerate both components. The vision encoder (Eagle2_5/deployment/export_vision_onnx.py) is the primary target for TensorRT optimization due to its compute intensity. For the language model, use export_llm_engine.py to generate TensorRT-LLM checkpoints, though most users find vision-only acceleration sufficient for latency reduction.
What batch sizes should I specify when building the TensorRT engine?
Configure the min, opt, and max batch sizes based on your expected traffic patterns. The example uses --minBS 1 --optBS 8 --maxBS 32, where the optimal batch size (8) represents your most common inference volume. These values feed into the builder profile setup (lines 66-73 in export_vision_onnx.py) that defines the memory allocation strategy.
How do I handle the bfloat16 to float16 conversion for TensorRT?
Eagle models typically use bfloat16 for weights, but TensorRT engines often run in FP16 for maximum performance. Convert your input tensors using .to(torch.float16) before creating the numpy buffer, then convert the output back to bfloat16 with .to(torch.bfloat16) before feeding features into the language model. This precision conversion occurs in the inference bindings setup shown in the code examples above.
Where can I find the image size configuration for Eagle models?
The default image size is defined in Eagle/eagle/constants.py within the vision_config settings, typically set to 448 pixels for Eagle 2.5 models. When preparing input tensors for TensorRT, ensure your dummy data matches this image_size (e.g., torch.randn(batch, 3, 448, 448)) to maintain alignment with the exported ONNX graph.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →