How Eagle Handles Variable-Sized Images in Its Encoding Pipeline
Eagle processes images of arbitrary resolution through a dynamic process_images routine in Eagle/eagle/mm_utils.py that either pads images to squares or selects optimal grid resolutions based on the model's image_aspect_ratio configuration.
The NVlabs/Eagle repository implements a flexible image encoding pipeline designed to ingest variable-sized images without manual preprocessing. Unlike rigid vision models that require fixed input dimensions, Eagle's multimodal architecture automatically adapts to any image resolution through intelligent padding or patch-based grid selection. This capability is centralized in the mm_utils.py module, which orchestrates aspect ratio preservation and efficient Vision Transformer tokenization.
The Core Pipeline: process_images
All image preprocessing flows through the process_images function in Eagle/eagle/mm_utils.py. This entry point inspects the model configuration to determine which aspect ratio handling strategy to apply. The function accepts a list of PIL images, a Hugging Face processor, and the model configuration object, then returns either a stacked tensor or a list of tensors depending on whether the batch contains homogeneous or heterogeneous image sizes.
The pipeline branches based on model_cfg.image_aspect_ratio, supporting three distinct processing modes:
"pad"– Expands images to squares usingexpand2square"anyres"– Enables dynamic resolution selection viaprocess_anyres_image- Default – Forwards images directly to the standard Hugging Face processor
Aspect Ratio Handling Modes
Square Padding Mode ("pad")
When the configuration specifies image_aspect_ratio = "pad", the pipeline invokes expand2square (line 81 in mm_utils.py). This function resizes the image to fit within a square canvas while preserving aspect ratio, then pads the remaining area with a background color (black: 0,0,0). This ensures all images conform to the Vision Transformer's expected square input dimensions without distortion.
Any-Resolution Mode ("anyres")
The "anyres" mode activates Eagle's advanced variable-resolution handling through process_anyres_image (line 48). This path is designed to maximize effective resolution while minimizing wasted pixels by dynamically selecting grid layouts defined in the model configuration.
Deep Dive: The Any-Resolution Processing Path
When processing variable-sized images with image_aspect_ratio == "anyres", the pipeline executes a five-step sequence to prepare inputs for the Vision Transformer backbone:
Step 1: Selecting the Best Grid Resolution
The select_best_resolution function (line 41) scans a list of candidate resolutions defined by model_cfg.image_grid_pinpoints. It calculates which grid configuration maximizes the effective resolution of the input image while minimizing padding waste. This selection determines how the image will be subdivided into patches.
Step 2: Resizing and Padding
Once the optimal resolution is selected, resize_and_pad_image (line 71) rescales the image to fit within the chosen grid dimensions while strictly preserving the original aspect ratio. The function pads any remaining space with black pixels (0,0,0), creating a tensor-compatible canvas that aligns with the Vision Transformer's patch boundaries.
Step 3: Dividing into Patches
The padded image is split into non-overlapping patches using divide_to_patches (line 6). Each patch matches the size specified by processor.crop_size['height'], which corresponds to the Vision Transformer's native patch dimensions. This division allows the model to process high-resolution details through localized attention mechanisms.
Step 4: Global View Concatenation
The pipeline creates a low-resolution fallback by generating image_original_resize (lines 69-70), a short-edge-scaled copy of the original image. This global view is concatenated with the patch list, ensuring the model receives both detailed local patches and a complete scene overview. This dual-scale approach prevents information loss in variable-sized images that exceed the grid boundaries.
Step 5: Tensor Encoding and Stacking
Each patch (including the global view) is processed through processor.preprocess(... )['pixel_values'][0]. The resulting tensors are stacked along a new batch dimension, yielding a tensor of shape (num_patches, C, H, W) as implemented in process_anyres_image (line 64). This format allows the Vision Transformer to process multiple patches from a single image as separate tokens.
Batch Processing Behavior
The process_images function (line 95) handles batch aggregation intelligently. When all images in a batch share identical dimensions, they are stacked into a single tensor for efficient parallel processing. However, when processing variable-sized images with differing resolutions, the function returns a list of tensors (one per image), allowing the model to process each image's unique patch count separately.
Configuration Parameters
Two critical configuration fields control the variable-sized image pipeline:
model_cfg.image_aspect_ratio– Determines the processing mode ("pad","anyres", or default)model_cfg.image_grid_pinpoints– Defines the candidate resolution grids for any-resolution mode
These parameters are typically stored in the checkpoint's config.json and loaded automatically when initializing the model.
Practical Implementation Examples
Processing Single Variable-Sized Images
from PIL import Image
from eagle.mm_utils import process_images, tokenizer_image_token
from transformers import AutoProcessor, AutoModelForCausalLM
# Load model and processor
processor = AutoProcessor.from_pretrained("nv-eagle/vision-llama")
model = AutoModelForCausalLM.from_pretrained("nv-eagle/vision-llama")
# Load an arbitrary-resolution image
img = Image.open("high_res_photo.jpg").convert("RGB")
# Automatically handle variable size (anyres or pad based on config)
img_tensor = process_images([img], processor, model.config)[0]
# Prepare prompt with image token
prompt = "Describe the scene in <image>"
input_ids = tokenizer_image_token(prompt, model.get_input_embeddings().tokenizer)
# Generate response
outputs = model.generate(
input_ids=input_ids.unsqueeze(0),
images=img_tensor.unsqueeze(0),
max_new_tokens=100,
)
print(processor.decode(outputs[0], skip_special_tokens=True))
Batched Processing with Mixed Resolutions
import torch
# Load multiple images with different sizes
imgs = [Image.open(p).convert("RGB") for p in ["wide.jpg", "tall.jpg", "square.jpg"]]
# Process returns list when shapes differ, tensor when uniform
batch_tensors = process_images(imgs, processor, model.config)
# Handle mixed-size outputs
if isinstance(batch_tensors, torch.Tensor):
# Homogeneous batch - all same size
image_inputs = batch_tensors
else:
# Variable-sized images - list of tensors
image_inputs = batch_tensors
# Process each image individually when sizes vary
for img_tensor in image_inputs:
out = model.generate(
input_ids=tokenizer_image_token("<image>", model.get_input_embeddings().tokenizer).unsqueeze(0),
images=img_tensor.unsqueeze(0),
max_new_tokens=80,
)
print(processor.decode(out[0], skip_special_tokens=True))
Summary
- Dynamic resolution handling occurs in
Eagle/eagle/mm_utils.pythrough theprocess_imagesfunction. - Two primary modes support variable-sized images:
"pad"for square expansion and"anyres"for intelligent grid selection. - Any-resolution processing selects optimal grids via
select_best_resolution, resizes withresize_and_pad_image, and splits into patches usingdivide_to_patches. - Dual-scale encoding combines high-resolution patches with a low-resolution global view to preserve detail and context.
- Flexible batching returns stacked tensors for uniform images or lists for mixed-resolution batches.
Frequently Asked Questions
What configuration setting controls how Eagle handles variable image sizes?
The image_aspect_ratio field in the model configuration (model_cfg) determines the processing strategy. Set it to "anyres" to enable dynamic grid-based processing of variable-sized images, or "pad" to force square aspect ratios through expansion and padding.
How does the any-resolution mode minimize wasted pixels?
The select_best_resolution function evaluates candidate grids defined in image_grid_pinpoints to find the configuration that maximizes effective resolution while minimizing padding. It selects the grid that most closely matches the input image's aspect ratio, avoiding the excessive padding required by fixed square inputs.
Can Eagle process batches containing images of different sizes?
Yes. When process_images receives variable-sized images that result in different tensor shapes (common with "anyres" mode), it returns a list of tensors rather than a single stacked tensor. Each tensor maintains its unique patch count derived from its specific resolution, allowing the model to handle heterogeneous batches.
Why does the pipeline create a low-resolution fallback image?
The low-resolution image_original_resize serves as a global context vector that is concatenated with the high-resolution patches. This ensures the Vision Transformer maintains awareness of the overall scene composition while attending to fine details in the variable-sized patches, preventing the loss of holistic image understanding that can occur when processing only cropped regions.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →