How Eagle Handles Variable-Sized Images in Its Encoding Pipeline

Eagle processes images of arbitrary resolution through a dynamic process_images routine in Eagle/eagle/mm_utils.py that either pads images to squares or selects optimal grid resolutions based on the model's image_aspect_ratio configuration.

The NVlabs/Eagle repository implements a flexible image encoding pipeline designed to ingest variable-sized images without manual preprocessing. Unlike rigid vision models that require fixed input dimensions, Eagle's multimodal architecture automatically adapts to any image resolution through intelligent padding or patch-based grid selection. This capability is centralized in the mm_utils.py module, which orchestrates aspect ratio preservation and efficient Vision Transformer tokenization.

The Core Pipeline: process_images

All image preprocessing flows through the process_images function in Eagle/eagle/mm_utils.py. This entry point inspects the model configuration to determine which aspect ratio handling strategy to apply. The function accepts a list of PIL images, a Hugging Face processor, and the model configuration object, then returns either a stacked tensor or a list of tensors depending on whether the batch contains homogeneous or heterogeneous image sizes.

The pipeline branches based on model_cfg.image_aspect_ratio, supporting three distinct processing modes:

  • "pad" – Expands images to squares using expand2square
  • "anyres" – Enables dynamic resolution selection via process_anyres_image
  • Default – Forwards images directly to the standard Hugging Face processor

Aspect Ratio Handling Modes

Square Padding Mode ("pad")

When the configuration specifies image_aspect_ratio = "pad", the pipeline invokes expand2square (line 81 in mm_utils.py). This function resizes the image to fit within a square canvas while preserving aspect ratio, then pads the remaining area with a background color (black: 0,0,0). This ensures all images conform to the Vision Transformer's expected square input dimensions without distortion.

Any-Resolution Mode ("anyres")

The "anyres" mode activates Eagle's advanced variable-resolution handling through process_anyres_image (line 48). This path is designed to maximize effective resolution while minimizing wasted pixels by dynamically selecting grid layouts defined in the model configuration.

Deep Dive: The Any-Resolution Processing Path

When processing variable-sized images with image_aspect_ratio == "anyres", the pipeline executes a five-step sequence to prepare inputs for the Vision Transformer backbone:

Step 1: Selecting the Best Grid Resolution

The select_best_resolution function (line 41) scans a list of candidate resolutions defined by model_cfg.image_grid_pinpoints. It calculates which grid configuration maximizes the effective resolution of the input image while minimizing padding waste. This selection determines how the image will be subdivided into patches.

Step 2: Resizing and Padding

Once the optimal resolution is selected, resize_and_pad_image (line 71) rescales the image to fit within the chosen grid dimensions while strictly preserving the original aspect ratio. The function pads any remaining space with black pixels (0,0,0), creating a tensor-compatible canvas that aligns with the Vision Transformer's patch boundaries.

Step 3: Dividing into Patches

The padded image is split into non-overlapping patches using divide_to_patches (line 6). Each patch matches the size specified by processor.crop_size['height'], which corresponds to the Vision Transformer's native patch dimensions. This division allows the model to process high-resolution details through localized attention mechanisms.

Step 4: Global View Concatenation

The pipeline creates a low-resolution fallback by generating image_original_resize (lines 69-70), a short-edge-scaled copy of the original image. This global view is concatenated with the patch list, ensuring the model receives both detailed local patches and a complete scene overview. This dual-scale approach prevents information loss in variable-sized images that exceed the grid boundaries.

Step 5: Tensor Encoding and Stacking

Each patch (including the global view) is processed through processor.preprocess(... )['pixel_values'][0]. The resulting tensors are stacked along a new batch dimension, yielding a tensor of shape (num_patches, C, H, W) as implemented in process_anyres_image (line 64). This format allows the Vision Transformer to process multiple patches from a single image as separate tokens.

Batch Processing Behavior

The process_images function (line 95) handles batch aggregation intelligently. When all images in a batch share identical dimensions, they are stacked into a single tensor for efficient parallel processing. However, when processing variable-sized images with differing resolutions, the function returns a list of tensors (one per image), allowing the model to process each image's unique patch count separately.

Configuration Parameters

Two critical configuration fields control the variable-sized image pipeline:

  • model_cfg.image_aspect_ratio – Determines the processing mode ("pad", "anyres", or default)
  • model_cfg.image_grid_pinpoints – Defines the candidate resolution grids for any-resolution mode

These parameters are typically stored in the checkpoint's config.json and loaded automatically when initializing the model.

Practical Implementation Examples

Processing Single Variable-Sized Images

from PIL import Image
from eagle.mm_utils import process_images, tokenizer_image_token
from transformers import AutoProcessor, AutoModelForCausalLM

# Load model and processor

processor = AutoProcessor.from_pretrained("nv-eagle/vision-llama")
model = AutoModelForCausalLM.from_pretrained("nv-eagle/vision-llama")

# Load an arbitrary-resolution image

img = Image.open("high_res_photo.jpg").convert("RGB")

# Automatically handle variable size (anyres or pad based on config)

img_tensor = process_images([img], processor, model.config)[0]

# Prepare prompt with image token

prompt = "Describe the scene in <image>"
input_ids = tokenizer_image_token(prompt, model.get_input_embeddings().tokenizer)

# Generate response

outputs = model.generate(
    input_ids=input_ids.unsqueeze(0),
    images=img_tensor.unsqueeze(0),
    max_new_tokens=100,
)
print(processor.decode(outputs[0], skip_special_tokens=True))

Batched Processing with Mixed Resolutions

import torch

# Load multiple images with different sizes

imgs = [Image.open(p).convert("RGB") for p in ["wide.jpg", "tall.jpg", "square.jpg"]]

# Process returns list when shapes differ, tensor when uniform

batch_tensors = process_images(imgs, processor, model.config)

# Handle mixed-size outputs

if isinstance(batch_tensors, torch.Tensor):
    # Homogeneous batch - all same size

    image_inputs = batch_tensors
else:
    # Variable-sized images - list of tensors

    image_inputs = batch_tensors

# Process each image individually when sizes vary

for img_tensor in image_inputs:
    out = model.generate(
        input_ids=tokenizer_image_token("<image>", model.get_input_embeddings().tokenizer).unsqueeze(0),
        images=img_tensor.unsqueeze(0),
        max_new_tokens=80,
    )
    print(processor.decode(out[0], skip_special_tokens=True))

Summary

  • Dynamic resolution handling occurs in Eagle/eagle/mm_utils.py through the process_images function.
  • Two primary modes support variable-sized images: "pad" for square expansion and "anyres" for intelligent grid selection.
  • Any-resolution processing selects optimal grids via select_best_resolution, resizes with resize_and_pad_image, and splits into patches using divide_to_patches.
  • Dual-scale encoding combines high-resolution patches with a low-resolution global view to preserve detail and context.
  • Flexible batching returns stacked tensors for uniform images or lists for mixed-resolution batches.

Frequently Asked Questions

What configuration setting controls how Eagle handles variable image sizes?

The image_aspect_ratio field in the model configuration (model_cfg) determines the processing strategy. Set it to "anyres" to enable dynamic grid-based processing of variable-sized images, or "pad" to force square aspect ratios through expansion and padding.

How does the any-resolution mode minimize wasted pixels?

The select_best_resolution function evaluates candidate grids defined in image_grid_pinpoints to find the configuration that maximizes effective resolution while minimizing padding. It selects the grid that most closely matches the input image's aspect ratio, avoiding the excessive padding required by fixed square inputs.

Can Eagle process batches containing images of different sizes?

Yes. When process_images receives variable-sized images that result in different tensor shapes (common with "anyres" mode), it returns a list of tensors rather than a single stacked tensor. Each tensor maintains its unique patch count derived from its specific resolution, allowing the model to handle heterogeneous batches.

Why does the pipeline create a low-resolution fallback image?

The low-resolution image_original_resize serves as a global context vector that is concatenated with the high-resolution patches. This ensures the Vision Transformer maintains awareness of the overall scene composition while attending to fine details in the variable-sized patches, preventing the loss of holistic image understanding that can occur when processing only cropped regions.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →