What Is AnyRes and How It Handles Multiple Image Resolutions in Eagle VLM
AnyRes is a PyTorch dataset wrapper that enables Eagle VLM to train on batches containing images of different resolutions by keeping pixel tensors as a variable-length list and tracking image positions with token-level flags.
AnyRes (implemented as the class ConcatDatasetForOnlinePacking_AnyRes) is a critical component in the NVlabs/Eagle training pipeline that eliminates the traditional requirement to resize all images to a fixed resolution. According to the Eagle source code, this lightweight wrapper extends PyTorch's ConcatDataset to support heterogeneous image sizes within a single training batch, allowing the model to learn from visual corpora with varying aspect ratios and scales.
The AnyRes Architecture
The implementation lives in Embodied/eaglevl/train/dataset.py (lines 45‑55), where the class inherits from ConcatDataset and overrides key methods to handle resolution-agnostic batching.
Index Translation Across Sub-Datasets
The wrapper first maps global indices to the appropriate sub-dataset using __getitem_for_int_idx__ and __get_raw_data_for_int_idx__. Because each sub-dataset may contain images of a different resolution, these methods correctly locate samples without assuming uniform image dimensions. This allows you to concatenate datasets where one contains 224×224 images and another contains 384×384 images without conflict.
The Packing Mechanism
When the trainer requests a batch via __getitem__, the method receives a Packer object (a list of indices). It calls pack_data(ret_list) to aggregate samples from potentially different sub-datasets. Unlike standard dataloaders that require tensor stacks to share identical shapes, AnyRes preserves each image's native resolution by keeping pixel_values as a Python list rather than a single stacked tensor.
Variable-Resolution Tensor Handling
Inside pack_data, the concatenation logic specifically handles the image data as follows:
pixel_values = [image for each in ret_list for image in each['pixel_values']]
This list comprehension preserves each image tensor with its original spatial shape, meaning one element might have shape [1, 3, 224, 224] while another has [1, 3, 384, 384]. The method simultaneously builds image_flags—a tensor that indicates which token positions in the sequence correspond to image data—allowing the model to process each resolution independently during the forward pass.
Dynamic Padding Strategy
For the text components (input_ids, labels, and attention_masks), AnyRes applies dynamic padding to match the model's model_max_length. When a packed batch falls short of this length, the code pads with appropriate token IDs. This guarantees that the text side of the batch always fits the model's context window, while the image side can vary in resolution without constraint.
Dummy Image Fallback
When a batch contains no images, the implementation inserts a minimal dummy tensor with shape 1 × 3 × 28 × 28 and sets image_flags to 0. This ensures that downstream model code can always assume an image tensor exists in the batch, even if the resolution is trivial and the flags indicate no actual image content should be processed.
Tokenization Support for Variable Resolutions
AnyRes coordinates with the tokenizer to allocate variable numbers of image tokens per resolution. In the preprocess_mpt function (same file, lines 92‑122), the num_image_token parameter accepts either an integer for fixed token counts or a numpy.ndarray where each element specifies the token count for a specific image resolution.
elif type(num_image_token) == np.ndarray:
# Build a per-image token string
image_tokens_list.append(
f'<image {idx+1}>{IMG_START_TOKEN}'
f'{IMG_CONTEXT_TOKEN * int(num_token_per_image)}{IMG_END_TOKEN}'
)
When a dataset sample contains a larger-resolution image, the tokenizer allocates more image tokens for that sample, while smaller-resolution images receive fewer tokens. The packing logic in ConcatDatasetForOnlinePacking_AnyRes maintains this alignment by keeping the raw pixel tensors separate, ensuring the model receives the exact resolution each image was originally encoded with.
Source Code Implementation
The primary definition resides in:
Embodied/eaglevl/train/dataset.py— DefinesConcatDatasetForOnlinePacking_AnyRes(lines 45‑55) and thepack_datamethod that preserves per-sample image resolution.Eagle2_5/eaglevl/train/dataset.py— Contains the same implementation for the newer demo version, useful if you are exploring the Streamlit demo.
The preprocess_mpt function within these files demonstrates how the training pipeline handles per-resolution token counts using numpy arrays, enabling the variable token allocation described above.
Practical Usage Example
To use AnyRes in your training pipeline:
from torch.utils.data import ConcatDataset
from eaglevl.train.constants import IMG_START_TOKEN, IMG_CONTEXT_TOKEN, IMG_END_TOKEN
from eaglevl.train.dataset import ConcatDatasetForOnlinePacking_AnyRes
# Assume two sub-datasets with different image sizes
ds_small = MyImageDataset(root='data/small_res') # images 224×224
ds_large = MyImageDataset(root='data/large_res') # images 384×384
# Concatenate with AnyRes wrapper
mixed_res_dataset = ConcatDatasetForOnlinePacking_AnyRes([ds_small, ds_large])
# During training, the Packer can mix indices from both datasets
packer = Packer([0, 15, 30]) # arbitrary indices across both datasets
batch = mixed_res_dataset[packer] # Returns a dict with:
# - input_ids, labels, attention_mask (text tensors)
# - pixel_values # list of tensors with different H×W shapes
# - image_flags # indicates which positions refer to images
Summary
- AnyRes is implemented as
ConcatDatasetForOnlinePacking_AnyResinEmbodied/eaglevl/train/dataset.pyand enables multi-resolution training batches. - The wrapper preserves image resolution by storing
pixel_valuesas a Python list of tensors rather than a single batched tensor. - Image flags track token positions to guide the model in processing variable-resolution images independently.
- Dynamic padding ensures text sequences fit the model's maximum length while images maintain their original dimensions.
- The tokenizer supports per-resolution token allocation via
numpy.ndarrayinputs tonum_image_tokeninpreprocess_mpt. - A dummy tensor fallback (
1 × 3 × 28 × 28) ensures downstream code handles text-only batches without errors.
Frequently Asked Questions
How does AnyRes differ from the standard PyTorch ConcatDataset?
ConcatDatasetForOnlinePacking_AnyRes extends the standard ConcatDataset by adding resolution-aware packing logic. While the base ConcatDataset simply concatenates datasets and assumes uniform tensor shapes for batching, AnyRes overrides __getitem__ and implements pack_data to handle samples with different image resolutions, specifically by keeping pixel_values as a list and using image_flags to track positions.
Why does AnyRes keep pixel_values as a list instead of stacking tensors?
Keeping pixel_values as a list allows each image in the batch to maintain its original spatial dimensions. If the code stacked these into a single tensor, PyTorch would require all images to share identical height and width dimensions. By deferring the stacking operation and using a list, AnyRes supports heterogeneous resolutions where one image might be 224×224 and another 384×384 within the same batch.
How does the model know which tokens correspond to high-resolution vs. low-resolution images?
The model uses the image_flags tensor generated during packing to identify which token positions contain image data. Additionally, the tokenizer's preprocess_mpt function allocates a specific number of image tokens based on the resolution—passed as a numpy.ndarray to num_image_token—so high-resolution images receive more tokens in the sequence than low-resolution images. The model processes these tokens according to the flags and the known token counts per image.
Can AnyRes handle batches with no images at all?
Yes. When a batch contains only text samples, AnyRes inserts a dummy tensor with shape 1 × 3 × 28 × 28 and sets image_flags to 0. This ensures that the model's forward pass always receives an image tensor, but the zeroed flags indicate that no actual image processing should occur for those positions, preventing errors in downstream layers.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →