# What Is AnyRes and How It Handles Multiple Image Resolutions in Eagle VLM

> Discover AnyRes, a PyTorch dataset wrapper used by Eagle VLM to train on diverse image resolutions. Learn how it manages variable-length lists and token-level flags for efficient multi-resolution handling.

- Repository: [NVIDIA Research Projects/Eagle](https://github.com/NVlabs/Eagle)
- Tags: tutorial
- Published: 2026-06-28

---

**AnyRes is a PyTorch dataset wrapper that enables Eagle VLM to train on batches containing images of different resolutions by keeping pixel tensors as a variable-length list and tracking image positions with token-level flags.**

**AnyRes** (implemented as the class `ConcatDatasetForOnlinePacking_AnyRes`) is a critical component in the NVlabs/Eagle training pipeline that eliminates the traditional requirement to resize all images to a fixed resolution. According to the Eagle source code, this lightweight wrapper extends PyTorch's `ConcatDataset` to support heterogeneous image sizes within a single training batch, allowing the model to learn from visual corpora with varying aspect ratios and scales.

## The AnyRes Architecture

The implementation lives in [`Embodied/eaglevl/train/dataset.py`](https://github.com/NVlabs/Eagle/blob/main/Embodied/eaglevl/train/dataset.py) (lines 45‑55), where the class inherits from `ConcatDataset` and overrides key methods to handle resolution-agnostic batching.

### Index Translation Across Sub-Datasets

The wrapper first maps global indices to the appropriate sub-dataset using `__getitem_for_int_idx__` and `__get_raw_data_for_int_idx__`. Because each sub-dataset may contain images of a different resolution, these methods correctly locate samples without assuming uniform image dimensions. This allows you to concatenate datasets where one contains 224×224 images and another contains 384×384 images without conflict.

### The Packing Mechanism

When the trainer requests a batch via `__getitem__`, the method receives a `Packer` object (a list of indices). It calls `pack_data(ret_list)` to aggregate samples from potentially different sub-datasets. Unlike standard dataloaders that require tensor stacks to share identical shapes, **AnyRes preserves each image's native resolution** by keeping `pixel_values` as a Python list rather than a single stacked tensor.

### Variable-Resolution Tensor Handling

Inside `pack_data`, the concatenation logic specifically handles the image data as follows:

```python
pixel_values = [image for each in ret_list for image in each['pixel_values']]

```

This list comprehension preserves each image tensor with its original spatial shape, meaning one element might have shape `[1, 3, 224, 224]` while another has `[1, 3, 384, 384]`. The method simultaneously builds **image_flags**—a tensor that indicates which token positions in the sequence correspond to image data—allowing the model to process each resolution independently during the forward pass.

### Dynamic Padding Strategy

For the text components (`input_ids`, `labels`, and `attention_masks`), AnyRes applies dynamic padding to match the model's `model_max_length`. When a packed batch falls short of this length, the code pads with appropriate token IDs. This guarantees that the text side of the batch always fits the model's context window, while the image side can vary in resolution without constraint.

### Dummy Image Fallback

When a batch contains no images, the implementation inserts a minimal dummy tensor with shape `1 × 3 × 28 × 28` and sets `image_flags` to 0. This ensures that downstream model code can always assume an image tensor exists in the batch, even if the resolution is trivial and the flags indicate no actual image content should be processed.

## Tokenization Support for Variable Resolutions

AnyRes coordinates with the tokenizer to allocate variable numbers of image tokens per resolution. In the `preprocess_mpt` function (same file, lines 92‑122), the `num_image_token` parameter accepts either an integer for fixed token counts or a **`numpy.ndarray`** where each element specifies the token count for a specific image resolution.

```python
elif type(num_image_token) == np.ndarray:
    # Build a per-image token string

    image_tokens_list.append(
        f'<image {idx+1}>{IMG_START_TOKEN}'
        f'{IMG_CONTEXT_TOKEN * int(num_token_per_image)}{IMG_END_TOKEN}'
    )

```

When a dataset sample contains a larger-resolution image, the tokenizer allocates more image tokens for that sample, while smaller-resolution images receive fewer tokens. The packing logic in `ConcatDatasetForOnlinePacking_AnyRes` maintains this alignment by keeping the raw pixel tensors separate, ensuring the model receives the exact resolution each image was originally encoded with.

## Source Code Implementation

The primary definition resides in:

- **[`Embodied/eaglevl/train/dataset.py`](https://github.com/NVlabs/Eagle/blob/main/Embodied/eaglevl/train/dataset.py)** — Defines `ConcatDatasetForOnlinePacking_AnyRes` (lines 45‑55) and the `pack_data` method that preserves per-sample image resolution.
- **[`Eagle2_5/eaglevl/train/dataset.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle2_5/eaglevl/train/dataset.py)** — Contains the same implementation for the newer demo version, useful if you are exploring the Streamlit demo.

The `preprocess_mpt` function within these files demonstrates how the training pipeline handles **per-resolution token counts** using numpy arrays, enabling the variable token allocation described above.

## Practical Usage Example

To use AnyRes in your training pipeline:

```python
from torch.utils.data import ConcatDataset
from eaglevl.train.constants import IMG_START_TOKEN, IMG_CONTEXT_TOKEN, IMG_END_TOKEN
from eaglevl.train.dataset import ConcatDatasetForOnlinePacking_AnyRes

# Assume two sub-datasets with different image sizes

ds_small = MyImageDataset(root='data/small_res')   # images 224×224

ds_large = MyImageDataset(root='data/large_res')   # images 384×384

# Concatenate with AnyRes wrapper

mixed_res_dataset = ConcatDatasetForOnlinePacking_AnyRes([ds_small, ds_large])

# During training, the Packer can mix indices from both datasets

packer = Packer([0, 15, 30])          # arbitrary indices across both datasets

batch = mixed_res_dataset[packer]     # Returns a dict with:

#   - input_ids, labels, attention_mask (text tensors)

#   - pixel_values   # list of tensors with different H×W shapes

#   - image_flags    # indicates which positions refer to images

```

## Summary

- **AnyRes** is implemented as `ConcatDatasetForOnlinePacking_AnyRes` in [`Embodied/eaglevl/train/dataset.py`](https://github.com/NVlabs/Eagle/blob/main/Embodied/eaglevl/train/dataset.py) and enables multi-resolution training batches.
- The wrapper preserves image resolution by storing `pixel_values` as a Python list of tensors rather than a single batched tensor.
- **Image flags** track token positions to guide the model in processing variable-resolution images independently.
- Dynamic padding ensures text sequences fit the model's maximum length while images maintain their original dimensions.
- The tokenizer supports per-resolution token allocation via `numpy.ndarray` inputs to `num_image_token` in `preprocess_mpt`.
- A dummy tensor fallback (`1 × 3 × 28 × 28`) ensures downstream code handles text-only batches without errors.

## Frequently Asked Questions

### How does AnyRes differ from the standard PyTorch ConcatDataset?

**`ConcatDatasetForOnlinePacking_AnyRes`** extends the standard `ConcatDataset` by adding resolution-aware packing logic. While the base `ConcatDataset` simply concatenates datasets and assumes uniform tensor shapes for batching, AnyRes overrides `__getitem__` and implements `pack_data` to handle samples with different image resolutions, specifically by keeping `pixel_values` as a list and using `image_flags` to track positions.

### Why does AnyRes keep pixel_values as a list instead of stacking tensors?

Keeping **pixel_values** as a list allows each image in the batch to maintain its original spatial dimensions. If the code stacked these into a single tensor, PyTorch would require all images to share identical height and width dimensions. By deferring the stacking operation and using a list, AnyRes supports heterogeneous resolutions where one image might be 224×224 and another 384×384 within the same batch.

### How does the model know which tokens correspond to high-resolution vs. low-resolution images?

The model uses the **`image_flags`** tensor generated during packing to identify which token positions contain image data. Additionally, the tokenizer's `preprocess_mpt` function allocates a specific number of image tokens based on the resolution—passed as a `numpy.ndarray` to `num_image_token`—so high-resolution images receive more tokens in the sequence than low-resolution images. The model processes these tokens according to the flags and the known token counts per image.

### Can AnyRes handle batches with no images at all?

Yes. When a batch contains only text samples, AnyRes inserts a dummy tensor with shape **`1 × 3 × 28 × 28`** and sets `image_flags` to 0. This ensures that the model's forward pass always receives an image tensor, but the zeroed flags indicate that no actual image processing should occur for those positions, preventing errors in downstream layers.