Understanding Spatial and Flat Patch Merge Types in Eagle

Eagle multimodal models process visual patches using either a flat merge (1D sequence) or spatial merge (2D grid retention), selected via the mm_patch_merge_type configuration in eagle_arch.py, with the latter preserving geometric relationships for spatial reasoning tasks.

The NVlabs/Eagle repository implements distinct strategies for merging visual patch embeddings from the vision tower into the language model's input space. These spatial and flat patch merge types determine whether the model treats image tokens as a simple sequential list or maintains their two-dimensional spatial arrangement, directly impacting the model's ability to perform grounded visual reasoning.

What Is the mm_patch_merge_type Configuration?

The Eagle architecture controls patch merging behavior through the mm_patch_merge_type argument, typically defined in model configurations or training scripts. According to the source code in Eagle/train.py, the default value is 'flat', which produces a simple flattened sequence. When set to 'spatial' or 'spatial_unpad', the model preserves the 2D grid structure of image patches, enabling spatial awareness in downstream transformer layers.

Flat Patch Merge: Sequential Processing

The flat merge type treats visual patches as a one-dimensional token stream. In Eagle/eagle/model/eagle_arch.py (lines 180-182), the implementation flattens patch embeddings using x.flatten(0, 1), resulting in a tensor of shape [num_patches, hidden_dim]. This approach concatenates visual tokens with text tokens without preserving any geometric information about the original image layout.

Use this approach for standard image captioning, general visual question answering, or any task where the absolute spatial position of visual features is less critical than the semantic content.

Spatial Patch Merge: Preserving Geometric Layout

When mm_patch_merge_type starts with 'spatial', the model retains the 2D grid structure. The code in eagle_arch.py (lines 184-198) separates the CLS token from patch embeddings, then reshapes the remaining patches from [num_patches, hidden_dim] back to [H, W, hidden_dim] where H and W represent the height and width of the patch grid. This preservation of spatial relationships allows the transformer to attend to visual features based on their relative geometric positions, essential for object grounding and spatial reasoning tasks.

Spatial Unpad for Arbitrary Aspect Ratios

The spatial_unpad variant extends spatial merging by removing padding artifacts from resized images. When processing images with varying aspect ratios, the model calls unpad_image (defined in eagle_arch.py, lines 195-199) to strip padded borders added during the square-grid resizing process. The code also inserts a special image_newline token to demarcate row boundaries, helping the model distinguish between actual image content and padding regions. The helper function get_anyres_image_grid_shape in Eagle/eagle/mm_utils.py calculates the correct grid dimensions for arbitrary resolution images in this pathway.

Implementation Details in Eagle Source Code

The core logic resides in Eagle/eagle/model/eagle_arch.py. For flat merging, the implementation is straightforward:

if mm_patch_merge_type == 'flat':
    # Results in shape [num_patches, hidden_dim]

    image_features = [x.flatten(0, 1) for x in image_features]

For spatial merging, the process involves reshaping the patch grid and optionally unpadding:

elif mm_patch_merge_type.startswith('spatial'):
    new_image_features = []
    for img_idx, img_feat in enumerate(image_features):
        base = img_feat[0]  # CLS token

        grid = img_feat[1:]  # Patch embeddings

        
        # Reshape to 2D grid

        H = W = vision_tower.num_patches_per_side
        grid = grid.view(H, W, -1)
        
        if "unpad" in mm_patch_merge_type:
            # Remove padding and add newline token

            grid = grid.permute(2, 0, 1).unsqueeze(0)
            grid = unpad_image(grid, image_sizes[img_idx])
            grid = torch.cat([grid, 
                              model.image_newline[:, None, None].expand(-1, *grid.shape[2:])], 
                             dim=0)
        
        new_image_features.append(torch.cat([base, grid], dim=0))

Practical Configuration Examples

Configure the merge type in your model arguments:

from dataclasses import dataclass

@dataclass
class ModelArgs:
    vision_tower: str = "eva_vit_g"
    mm_patch_merge_type: str = "flat"  # Change to "spatial_unpad" for spatial merge

The resulting tensor shapes differ significantly between modes:

  • flat: [num_patches + 1, hidden] — A single flattened sequence including the CLS token.
  • spatial: [1 + H*W, hidden] — CLS token followed by spatially-ordered patches.
  • spatial_unpad: Variable length based on original image aspect ratio, with padding removed and newline tokens inserted.

When to Use Each Merge Type

  • Flat: Use for standard multimodal tasks like image captioning or generic VQA where spatial coordinates are irrelevant.
  • Spatial: Use when the model must reason about object locations, spatial relationships, or perform referring expression comprehension.
  • Spatial Unpad: Use when processing images with diverse aspect ratios where padding artifacts would degrade performance, such as in document understanding or panoramic image analysis.

Summary

  • Eagle provides two primary merge strategies via mm_patch_merge_type in eagle_arch.py: flat and spatial (including the spatial_unpad variant).
  • Flat merging (lines 180-182) creates a 1D sequence by flattening all patches, discarding geometric information.
  • Spatial merging (lines 184-198) reshapes patches back to a 2D grid [H, W, hidden_dim], preserving spatial relationships for geometric reasoning.
  • Spatial Unpad calls unpad_image (lines 195-199) to remove padding and inserts image_newline tokens for arbitrary aspect ratio handling.
  • Flat is the default in train.py, while spatial variants excel at grounded visual understanding tasks.

Frequently Asked Questions

What is the default patch merge type in Eagle?

The default value is 'flat', as specified in the training configuration within Eagle/train.py. This setting treats visual patches as a simple sequential token stream compatible with standard transformer language models.

How does spatial patch merge affect model performance?

Spatial merging enables superior performance on tasks requiring spatial reasoning, such as object grounding, referring expression comprehension, and geometric relationship prediction. By preserving the 2D grid structure in eagle_arch.py, the model can attend to visual features based on their relative positions, though it may require more careful handling of variable image sizes.

What is the purpose of the spatial_unpad variant?

The spatial_unpad option removes padding artifacts that are added when resizing images to a fixed square grid. According to the implementation in eagle_arch.py, it calls unpad_image to strip padded borders and inserts a special image_newline token to mark row boundaries, making it ideal for processing images with arbitrary aspect ratios without introducing padding noise.

Where is the patch merge logic implemented in the Eagle codebase?

The core logic resides in Eagle/eagle/model/eagle_arch.py, specifically between lines 180-198. The unpad_image function is defined in the same file around lines 195-199, while Eagle/eagle/mm_utils.py provides the get_anyres_image_grid_shape helper for calculating grid dimensions during spatial processing.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →