# Understanding Spatial and Flat Patch Merge Types in Eagle

> Discover the difference between spatial and flat patch merge types in Eagle models. Learn how mm_patch_merge_type configures 1D vs 2D patch processing for enhanced visual understanding.

- Repository: [NVIDIA Research Projects/Eagle](https://github.com/NVlabs/Eagle)
- Tags: deep-dive
- Published: 2026-06-28

---

**Eagle multimodal models process visual patches using either a `flat` merge (1D sequence) or `spatial` merge (2D grid retention), selected via the `mm_patch_merge_type` configuration in [`eagle_arch.py`](https://github.com/NVlabs/Eagle/blob/main/eagle_arch.py), with the latter preserving geometric relationships for spatial reasoning tasks.**

The NVlabs/Eagle repository implements distinct strategies for merging visual patch embeddings from the vision tower into the language model's input space. These spatial and flat patch merge types determine whether the model treats image tokens as a simple sequential list or maintains their two-dimensional spatial arrangement, directly impacting the model's ability to perform grounded visual reasoning.

## What Is the `mm_patch_merge_type` Configuration?

The Eagle architecture controls patch merging behavior through the **`mm_patch_merge_type`** argument, typically defined in model configurations or training scripts. According to the source code in [`Eagle/train.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/train.py), the default value is `'flat'`, which produces a simple flattened sequence. When set to `'spatial'` or `'spatial_unpad'`, the model preserves the 2D grid structure of image patches, enabling spatial awareness in downstream transformer layers.

## Flat Patch Merge: Sequential Processing

The **`flat`** merge type treats visual patches as a one-dimensional token stream. In [`Eagle/eagle/model/eagle_arch.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/eagle/model/eagle_arch.py) (lines 180-182), the implementation flattens patch embeddings using `x.flatten(0, 1)`, resulting in a tensor of shape `[num_patches, hidden_dim]`. This approach concatenates visual tokens with text tokens without preserving any geometric information about the original image layout.

Use this approach for standard image captioning, general visual question answering, or any task where the absolute spatial position of visual features is less critical than the semantic content.

## Spatial Patch Merge: Preserving Geometric Layout

When `mm_patch_merge_type` starts with `'spatial'`, the model retains the 2D grid structure. The code in [`eagle_arch.py`](https://github.com/NVlabs/Eagle/blob/main/eagle_arch.py) (lines 184-198) separates the CLS token from patch embeddings, then reshapes the remaining patches from `[num_patches, hidden_dim]` back to `[H, W, hidden_dim]` where `H` and `W` represent the height and width of the patch grid. This preservation of spatial relationships allows the transformer to attend to visual features based on their relative geometric positions, essential for object grounding and spatial reasoning tasks.

### Spatial Unpad for Arbitrary Aspect Ratios

The **`spatial_unpad`** variant extends spatial merging by removing padding artifacts from resized images. When processing images with varying aspect ratios, the model calls **`unpad_image`** (defined in [`eagle_arch.py`](https://github.com/NVlabs/Eagle/blob/main/eagle_arch.py), lines 195-199) to strip padded borders added during the square-grid resizing process. The code also inserts a special **`image_newline`** token to demarcate row boundaries, helping the model distinguish between actual image content and padding regions. The helper function **`get_anyres_image_grid_shape`** in [`Eagle/eagle/mm_utils.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/eagle/mm_utils.py) calculates the correct grid dimensions for arbitrary resolution images in this pathway.

## Implementation Details in Eagle Source Code

The core logic resides in [`Eagle/eagle/model/eagle_arch.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/eagle/model/eagle_arch.py). For flat merging, the implementation is straightforward:

```python
if mm_patch_merge_type == 'flat':
    # Results in shape [num_patches, hidden_dim]

    image_features = [x.flatten(0, 1) for x in image_features]

```

For spatial merging, the process involves reshaping the patch grid and optionally unpadding:

```python
elif mm_patch_merge_type.startswith('spatial'):
    new_image_features = []
    for img_idx, img_feat in enumerate(image_features):
        base = img_feat[0]  # CLS token

        grid = img_feat[1:]  # Patch embeddings

        
        # Reshape to 2D grid

        H = W = vision_tower.num_patches_per_side
        grid = grid.view(H, W, -1)
        
        if "unpad" in mm_patch_merge_type:
            # Remove padding and add newline token

            grid = grid.permute(2, 0, 1).unsqueeze(0)
            grid = unpad_image(grid, image_sizes[img_idx])
            grid = torch.cat([grid, 
                              model.image_newline[:, None, None].expand(-1, *grid.shape[2:])], 
                             dim=0)
        
        new_image_features.append(torch.cat([base, grid], dim=0))

```

## Practical Configuration Examples

Configure the merge type in your model arguments:

```python
from dataclasses import dataclass

@dataclass
class ModelArgs:
    vision_tower: str = "eva_vit_g"
    mm_patch_merge_type: str = "flat"  # Change to "spatial_unpad" for spatial merge

```

The resulting tensor shapes differ significantly between modes:

- **`flat`**: `[num_patches + 1, hidden]` — A single flattened sequence including the CLS token.
- **`spatial`**: `[1 + H*W, hidden]` — CLS token followed by spatially-ordered patches.
- **`spatial_unpad`**: Variable length based on original image aspect ratio, with padding removed and newline tokens inserted.

## When to Use Each Merge Type

- **Flat**: Use for standard multimodal tasks like image captioning or generic VQA where spatial coordinates are irrelevant.
- **Spatial**: Use when the model must reason about object locations, spatial relationships, or perform referring expression comprehension.
- **Spatial Unpad**: Use when processing images with diverse aspect ratios where padding artifacts would degrade performance, such as in document understanding or panoramic image analysis.

## Summary

- **Eagle** provides two primary merge strategies via `mm_patch_merge_type` in [`eagle_arch.py`](https://github.com/NVlabs/Eagle/blob/main/eagle_arch.py): `flat` and `spatial` (including the `spatial_unpad` variant).
- **Flat** merging (lines 180-182) creates a 1D sequence by flattening all patches, discarding geometric information.
- **Spatial** merging (lines 184-198) reshapes patches back to a 2D grid `[H, W, hidden_dim]`, preserving spatial relationships for geometric reasoning.
- **Spatial Unpad** calls `unpad_image` (lines 195-199) to remove padding and inserts `image_newline` tokens for arbitrary aspect ratio handling.
- Flat is the default in [`train.py`](https://github.com/NVlabs/Eagle/blob/main/train.py), while spatial variants excel at grounded visual understanding tasks.

## Frequently Asked Questions

### What is the default patch merge type in Eagle?

The default value is **`'flat'`**, as specified in the training configuration within [`Eagle/train.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/train.py). This setting treats visual patches as a simple sequential token stream compatible with standard transformer language models.

### How does spatial patch merge affect model performance?

Spatial merging enables superior performance on tasks requiring **spatial reasoning**, such as object grounding, referring expression comprehension, and geometric relationship prediction. By preserving the 2D grid structure in [`eagle_arch.py`](https://github.com/NVlabs/Eagle/blob/main/eagle_arch.py), the model can attend to visual features based on their relative positions, though it may require more careful handling of variable image sizes.

### What is the purpose of the `spatial_unpad` variant?

The `spatial_unpad` option removes padding artifacts that are added when resizing images to a fixed square grid. According to the implementation in [`eagle_arch.py`](https://github.com/NVlabs/Eagle/blob/main/eagle_arch.py), it calls `unpad_image` to strip padded borders and inserts a special `image_newline` token to mark row boundaries, making it ideal for processing images with **arbitrary aspect ratios** without introducing padding noise.

### Where is the patch merge logic implemented in the Eagle codebase?

The core logic resides in **[`Eagle/eagle/model/eagle_arch.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/eagle/model/eagle_arch.py)**, specifically between lines 180-198. The `unpad_image` function is defined in the same file around lines 195-199, while [`Eagle/eagle/mm_utils.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/eagle/mm_utils.py) provides the `get_anyres_image_grid_shape` helper for calculating grid dimensions during spatial processing.