# How Eagle's Mixture-of-Encoders Combines Multiple Vision Encoders

> Learn how Eagle's mixture-of-encoders aggregates heterogeneous vision backbones. This strategy dynamically loads multiple encoders, aligns spatial resolutions, and concatenates features for a unified visual representation.

- Repository: [NVIDIA Research Projects/Eagle](https://github.com/NVlabs/Eagle)
- Tags: internals
- Published: 2026-06-28

---

**Eagle's mixture-of-encoders strategy aggregates heterogeneous vision backbones by dynamically loading multiple encoders, aligning their spatial resolutions to a common grid, and concatenating features along the channel dimension to create a unified visual representation.**

The NVlabs/Eagle open-source repository implements a flexible **mixture-of-encoders** architecture that enables large multimodal models to leverage complementary visual capabilities from different architectures simultaneously. Instead of relying on a single vision backbone, this mechanism combines detection-oriented, segmentation-focused, and CLIP-style encoders into a single tower. Understanding how Eagle concatenates these diverse representations reveals why the architecture achieves robust performance across varied visual-language tasks.

## How the Mixture-of-Encoders Architecture Works

### Configuration and Builder Detection

The entry point for creating a multi-encoder vision tower resides in [`Eagle/eagle/model/multimodal_encoder/builder.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/eagle/model/multimodal_encoder/builder.py). The function `build_vision_tower` inspects the `vision_tower` configuration string, and when it detects a semicolon-separated list of identifiers—for example, `"det-1024;convnext-1024;sam-1024"`—it returns an instance of `MultiBackboneChannelConcatenationVisionTower` at line 44 rather than a single encoder model.

### Dynamic Encoder Loading

Inside [`Eagle/eagle/model/multimodal_encoder/multi_backbone_channel_concatenation_encoder.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/eagle/model/multimodal_encoder/multi_backbone_channel_concatenation_encoder.py), the `load_vision_towers` method parses each semicolon-delimited token and maps it to a concrete encoder class:

- `det-1024` → `EVAVITVisionTower` (EVA-ViT trained on detection data)
- `convnext-1024` → `ConvNextVisionTower` (ConvNeXt-XXLarge)
- `sam-1024` → `SAMVisionTower` (Segment-Anything Model)
- `pix2struct-1024` → `Pix2StructLargeVisionTower`
- `clip-448` → `HRCLIPVisionTower`

For each matched encoder, the method creates a deep-copied argument namespace, instantiates the model, and calls `load_model()`. The tower also records a common `input_image_size` (hard-coded to 1024 pixels) and stores the image processor from the ConvNeXt encoder to ensure consistent preprocessing across all backbones.

### Feature Alignment and Concatenation

During the forward pass, the `MultiBackboneChannelConcatenationVisionTower` iterates over its loaded encoders and processes inputs through each sequentially. If an encoder expects a different input resolution, the input tensor is resized using `F.interpolate`.

Each encoder produces features that are checked for shape consistency. Outputs arriving as `(B, N, C)` are reshaped to `(B, C, H, W)` if necessary. The code ensures spatial alignment by interpolating all feature maps to the target `grid_size` of 32 (yielding `num_tokens = grid_size ** 2 = 1024`). Finally, the tensors are flattened to `(B, N, C_i)` and concatenated along the channel dimension via `torch.cat(features, dim=-1)`. The resulting `hidden_size` equals the sum of all individual encoder hidden sizes, producing a fused representation that downstream components treat exactly like single-encoder output.

## Implementation Example

The following example demonstrates how to configure and instantiate a mixture-of-encoders vision tower:

```python

# -------------------------------------------------

# 1) Define a mixed-vision tower in the model config

# -------------------------------------------------

# In a yaml / json config file (or programmatically):

vision_cfg = dict(
    vision_tower="det-1024;convnext-1024;sam-1024",  # semi-colon separates encoders

    input_image_size=1024,
    freeze_vision=False,
)

# -------------------------------------------------

# 2) Build the tower via the builder

# -------------------------------------------------

from Eagle.eagle.model.multimodal_encoder.builder import build_vision_tower
vision_tower = build_vision_tower(vision_cfg)

# -------------------------------------------------

# 3) Use it in a forward pass

# -------------------------------------------------

import torch
dummy_img = torch.randn(2, 3, 1024, 1024)   # B=2, C=3, H=W=1024

vision_emb = vision_tower(dummy_img)       # shape: (2, N, C_total)

print(vision_emb.shape)

# Example output: torch.Size([2, 1024, 2304])

#   where 1024 = grid_size**2 and 2304 = sum of hidden sizes

```

## Key Files and Components

- **[`Eagle/eagle/model/multimodal_encoder/multi_backbone_channel_concatenation_encoder.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/eagle/model/multimodal_encoder/multi_backbone_channel_concatenation_encoder.py)** — Implements the `MultiBackboneChannelConcatenationVisionTower` class that loads, aligns, and concatenates multiple vision encoders.
- **[`Eagle/eagle/model/multimodal_encoder/builder.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/eagle/model/multimodal_encoder/builder.py)** — Contains `build_vision_tower` which detects semicolon-separated encoder lists and routes to the multi-backbone implementation.
- **[`Eagle2_5/eaglevl/model/multimodal_encoder/multi_backbone_channel_concatentation_model.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle2_5/eaglevl/model/multimodal_encoder/multi_backbone_channel_concatentation_model.py)** — Demonstrates how the mixed vision tower embeds inside the full multimodal model.
- **[`Eagle2_5/eaglevl/model/multimodal_encoder/configuration_multi_backbone_channel_concatentation_model.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle2_5/eaglevl/model/multimodal_encoder/configuration_multi_backbone_channel_concatentation_model.py)** — Defines the configuration schema for mixture-of-encoders models.

## Summary

- **Dynamic Loading:** The builder detects semicolon-separated encoder identifiers (e.g., `"det-1024;convnext-1024"`) and instantiates a `MultiBackboneChannelConcatenationVisionTower` to manage heterogeneous backbones.
- **Resolution Alignment:** The tower standardizes inputs to 1024 pixels and interpolates encoder outputs to a common `grid_size` of 32x32 (1024 tokens) using `F.interpolate`.
- **Channel Concatenation:** Features from all encoders are concatenated along the channel dimension (`dim=-1`), yielding a combined `hidden_size` equal to the sum of individual encoder dimensions.
- **Unified Interface:** The resulting tensor is treated as a single vision embedding by downstream multimodal transformers, requiring no architectural modifications to the language model or projector.

## Frequently Asked Questions

### What vision encoders does Eagle support in its mixture-of-encoders framework?

Eagle supports several heterogeneous backbones including EVA-ViT for detection (`det-1024`), ConvNeXt-XXLarge (`convnext-1024`), SAM (`sam-1024`), Pix2Struct-Large (`pix2struct-1024`), and HR-CLIP (`clip-448`). You can extend support by adding new identity-to-class mappings in the `load_vision_towers` method of the multi-backbone encoder.

### How does the model handle different input resolutions across encoders?

The `MultiBackboneChannelConcatenationVisionTower` uses `F.interpolate` to resize inputs when an encoder requires a resolution different from the standard 1024 pixels. During the forward pass, all spatial outputs are interpolated to a common grid size of 32x32 (1024 tokens) before concatenation occurs.

### Why concatenate features along the channel dimension instead of averaging them?

Channel-wise concatenation preserves the distinct visual signatures of each encoder—such as detection-oriented features from EVA-ViT and segmentation masks from SAM—allowing the downstream multimodal transformer to learn optimal fusion strategies through its attention mechanisms. Averaging would dilute these complementary cues rather than preserving them.

### Do I need to modify the LLM architecture to use multiple vision encoders?

No. The concatenated output presents a unified interface where the channel dimension equals the sum of individual hidden sizes. Downstream components like the multimodal projector receive a single tensor and require no special logic to handle the mixture-of-encoders configuration.