How Eagle's Mixture-of-Encoders Combines Multiple Vision Encoders

Eagle's mixture-of-encoders strategy aggregates heterogeneous vision backbones by dynamically loading multiple encoders, aligning their spatial resolutions to a common grid, and concatenating features along the channel dimension to create a unified visual representation.

The NVlabs/Eagle open-source repository implements a flexible mixture-of-encoders architecture that enables large multimodal models to leverage complementary visual capabilities from different architectures simultaneously. Instead of relying on a single vision backbone, this mechanism combines detection-oriented, segmentation-focused, and CLIP-style encoders into a single tower. Understanding how Eagle concatenates these diverse representations reveals why the architecture achieves robust performance across varied visual-language tasks.

How the Mixture-of-Encoders Architecture Works

Configuration and Builder Detection

The entry point for creating a multi-encoder vision tower resides in Eagle/eagle/model/multimodal_encoder/builder.py. The function build_vision_tower inspects the vision_tower configuration string, and when it detects a semicolon-separated list of identifiers—for example, "det-1024;convnext-1024;sam-1024"—it returns an instance of MultiBackboneChannelConcatenationVisionTower at line 44 rather than a single encoder model.

Dynamic Encoder Loading

Inside Eagle/eagle/model/multimodal_encoder/multi_backbone_channel_concatenation_encoder.py, the load_vision_towers method parses each semicolon-delimited token and maps it to a concrete encoder class:

  • det-1024 → EVAVITVisionTower (EVA-ViT trained on detection data)
  • convnext-1024 → ConvNextVisionTower (ConvNeXt-XXLarge)
  • sam-1024 → SAMVisionTower (Segment-Anything Model)
  • pix2struct-1024 → Pix2StructLargeVisionTower
  • clip-448 → HRCLIPVisionTower

For each matched encoder, the method creates a deep-copied argument namespace, instantiates the model, and calls load_model(). The tower also records a common input_image_size (hard-coded to 1024 pixels) and stores the image processor from the ConvNeXt encoder to ensure consistent preprocessing across all backbones.

Feature Alignment and Concatenation

During the forward pass, the MultiBackboneChannelConcatenationVisionTower iterates over its loaded encoders and processes inputs through each sequentially. If an encoder expects a different input resolution, the input tensor is resized using F.interpolate.

Each encoder produces features that are checked for shape consistency. Outputs arriving as (B, N, C) are reshaped to (B, C, H, W) if necessary. The code ensures spatial alignment by interpolating all feature maps to the target grid_size of 32 (yielding num_tokens = grid_size ** 2 = 1024). Finally, the tensors are flattened to (B, N, C_i) and concatenated along the channel dimension via torch.cat(features, dim=-1). The resulting hidden_size equals the sum of all individual encoder hidden sizes, producing a fused representation that downstream components treat exactly like single-encoder output.

Implementation Example

The following example demonstrates how to configure and instantiate a mixture-of-encoders vision tower:


# -------------------------------------------------

# 1) Define a mixed-vision tower in the model config

# -------------------------------------------------

# In a yaml / json config file (or programmatically):

vision_cfg = dict(
    vision_tower="det-1024;convnext-1024;sam-1024",  # semi-colon separates encoders

    input_image_size=1024,
    freeze_vision=False,
)

# -------------------------------------------------

# 2) Build the tower via the builder

# -------------------------------------------------

from Eagle.eagle.model.multimodal_encoder.builder import build_vision_tower
vision_tower = build_vision_tower(vision_cfg)

# -------------------------------------------------

# 3) Use it in a forward pass

# -------------------------------------------------

import torch
dummy_img = torch.randn(2, 3, 1024, 1024)   # B=2, C=3, H=W=1024

vision_emb = vision_tower(dummy_img)       # shape: (2, N, C_total)

print(vision_emb.shape)

# Example output: torch.Size([2, 1024, 2304])

#   where 1024 = grid_size**2 and 2304 = sum of hidden sizes

Key Files and Components

Summary

  • Dynamic Loading: The builder detects semicolon-separated encoder identifiers (e.g., "det-1024;convnext-1024") and instantiates a MultiBackboneChannelConcatenationVisionTower to manage heterogeneous backbones.
  • Resolution Alignment: The tower standardizes inputs to 1024 pixels and interpolates encoder outputs to a common grid_size of 32x32 (1024 tokens) using F.interpolate.
  • Channel Concatenation: Features from all encoders are concatenated along the channel dimension (dim=-1), yielding a combined hidden_size equal to the sum of individual encoder dimensions.
  • Unified Interface: The resulting tensor is treated as a single vision embedding by downstream multimodal transformers, requiring no architectural modifications to the language model or projector.

Frequently Asked Questions

What vision encoders does Eagle support in its mixture-of-encoders framework?

Eagle supports several heterogeneous backbones including EVA-ViT for detection (det-1024), ConvNeXt-XXLarge (convnext-1024), SAM (sam-1024), Pix2Struct-Large (pix2struct-1024), and HR-CLIP (clip-448). You can extend support by adding new identity-to-class mappings in the load_vision_towers method of the multi-backbone encoder.

How does the model handle different input resolutions across encoders?

The MultiBackboneChannelConcatenationVisionTower uses F.interpolate to resize inputs when an encoder requires a resolution different from the standard 1024 pixels. During the forward pass, all spatial outputs are interpolated to a common grid size of 32x32 (1024 tokens) before concatenation occurs.

Why concatenate features along the channel dimension instead of averaging them?

Channel-wise concatenation preserves the distinct visual signatures of each encoder—such as detection-oriented features from EVA-ViT and segmentation masks from SAM—allowing the downstream multimodal transformer to learn optimal fusion strategies through its attention mechanisms. Averaging would dilute these complementary cues rather than preserving them.

Do I need to modify the LLM architecture to use multiple vision encoders?

No. The concatenated output presents a unified interface where the channel dimension equals the sum of individual hidden sizes. Downstream components like the multimodal projector receive a single tensor and require no special logic to handle the mixture-of-encoders configuration.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →