# How Eagle's Vision-Language Projector Handles Feature Fusion from Multiple Encoders

> Discover how Eagle fuses features from multiple vision encoders. Learn about its configurable linear layer or MLP projection into the LLM's hidden space.

- Repository: [NVIDIA Research Projects/Eagle](https://github.com/NVlabs/Eagle)
- Tags: internals
- Published: 2026-06-28

---

**Eagle concatenates channel dimensions from multiple vision encoders and projects the combined vector into the LLM's hidden space using a configurable linear layer or MLP.**

The vision-language projector in NVlabs/Eagle serves as the critical bridge between diverse visual encoders and the language model. When working with multiple vision towers—such as combining ConvNeXt and SigLIP backbones—the projector must fuse heterogeneous feature representations into a unified embedding space. Understanding this fusion mechanism reveals how Eagle achieves flexible multi-modal integration while maintaining a clean interface to the underlying LLM.

## Multi-Encoder Feature Concatenation

Eagle handles multiple vision encoders through channel-wise concatenation before projection. When configured with multiple vision towers, each encoder emits a feature tensor of shape `N × C_i`, where `C_i` represents the channel dimension of the *i-th* encoder.

The `MultiBackboneChannelConcatenationEncoder` stacks these tensors along the channel dimension, producing a single tensor of shape `N × Σ C_i`. This concatenated representation preserves information from all visual streams while creating a unified input for the projector. The sum of channel dimensions is stored in `fpn_input_dim` and passed to the projector builder to ensure the input dimension matches the combined visual feature size.

## Configurable Projector Architecture

The projector construction logic resides in [`Eagle/eagle/model/multimodal_projector/builder.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/eagle/model/multimodal_projector/builder.py) (lines 42-55). The architecture adapts based on the `mm_projector_type` configuration parameter:

- **`linear`** (default): A single `nn.Linear` layer mapping from `config.mm_hidden_size` (the summed visual dimension) to the LLM's hidden size
- **`mlp2x_gelu`**: A two-layer MLP with `Linear → GELU → Linear` that provides non-linear capacity for more complex fusion patterns
- **`identity`**: A pass-through layer that forwards the concatenated visual tensor unchanged

Regardless of the specific type, the projector's input dimension always equals `fpn_input_dim` (the total channel count Σ C_i), ensuring compatibility with the concatenated multi-encoder output.

## Fusion Mechanism in the Forward Pass

Feature fusion occurs implicitly within the projector's learned transformation. In [`Eagle/eagle/model/eagle_arch.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/eagle/model/eagle_arch.py) (line 60), the forward pass executes `self.mm_projector(image_features)`, where `image_features` represents the already-concatenated visual tensor.

Because the projector receives the full concatenated vector of dimension Σ C_i, the linear or MLP weights learn optimal combinations across all encoder channels simultaneously. This design treats fusion as a projection problem: the network learns to compress the high-dimensional visual representation into the LLM's hidden space through trained weights that inherently mix features from all input encoders.

## Training and Checkpoint Loading

The projector supports independent pretraining through checkpoint loading. When a pretrained adapter path is specified via `pretrain_mm_mlp_adapter`, the system loads stored weights into the projector using `self.mm_projector.load_state_dict(...)` in [`Eagle/eagle/model/eagle_arch.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/eagle/model/eagle_arch.py) (lines 110-116).

This separation allows the fusion mapping to be trained independently from the language model, then frozen or fine-tuned during full model training. The checkpoint contains the learned projection weights that encode the optimal fusion strategy for the specific combination of vision encoders.

## Implementation Example

The following workflow demonstrates how the vision tower and projector collaborate to fuse multi-encoder features:

```python

# Build the vision tower(s) - may return a list of towers or concatenated encoder

from eagle.model.multimodal_encoder.builder import build_vision_tower
vision_tower = build_vision_tower(config, delay_load=True)

# Determine total visual channel size (e.g., [768, 1536] -> 2304)

fpn_input_dim = [] if not hasattr(vision_tower, "fpn_input_dim") \
                 else vision_tower.fpn_input_dim

# Build the projector that will fuse concatenated features

from eagle.model.multimodal_projector.builder import build_vision_projector
mm_projector = build_vision_projector(
    config,                     # contains mm_hidden_size & hidden_size

    fpn_input_dim=fpn_input_dim # summed channel dimension

)

# Forward pass: concatenated features → projector → LLM hidden space

visual_feats = vision_tower(images)     # shape: [B, Σ C_i]

projected = mm_projector(visual_feats) # shape: [B, hidden_size]

```

The projector automatically adapts to the total visual dimension. Adding or removing encoders only requires updating `fpn_input_dim`, leaving the projection pipeline architecture unchanged.

## Key Source Files

Understanding the complete fusion pipeline requires examining three critical components:

- **[`Eagle/eagle/model/eagle_arch.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/eagle/model/eagle_arch.py)**: Instantiates vision towers and executes the forward projection via `self.mm_projector(image_features)`
- **[`Eagle/eagle/model/multimodal_projector/builder.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/eagle/model/multimodal_projector/builder.py)**: Constructs the projector architecture based on `mm_projector_type`
- **[`Eagle2_5/eaglevl/model/multimodal_encoder/multi_backbone_channel_concatenation_encoder.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle2_5/eaglevl/model/multimodal_encoder/multi_backbone_channel_concatenation_encoder.py)**: Performs channel-wise concatenation of multiple encoder outputs

## Summary

Eagle's vision-language projector achieves multi-encoder feature fusion through a streamlined two-stage process:

- **Concatenation**: Channel dimensions from multiple vision towers are stacked before reaching the projector
- **Learned Projection**: A configurable linear layer or MLP learns to map the combined high-dimensional visual vector into the LLM's hidden space
- **Modular Design**: The `mm_projector_type` parameter controls fusion complexity without requiring architectural changes to the encoder or LLM components
- **Independent Training**: Pretrained projector checkpoints can be loaded separately, allowing the fusion weights to be optimized before full model integration

## Frequently Asked Questions

### How does Eagle combine features from different vision encoders with varying output dimensions?

Eagle concatenates features along the channel dimension using `MultiBackboneChannelConcatenationEncoder`. Each encoder's output of shape `N × C_i` is stacked to create a tensor of shape `N × Σ C_i`. The projector then learns to compress this combined representation into the LLM's hidden dimension through its linear or MLP weights.

### What projector types are available for feature fusion in Eagle?

Eagle supports three projector configurations controlled by `mm_projector_type`: `linear` (default single-layer projection), `mlp2x_gelu` (two-layer MLP with GELU activation for non-linear fusion), and `identity` (direct passthrough without transformation). All types accept the concatenated multi-encoder features as input.

### Can I use a pretrained projector checkpoint with different vision encoders?

You can load a pretrained projector using `pretrain_mm_mlp_adapter` in [`eagle_arch.py`](https://github.com/NVlabs/Eagle/blob/main/eagle_arch.py) (lines 110-116), but the checkpoint must match the target encoder configuration. Since the projector's input dimension equals the sum of all encoder channel dimensions (`fpn_input_dim`), changing the encoder set requires retraining or careful dimension matching in the projector weights.

### Where does the actual feature fusion occur in the code?

The fusion happens implicitly in [`Eagle/eagle/model/eagle_arch.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/eagle/model/eagle_arch.py) at line 60 when `self.mm_projector(image_features)` processes the concatenated tensor. The learning occurs in the projector's weights, which optimally combine channels from all encoders while projecting into the LLM's embedding space.