How to Configure the Multimodal Projector for Different Vision Encoders in Eagle
Configure Eagle's multimodal projector by setting mm_projector_type to linear, mlp<N>x_gelu, or identity, and ensure mm_hidden_size matches your vision encoder's output dimension (or the sum when using multiple towers).
Eagle, developed by NVlabs, is a multimodal large language model architecture that connects vision encoders to an LLM through a learnable projection layer. When you configure the multimodal projector for different vision encoders in Eagle, you control how visual features are mapped into the language model's hidden space. The projector type and dimensions must align with your specific vision tower configuration to ensure proper feature alignment during training and inference.
Understanding the Multimodal Projector Architecture
The multimodal projector (mm_projector) serves as the bridge between Eagle's vision encoder and the language model, transforming visual feature representations into the hidden-space expected by the LLM.
Where the Projector is Built
The projector instantiation occurs in Eagle/eagle/model/eagle_arch.py, which calls the factory function build_vision_projector defined in Eagle/eagle/model/multimodal_projector/builder.py. The builder creates the specific projection module based on the mm_projector_type configuration field.
# From builder.py - factory logic
if projector_type == 'linear':
return nn.Linear(mm_hidden_size, hidden_size)
elif 'mlp' in projector_type:
# Parses mlp<N>x_gelu pattern
mlp_gelu_match = re.match(r'^mlp(\d+)x_gelu$', projector_type)
# ... creates stacked Linear+GELU layers
elif projector_type == 'identity':
return IdentityMap()
else:
raise ValueError(f'Unknown projector type: {projector_type}')
Available Projector Types
Eagle supports three distinct projector architectures, selected via the mm_projector_type parameter:
linear– A singlenn.Linear(mm_hidden_size, hidden_size)layer. This is the default option and works well for standard vision-language alignment tasks.mlp<N>x_gelu– A multi-layer perceptron with N stacked linear-GELU layers (where N ≥ 1). Use this when the vision encoder outputs require deeper non-linear transformation before entering the LLM.identity– AnIdentityMap()module that passes features through unchanged. Select this when the vision encoder already outputs features in the same dimension as the LLM's hidden size, such as after concatenating multiple compatible towers.
If an unrecognized string is passed to mm_projector_type, the builder raises a ValueError to prevent silent misconfiguration.
Configuring for Single and Multiple Vision Encoders
The projector's input dimension (mm_hidden_size) depends on whether you use a single vision tower or multiple concatenated encoders.
Single Vision Tower Configuration
When using one vision encoder, set mm_hidden_size equal to the encoder's hidden size (vision_tower.hidden_size). For example:
- Use
768for ViT-B/16 models - Use
1024for SigLIP-Base models
The encoder returns a tensor of shape (B, C) where C is the encoder's hidden dimension, and the projector expects this exact size as input.
Multiple Vision Towers and Concatenation
For multi-encoder setups, Eagle uses MultiBackboneChannelConcatenationEncoder defined in Eagle/eagle/model/multimodal_encoder/multi_backbone_channel_concatenation_encoder.py. This module concatenates outputs from each vision tower along the channel dimension and automatically sets mm_hidden_size to the sum of all tower hidden sizes.
# From multi_backbone_channel_concatenation_encoder.py
# mm_hidden_size = sum of individual tower hidden sizes
When configuring multiple towers via the command line, separate encoder names with semicolons (e.g., siglip_base_patch16;convnext_xxlarge). The concatenation encoder handles the dimension calculation, but you can override mm_hidden_size manually if implementing custom tower combinations.
Setting Configuration Parameters
You can configure the multimodal projector through command-line arguments when launching training or programmatically when building models manually.
Command-Line Configuration
The training script Eagle/train.py defines the relevant arguments in the ModelArguments dataclass:
@dataclass
class ModelArguments:
mm_projector_type: Optional[str] = field(default='linear')
vision_tower: Optional[str] = field(default=None)
# mm_hidden_size is inferred from vision_tower or set manually
Launch training with your desired projector configuration:
# Single tower with linear projector
torchrun --nproc_per_node=8 \
Eagle/train.py \
--model_name_or_path llama2-7b \
--vision_tower siglip_base_patch16 \
--mm_projector_type linear \
--output_dir ./output_linear
# Single tower with 2-layer MLP projector
torchrun --nproc_per_node=8 \
Eagle/train.py \
--model_name_or_path llama2-7b \
--vision_tower siglip_base_patch16 \
--mm_projector_type mlp2x_gelu \
--output_dir ./output_mlp
# Identity projector when dimensions already match
torchrun --nproc_per_node=2 \
Eagle/train.py \
--model_name_or_path llama2-7b \
--vision_tower custom_encoder_with_4096_dim \
--mm_projector_type identity \
--output_dir ./output_identity
Programmatic Configuration
When building an Eagle model directly in Python, set the configuration attributes before instantiation:
from Eagle.eagle.model.eagle_arch import EagleModel
cfg = EagleModel.get_default_config()
cfg.vision_tower = "siglip_base_patch16"
cfg.mm_projector_type = "mlp3x_gelu"
cfg.mm_hidden_size = 1024 # Output dim of vision tower
cfg.hidden_size = 4096 # LLM hidden size
model = EagleModel(cfg) # Projector built automatically during __init__
For multiple towers, specify semicolon-separated values:
cfg.vision_tower = "siglip_base_patch16;convnext_xxlarge"
cfg.mm_projector_type = "mlp2x_gelu"
# mm_hidden_size automatically becomes 1024 + 3072 = 4096
Optimizer Considerations for Projector Training
In Eagle/eagle/train/eagle_trainer.py, the optimizer treats projector parameters separately from the vision tower and LLM weights. This allows you to specify a dedicated learning rate for the projection layer using the mm_projector_lr argument.
# From eagle_trainer.py - parameter group separation
# Projector weights use mm_projector_lr if specified,
# otherwise default learning rate
This separation enables fine-tuning the visual mapping independently from the frozen or lightly-tuned vision encoder and LLM components.
Summary
- Select projector type using
mm_projector_type: chooselinearfor simple mapping,mlp<N>x_gelufor deep non-linear projection, oridentitywhen dimensions already align. - Match dimensions by setting
mm_hidden_sizeto the vision encoder's output dimension for single towers, or the sum of dimensions for multiple concatenated towers. - Configure via CLI using arguments in
Eagle/train.py, or programmatically by setting attributes on the Eagle configuration object before model instantiation. - Optimize separately using
mm_projector_lrin the trainer to apply distinct learning rates to projection parameters versus the LLM and vision tower.
Frequently Asked Questions
What happens if I specify an invalid projector type?
The build_vision_projector factory in Eagle/eagle/model/multimodal_projector/builder.py raises a ValueError with the message "Unknown projector type" if you provide a string that does not match linear, identity, or the mlp<N>x_gelu pattern. This prevents the model from initializing with an undefined projection layer.
How do I calculate mm_hidden_size for multiple vision encoders?
When using multiple vision towers separated by semicolons in the vision_tower argument, the MultiBackboneChannelConcatenationEncoder automatically calculates mm_hidden_size as the sum of each individual encoder's hidden dimension. For manual configuration, add the hidden sizes of all towers you plan to concatenate (e.g., SigLIP-Base at 1024 plus ConvNeXt-XXLarge at 3072 equals 4096).
Can I use different learning rates for the projector and the LLM?
Yes. The EagleTrainer class in Eagle/eagle/train/eagle_trainer.py separates projector parameters into a distinct optimizer group. Pass the mm_projector_lr command-line argument to apply a specific learning rate to the projection layers while using the default rate for other components.
When should I use the identity projector type?
Use identity when your vision encoder already outputs features in the same dimension as the LLM's hidden_size, making linear transformation unnecessary. This commonly occurs when concatenating multiple towers whose combined dimension matches the LLM embedding size, or when using a custom vision encoder specifically designed to match the language model's dimensions.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →