Supported LLM Backbones in Eagle and How to Switch Between Them
Eagle supports any HuggingFace causal language model—including LLaMA, OPT, Mistral, and Falcon—and allows seamless switching via the --model_name_or_path or --llm_path command-line arguments without modifying source code.
The NVlabs/Eagle multimodal architecture is backbone-agnostic, loading language models through the standard HuggingFace AutoModelForCausalLM interface. This design means supported LLM backbones in Eagle include any causal LM available on the HuggingFace Hub, with dedicated wrappers provided for the LLaMA family and out-of-the-box compatibility with OPT models.
Supported LLM Backbones
Eagle ships with infrastructure for two primary LLM families, while remaining compatible with any causal language model following the 🤗 Transformers API.
LLaMA Family Models
For LLaMA-1, LLaMA-2, LLaMA-3, and compatible variants, Eagle provides the EagleLlamaForCausalLM wrapper class. This wrapper inherits from LlamaForCausalLM and is defined in Eagle/eagle/model/language_model/eagle_llama.py. When you specify a LLaMA checkpoint, Eagle automatically instantiates this wrapper to handle multimodal projections.
OPT and Generic Causal LMs
By default, Eagle uses the vanilla HuggingFace OPT implementation (facebook/opt-125m) as specified in Eagle/train.py within the ModelArguments dataclass. Because the framework calls AutoModelForCausalLM.from_pretrained, you can substitute any causal LM—including Mistral, Falcon, Gemma, or Qwen2—by providing the appropriate model identifier.
Where Backbone Configuration Is Defined
The LLM backbone selection is exposed through several argument parsers across the repository:
Eagle/train.py: DefinesModelArguments.model_name_or_path(default:facebook/opt-125m) for generic training scripts.Embodied/eaglevl/train/arguments.py: ExposesModelArguments.llm_pathfor "locate-anything" finetuning workflows.Eagle2_5/eaglevl/train/eagle_2_5_vl_finetune.py: Usesmodel_args.llm_pathfor the Eagle 2.5 variant training pipeline.
These arguments feed directly into the model instantiation logic:
# Simplified fragment from Eagle/train.py
if model_args.llm_path is not None:
self.llm = AutoModelForCausalLM.from_pretrained(
model_args.llm_path,
trust_remote_code=True,
**llm_kwargs,
)
How to Switch the LLM Backbone
Switching requires only changing the model identifier passed to the training script. No source code modification is necessary.
Command-Line Interface
To switch to a different backbone, pass the HuggingFace model ID to the appropriate argument:
# Switch to LLaMA-2-7B in generic training
python Eagle/train.py \
--model_name_or_path meta-llama/Llama-2-7b-hf \
--vision_path <vision-ckpt> \
--freeze_llm False
For "locate-anything" or Eagle 2.5 finetuning scripts, use --llm_path instead:
python Embodied/eaglevl/train/locany_finetune_magi_stream.py \
--llm_path meta-llama/Llama-2-7b-hf \
--vision_path <vision-ckpt>
Programmatic Configuration
When instantiating Eagle within Python, set the path fields in the arguments dataclass:
from Eagle.train import ModelArguments, EagleTrainer
args = ModelArguments(
model_name_or_path="meta-llama/Llama-2-7b-hf",
llm_path="meta-llama/Llama-2-7b-hf",
vision_path="google/vit-base-patch16-224",
freeze_backbone=False,
use_backbone_lora=0,
)
trainer = EagleTrainer(args)
trainer.train()
Adding LoRA Adapters
To keep the backbone frozen but train a LoRA adapter, specify the rank using --use_llm_lora:
python Eagle/train.py \
--model_name_or_path meta-llama/Llama-2-7b-hf \
--use_llm_lora 128 \
--freeze_llm True
This triggers model.wrap_llm_lora(r=128, ...) within the LLM wrapper, attaching a rank-128 adapter while keeping the base weights frozen.
Summary
- Eagle supports any HuggingFace causal LM, including dedicated wrappers for LLaMA (
EagleLlamaForCausalLM) and default support for OPT. - Switch backbones via CLI arguments: Use
--model_name_or_pathfor generic scripts or--llm_pathfor Eagle 2.5 and embodied training scripts. - LoRA integration is available through the
--use_llm_loraflag without modifyingEagle/eagle/model/language_model/eagle_llama.pyor other source files. - Configuration files in
Eagle/train.pyandEmbodied/eaglevl/train/arguments.pyexpose these options through standard HuggingFace argument parsers.
Frequently Asked Questions
Can I use Mistral or Falcon models with Eagle?
Yes. Because Eagle loads models via AutoModelForCausalLM.from_pretrained, any causal language model on the HuggingFace Hub—including Mistral, Falcon, Gemma, and Qwen2—can serve as the backbone. Simply pass the model identifier to --model_name_or_path or --llm_path as you would with LLaMA or OPT.
What is the default LLM backbone in Eagle?
The default backbone is facebook/opt-125m, defined in the ModelArguments dataclass within Eagle/train.py. This default applies to generic training scripts unless overridden by user-specified arguments.
How do I freeze the LLM backbone during training?
Pass --freeze_llm True (or freeze_backbone=True programmatically) to prevent weight updates in the language model. This is commonly combined with --use_llm_lora to train only the adapter parameters while keeping the backbone frozen.
Where is the LLaMA-specific implementation located?
The LLaMA wrapper class EagleLlamaForCausalLM is implemented in Eagle/eagle/model/language_model/eagle_llama.py. This file inherits from LlamaForCausalLM and handles the multimodal projection layers specific to the LLaMA architecture within the Eagle framework.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →