Supported LLM Backbones in Eagle and How to Switch Between Them

Eagle supports any HuggingFace causal language model—including LLaMA, OPT, Mistral, and Falcon—and allows seamless switching via the --model_name_or_path or --llm_path command-line arguments without modifying source code.

The NVlabs/Eagle multimodal architecture is backbone-agnostic, loading language models through the standard HuggingFace AutoModelForCausalLM interface. This design means supported LLM backbones in Eagle include any causal LM available on the HuggingFace Hub, with dedicated wrappers provided for the LLaMA family and out-of-the-box compatibility with OPT models.

Supported LLM Backbones

Eagle ships with infrastructure for two primary LLM families, while remaining compatible with any causal language model following the 🤗 Transformers API.

LLaMA Family Models

For LLaMA-1, LLaMA-2, LLaMA-3, and compatible variants, Eagle provides the EagleLlamaForCausalLM wrapper class. This wrapper inherits from LlamaForCausalLM and is defined in Eagle/eagle/model/language_model/eagle_llama.py. When you specify a LLaMA checkpoint, Eagle automatically instantiates this wrapper to handle multimodal projections.

OPT and Generic Causal LMs

By default, Eagle uses the vanilla HuggingFace OPT implementation (facebook/opt-125m) as specified in Eagle/train.py within the ModelArguments dataclass. Because the framework calls AutoModelForCausalLM.from_pretrained, you can substitute any causal LM—including Mistral, Falcon, Gemma, or Qwen2—by providing the appropriate model identifier.

Where Backbone Configuration Is Defined

The LLM backbone selection is exposed through several argument parsers across the repository:

These arguments feed directly into the model instantiation logic:


# Simplified fragment from Eagle/train.py

if model_args.llm_path is not None:
    self.llm = AutoModelForCausalLM.from_pretrained(
        model_args.llm_path,
        trust_remote_code=True,
        **llm_kwargs,
    )

How to Switch the LLM Backbone

Switching requires only changing the model identifier passed to the training script. No source code modification is necessary.

Command-Line Interface

To switch to a different backbone, pass the HuggingFace model ID to the appropriate argument:


# Switch to LLaMA-2-7B in generic training

python Eagle/train.py \
    --model_name_or_path meta-llama/Llama-2-7b-hf \
    --vision_path <vision-ckpt> \
    --freeze_llm False

For "locate-anything" or Eagle 2.5 finetuning scripts, use --llm_path instead:

python Embodied/eaglevl/train/locany_finetune_magi_stream.py \
    --llm_path meta-llama/Llama-2-7b-hf \
    --vision_path <vision-ckpt>

Programmatic Configuration

When instantiating Eagle within Python, set the path fields in the arguments dataclass:

from Eagle.train import ModelArguments, EagleTrainer

args = ModelArguments(
    model_name_or_path="meta-llama/Llama-2-7b-hf",
    llm_path="meta-llama/Llama-2-7b-hf",
    vision_path="google/vit-base-patch16-224",
    freeze_backbone=False,
    use_backbone_lora=0,
)

trainer = EagleTrainer(args)
trainer.train()

Adding LoRA Adapters

To keep the backbone frozen but train a LoRA adapter, specify the rank using --use_llm_lora:

python Eagle/train.py \
    --model_name_or_path meta-llama/Llama-2-7b-hf \
    --use_llm_lora 128 \
    --freeze_llm True

This triggers model.wrap_llm_lora(r=128, ...) within the LLM wrapper, attaching a rank-128 adapter while keeping the base weights frozen.

Summary

  • Eagle supports any HuggingFace causal LM, including dedicated wrappers for LLaMA (EagleLlamaForCausalLM) and default support for OPT.
  • Switch backbones via CLI arguments: Use --model_name_or_path for generic scripts or --llm_path for Eagle 2.5 and embodied training scripts.
  • LoRA integration is available through the --use_llm_lora flag without modifying Eagle/eagle/model/language_model/eagle_llama.py or other source files.
  • Configuration files in Eagle/train.py and Embodied/eaglevl/train/arguments.py expose these options through standard HuggingFace argument parsers.

Frequently Asked Questions

Can I use Mistral or Falcon models with Eagle?

Yes. Because Eagle loads models via AutoModelForCausalLM.from_pretrained, any causal language model on the HuggingFace Hub—including Mistral, Falcon, Gemma, and Qwen2—can serve as the backbone. Simply pass the model identifier to --model_name_or_path or --llm_path as you would with LLaMA or OPT.

What is the default LLM backbone in Eagle?

The default backbone is facebook/opt-125m, defined in the ModelArguments dataclass within Eagle/train.py. This default applies to generic training scripts unless overridden by user-specified arguments.

How do I freeze the LLM backbone during training?

Pass --freeze_llm True (or freeze_backbone=True programmatically) to prevent weight updates in the language model. This is commonly combined with --use_llm_lora to train only the adapter parameters while keeping the backbone frozen.

Where is the LLaMA-specific implementation located?

The LLaMA wrapper class EagleLlamaForCausalLM is implemented in Eagle/eagle/model/language_model/eagle_llama.py. This file inherits from LlamaForCausalLM and handles the multimodal projection layers specific to the LLaMA architecture within the Eagle framework.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →