How to Fine-Tune Eagle Models with LoRA for Efficient Training
You can fine-tune Eagle vision-language models efficiently by setting training_args.lora_enable=True and configuring adapter ranks via use_llm_lora and use_backbone_lora, which wraps the base model with PEFT LoRA adapters while keeping original weights frozen to minimize GPU memory usage.
Eagle is a family of vision-language models developed by NVIDIA that supports parameter-efficient fine-tuning through LoRA (Low-Rank Adaptation). The implementation integrates directly with the Hugging Face PEFT library and is exposed through high-level training arguments. This guide covers the exact source code paths, configuration parameters, and runnable examples for fine-tuning Eagle models with LoRA.
LoRA Integration in the Training Pipeline
The LoRA implementation lives in the core training loop and is activated through specific configuration flags.
Training Entry Point (Eagle/train.py)
The primary training script detects LoRA mode via training_args.lora_enable. When enabled, it constructs a peft.LoraConfig and wraps the base model using peft.get_peft_model according to lines 1060–1078 in Eagle/train.py. This process injects trainable low-rank matrices into targeted linear layers while freezing the original pre-trained weights.
The configuration automatically identifies target modules through the find_all_linear_names helper (lines 1089–1100), which selects linear layers while skipping vision-specific modules that should remain frozen during language-model fine-tuning.
Model Arguments for Adapter Configuration
The ModelArguments dataclass in Embodied/eaglevl/train/arguments.py exposes two critical fields for controlling adapter placement:
use_backbone_lora(default:0) – Sets the rank for the vision backbone adapter (lines 53–57).use_llm_lora(default:0) – Sets the rank for the language model head adapter (lines 57–60).
When either value is non-zero, the trainer passes these ranks to the LoRA configuration builder, allowing independent control over which modality receives trainable adapters.
LoRA Hyperparameters
The LoraConfig receives the following parameters during initialization:
r– Adapter rank (default:64), determining the dimensionality of the low-rank decomposition.lora_alpha– Scaling factor (default:16) applied to the adapter outputs.lora_dropout– Dropout probability (default:0.05) applied before the low-rank projection.target_modules– List of linear layer names selected by the automatic finder.bias– Bias training strategy (none,all, orlora_only).
End-to-End Fine-Tuning Workflow
The training process follows these sequential steps:
-
Environment Setup – Set
USE_LLM_LORAand/orUSE_BACKBONE_LORAenvironment variables to specify ranks (e.g.,64for a 64-dimensional adapter), and ensurelora_enable=Truein the training arguments. -
Argument Parsing – The training script
locany_finetune_magi_stream.pyparses CLI arguments intoModelArguments, forwarding the LoRA ranks to the core trainer. -
Adapter Injection –
Eagle/train.pydetectslora_enable, builds theLoraConfigwith the supplied ranks, and callsget_peft_modelto insert adapters. -
Frozen Base Weights – The PEFT runtime freezes the original model parameters, leaving only the low-rank matrices trainable.
-
Training Loop – The Hugging Face Trainer updates only LoRA parameters, keeping the base model in FP16/BF16 without gradients, significantly reducing VRAM requirements.
Practical Code Examples
Fine-Tuning with the Provided Shell Script
The repository includes a ready-to-run script for visual-prompt fine-tuning:
export HF_TOKEN=YOUR_HF_TOKEN
export META_PATH=path/to/your/meta.json
bash Embodied/shell/locate-anything-lora-visual-prompt.sh \
1 \
work_dirs/locany_lora_prompt
This script configures LoRA through environment variables (lines 49–80):
USE_LLM_LORA=64– Attaches a 64-dim adapter to the language model.USE_BACKBONE_LORA=0– Keeps the vision backbone frozen (set to a positive integer to enable vision fine-tuning).FREEZE_LLM=TrueandFREEZE_BACKBONE=True– Ensures original weights remain frozen while only adapters update.
Manual Python Invocation
For programmatic control, instantiate the arguments directly:
from transformers import HfArgumentParser
from eaglevl.train.arguments import ModelArguments, DataTrainingArguments
from Eagle.train import main as train_main
parser = HfArgumentParser((ModelArguments, DataTrainingArguments))
model_args, data_args = parser.parse_args_into_dataclasses()
# Configure LoRA ranks
model_args.use_llm_lora = 64 # LLM adapter rank
model_args.use_backbone_lora = 32 # Vision backbone adapter rank
model_args.freeze_llm = True
model_args.freeze_backbone = True
# Training configuration
training_args = {
"lora_enable": True,
"lora_r": model_args.use_llm_lora,
"lora_alpha": 16,
"lora_dropout": 0.05,
"lora_bias": "none",
"bits": 16,
"bf16": True,
# ... additional Trainer arguments
}
train_main(model_args, data_args, training_args)
The train_main function handles the PEFT integration, creating the LoraConfig and applying it via get_peft_model before starting the training loop.
Inspecting and Exporting LoRA Weights
After training, extract only the adapter weights for efficient storage or distribution:
from peft import PeftModel
import torch
model = PeftModel.from_pretrained("work_dirs/locany_lora_prompt/checkpoint-500")
lora_state = model.state_dict()
# Save only LoRA parameters
torch.save(
{k: v for k, v in lora_state.items() if "lora_" in k},
"locany_lora_adapters.pt"
)
To create a single-file deployment model, merge the adapters back into the base checkpoint using peft.merge_and_save.
Key Source Files and Implementation Details
Understanding these files helps with custom modifications:
Eagle/train.py– Core training loop containing the LoRA detection logic (lines 1060–1078) andLoraConfigconstruction.Embodied/eaglevl/train/arguments.py– Definesuse_backbone_loraanduse_llm_lorafields in theModelArgumentsdataclass.Embodied/shell/locate-anything-lora-visual-prompt.sh– Reference implementation setting LoRA environment variables and launching distributed training.Embodied/eaglevl/train/locany_finetune_magi_stream.py– Entry point that parses LoRA-specific CLI arguments and forwards them to the trainer.Embodied/eaglevl/model/locany/modeling_locateanything.py– Model architecture compatible with PEFT-injected modules.
Summary
- Enable LoRA by setting
training_args.lora_enable=Truein the training configuration. - Control adapter placement separately for the vision backbone (
use_backbone_lora) and language model (use_llm_lora) throughModelArguments. - Default hyperparameters use rank
64, alpha16, and dropout0.05, targeting only linear layers while skipping vision-specific modules. - The PEFT integration freezes base weights automatically, reducing memory consumption to adapter parameters only (approximately
2 × r × hidden_dimper layer). - Use the provided shell scripts for quick experiments or the Python API for custom training pipelines.
Frequently Asked Questions
What are the default LoRA hyperparameters in Eagle?
According to the source code in Eagle/train.py, the default configuration uses a rank (r) of 64, a scaling factor (lora_alpha) of 16, and a dropout rate of 0.05. The bias parameter defaults to "none", meaning bias terms are not trained unless explicitly configured otherwise.
Can I fine-tune only the vision backbone or only the LLM?
Yes. The ModelArguments dataclass provides independent control through use_backbone_lora and use_llm_lora. Set the respective field to a positive integer (the desired rank) to enable adapters for that component, or leave it at 0 to keep that modality frozen. This flexibility allows visual prompt tuning without modifying the language model, or vice versa.
How much memory does LoRA save compared to full fine-tuning?
LoRA adapters add only approximately 2 × r × hidden_dim parameters per target layer. For a typical configuration with rank 64, this amounts to a few megabytes per layer rather than gigabytes for full weight matrices. Since the base model weights remain frozen in FP16 or BF16 without gradients, VRAM usage drops significantly, often allowing fine-tuning on consumer GPUs that could only run inference with full fine-tuning.
Where are the LoRA adapters saved after training?
Adapter weights are saved within the standard Hugging Face checkpoint directories (e.g., work_dirs/locany_lora_prompt/checkpoint-500). To extract only the LoRA parameters, filter the state dictionary for keys containing "lora_" as shown in the Python example above. You can reload these adapters using PeftModel.from_pretrained() or merge them permanently into the base model using PEFT's merge utilities.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →