How to Evaluate Eagle on VLM Benchmarks Using lmms-eval: A Complete Guide
Eagle registers a model wrapper named eagle in the lmms-eval harness that loads pretrained checkpoints, prepares the vision tower, and runs batched inference across vision-language tasks through the simple_evaluate loop in Eagle/lmms_eval/evaluator.py.
The NVlabs/Eagle repository provides a native integration with the lmms-eval framework, allowing you to evaluate Eagle vision-language models on standard benchmarks like VQAv2, MM-Bench, and MME. This integration leverages a custom model wrapper that conforms to the lmms abstract class requirements while handling Eagle-specific tokenization and image preprocessing. Understanding how to evaluate Eagle on VLM benchmarks using lmms-eval enables reproducible benchmarking and fair comparison against other vision-language models.
Understanding the lmms-eval Integration Architecture
Model Registration and Loading
The evaluation harness recognizes Eagle through the @register_model("eagle") decorator in Eagle/lmms_eval/models/eagle.py. When you specify --model eagle on the command line, the wrapper invokes load_pretrained_model from Eagle/eagle/model/builder.py to initialize the full inference stack.
This builder function performs three critical operations:
- Instantiates the tokenizer, language model, and optional LoRA weights.
- Adds special image-patch, start, and end tokens to the tokenizer vocabulary.
- Loads the vision tower (CLIP-style image encoder) and attaches its image processor.
The wrapper then returns a model instance that satisfies the lmms-eval abstract class interface, enabling the harness to treat Eagle like any other supported vision-language model.
Task Initialization and Data Pipeline
Before inference begins, initialize_tasks (defined in Eagle/lmms_eval/__init__.py) registers every available VLM benchmark under the ALL_TASKS registry. Each task definition in Eagle/lmms_eval/tasks/<benchmark>/ supplies:
- A
data_pathpointing to the dataset (downloaded automatically by lmms-eval). - A prompt template that combines visual tokens with language model input.
- Metric calculation functions (e.g., accuracy for VQAv2, BLEU for captioning).
Step-by-Step Evaluation Workflow
The evaluation process follows three distinct phases orchestrated by evaluator.simple_evaluate in Eagle/lmms_eval/evaluator.py:
-
Preprocessing: The harness loads images and calls
process_imagesfromEagle/eagle/mm_utils.pyto convert visual inputs into tensors compatible with the vision tower. -
Encoding and Inference: The tokenizer encodes the prompt with appended special image tokens, padding to the model's
max_length. The model runs batched inference according to thebatch_sizeparameter on the specified device (CPU or GPU). -
Post-processing and Scoring: Generated answers undergo normalization (case conversion, punctuation removal) before the task-specific metric functions compute final scores.
Running Evaluations: Code Examples
All commands assume execution from the repository root. The entry point is Eagle/evaluate_lmms_eval.py, which parses arguments and delegates to the evaluation loop.
Single Benchmark Evaluation (VQAv2)
Run a quick evaluation on VQAv2 to verify your setup:
python -m Eagle.evaluate_lmms_eval \
--model eagle \
--model_args "pretrained=NVEagle/Eagle-X5-7B,device=cuda" \
--tasks vqav2 \
--batch_size 4 \
--output_path ./eval_results \
--log_samples
The --model eagle flag selects the wrapper defined in Eagle/lmms_eval/models/eagle.py. The pretrained argument accepts any HuggingFace hub checkpoint or local path.
Multi-Benchmark Suite Evaluation
Evaluate across multiple benchmarks simultaneously by comma-separating task names:
python -m Eagle.evaluate_lmms_eval \
--model eagle \
--model_args "pretrained=NVEagle/Eagle-X5-7B,device=cuda" \
--tasks vqav2,mmbench,mmmu,mme \
--batch_size 2 \
--output_path ./full_eval \
--log_samples \
--wandb_log_samples \
--wandb_args "project=EagleVLM,entity=your_wandb_user"
The harness supports wildcard expansion (*) for task selection. Weights & Biases integration is optional and configured via --wandb_args.
YAML Configuration for Reproducibility
For reproducible experiments, store parameters in a YAML file:
model: eagle
model_args: "pretrained=NVEagle/Eagle-X5-7B,device=cuda"
tasks: vqav2,mmbench,mmmu,mme
batch_size: 2
log_samples: true
output_path: ./config_eval
Execute with:
python -m Eagle.evaluate_lmms_eval --config config.yaml
The script parses the YAML (lines 199-205 in Eagle/evaluate_lmms_eval.py) and constructs an argparse.Namespace for each configuration entry.
Key Source Files and Their Roles
Understanding the source structure helps when debugging or extending the evaluation pipeline:
Eagle/lmms_eval/models/eagle.py: Registers the Eagle model wrapper and implements the interface required by the lmms-eval harness.Eagle/eagle/model/builder.py: Loads pretrained checkpoints, adds image tokens to the tokenizer, and initializes the vision tower.Eagle/evaluate_lmms_eval.py: CLI entry point that parses arguments (including YAML configs) and invokes the evaluation loop.Eagle/lmms_eval/evaluator.py: Containssimple_evaluate, the core evaluation loop that runs inference and aggregates metrics across tasks.Eagle/lmms_eval/tasks/: Directory containing individual benchmark definitions (e.g.,vqav2,mmbench,mme), each with data loading utilities and metric functions.Eagle/eagle/mm_utils.py: Utility functions for inserting image tokens into prompts and handling image processor transformations.
Summary
- Eagle integrates natively with lmms-eval through a registered model wrapper in
Eagle/lmms_eval/models/eagle.pythat handles checkpoint loading and vision tower initialization. - The evaluation workflow involves model registration, task initialization, and the
simple_evaluateinference loop that preprocesses images, encodes prompts, and computes task-specific metrics. - You can run evaluations via command-line arguments or YAML configuration files, supporting both single benchmarks and multi-benchmark suites with optional Weights & Biases logging.
- Key files include the model builder (
builder.py), evaluation entry point (evaluate_lmms_eval.py), and task definitions underlmms_eval/tasks/.
Frequently Asked Questions
What checkpoints are compatible with the Eagle lmms-eval wrapper?
Any Eagle checkpoint hosted on Hugging Face Hub or stored locally works with the wrapper. Pass the checkpoint identifier via the pretrained argument in --model_args, such as pretrained=NVEagle/Eagle-X5-7B or pretrained=/path/to/local/checkpoint. The builder automatically detects and loads associated vision tower weights.
How do I add a custom VLM benchmark to the evaluation harness?
Create a new directory under Eagle/lmms_eval/tasks/ containing a task definition file that specifies the dataset path, prompt template, and metric functions. The initialize_tasks function in Eagle/lmms_eval/__init__.py automatically discovers new task modules, making them available via the --tasks flag without modifying core evaluation code.
Can I run evaluation on CPU or multiple GPUs?
Yes. Specify device=cpu in --model_args for CPU inference, or device=cuda for GPU acceleration. The lmms-eval harness handles device placement through the model wrapper. For multi-GPU evaluation, the harness supports data parallelism, though you should verify that your specific Eagle checkpoint fits within the memory constraints of your target device.
Where are the evaluation results and sample logs stored?
Results are written to the directory specified by --output_path, containing JSON files with aggregated metrics and individual sample predictions. When using --log_samples, the harness saves detailed inputs, model outputs, and ground-truth labels for each example, enabling detailed error analysis and result verification.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →