# How to Evaluate Eagle on VLM Benchmarks Using lmms-eval: A Complete Guide

> Learn how to evaluate Eagle on VLM benchmarks with lmms-eval. Our guide covers loading checkpoints, preparing vision towers, and running batched inference for comprehensive vision-language task analysis.

- Repository: [NVIDIA Research Projects/Eagle](https://github.com/NVlabs/Eagle)
- Tags: how-to-guide
- Published: 2026-06-28

---

**Eagle registers a model wrapper named `eagle` in the lmms-eval harness that loads pretrained checkpoints, prepares the vision tower, and runs batched inference across vision-language tasks through the `simple_evaluate` loop in [`Eagle/lmms_eval/evaluator.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/lmms_eval/evaluator.py).**

The **NVlabs/Eagle** repository provides a native integration with the **lmms-eval** framework, allowing you to evaluate Eagle vision-language models on standard benchmarks like VQAv2, MM-Bench, and MME. This integration leverages a custom model wrapper that conforms to the lmms abstract class requirements while handling Eagle-specific tokenization and image preprocessing. Understanding how to evaluate Eagle on VLM benchmarks using lmms-eval enables reproducible benchmarking and fair comparison against other vision-language models.

## Understanding the lmms-eval Integration Architecture

### Model Registration and Loading

The evaluation harness recognizes Eagle through the `@register_model("eagle")` decorator in [`Eagle/lmms_eval/models/eagle.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/lmms_eval/models/eagle.py). When you specify `--model eagle` on the command line, the wrapper invokes `load_pretrained_model` from [`Eagle/eagle/model/builder.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/eagle/model/builder.py) to initialize the full inference stack.

This builder function performs three critical operations:
1. Instantiates the tokenizer, language model, and optional LoRA weights.
2. Adds special image-patch, start, and end tokens to the tokenizer vocabulary.
3. Loads the vision tower (CLIP-style image encoder) and attaches its image processor.

The wrapper then returns a model instance that satisfies the lmms-eval abstract class interface, enabling the harness to treat Eagle like any other supported vision-language model.

### Task Initialization and Data Pipeline

Before inference begins, `initialize_tasks` (defined in [`Eagle/lmms_eval/__init__.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/lmms_eval/__init__.py)) registers every available VLM benchmark under the `ALL_TASKS` registry. Each task definition in `Eagle/lmms_eval/tasks/<benchmark>/` supplies:
- A `data_path` pointing to the dataset (downloaded automatically by lmms-eval).
- A prompt template that combines visual tokens with language model input.
- Metric calculation functions (e.g., accuracy for VQAv2, BLEU for captioning).

## Step-by-Step Evaluation Workflow

The evaluation process follows three distinct phases orchestrated by `evaluator.simple_evaluate` in [`Eagle/lmms_eval/evaluator.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/lmms_eval/evaluator.py):

1. **Preprocessing**: The harness loads images and calls `process_images` from [`Eagle/eagle/mm_utils.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/eagle/mm_utils.py) to convert visual inputs into tensors compatible with the vision tower.

2. **Encoding and Inference**: The tokenizer encodes the prompt with appended special image tokens, padding to the model's `max_length`. The model runs batched inference according to the `batch_size` parameter on the specified device (CPU or GPU).

3. **Post-processing and Scoring**: Generated answers undergo normalization (case conversion, punctuation removal) before the task-specific metric functions compute final scores.

## Running Evaluations: Code Examples

All commands assume execution from the repository root. The entry point is [`Eagle/evaluate_lmms_eval.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/evaluate_lmms_eval.py), which parses arguments and delegates to the evaluation loop.

### Single Benchmark Evaluation (VQAv2)

Run a quick evaluation on VQAv2 to verify your setup:

```bash
python -m Eagle.evaluate_lmms_eval \
    --model eagle \
    --model_args "pretrained=NVEagle/Eagle-X5-7B,device=cuda" \
    --tasks vqav2 \
    --batch_size 4 \
    --output_path ./eval_results \
    --log_samples

```

The `--model eagle` flag selects the wrapper defined in [`Eagle/lmms_eval/models/eagle.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/lmms_eval/models/eagle.py). The `pretrained` argument accepts any HuggingFace hub checkpoint or local path.

### Multi-Benchmark Suite Evaluation

Evaluate across multiple benchmarks simultaneously by comma-separating task names:

```bash
python -m Eagle.evaluate_lmms_eval \
    --model eagle \
    --model_args "pretrained=NVEagle/Eagle-X5-7B,device=cuda" \
    --tasks vqav2,mmbench,mmmu,mme \
    --batch_size 2 \
    --output_path ./full_eval \
    --log_samples \
    --wandb_log_samples \
    --wandb_args "project=EagleVLM,entity=your_wandb_user"

```

The harness supports wildcard expansion (`*`) for task selection. Weights & Biases integration is optional and configured via `--wandb_args`.

### YAML Configuration for Reproducibility

For reproducible experiments, store parameters in a YAML file:

```yaml
model: eagle
model_args: "pretrained=NVEagle/Eagle-X5-7B,device=cuda"
tasks: vqav2,mmbench,mmmu,mme
batch_size: 2
log_samples: true
output_path: ./config_eval

```

Execute with:

```bash
python -m Eagle.evaluate_lmms_eval --config config.yaml

```

The script parses the YAML (lines 199-205 in [`Eagle/evaluate_lmms_eval.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/evaluate_lmms_eval.py)) and constructs an `argparse.Namespace` for each configuration entry.

## Key Source Files and Their Roles

Understanding the source structure helps when debugging or extending the evaluation pipeline:

- **[`Eagle/lmms_eval/models/eagle.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/lmms_eval/models/eagle.py)**: Registers the Eagle model wrapper and implements the interface required by the lmms-eval harness.
- **[`Eagle/eagle/model/builder.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/eagle/model/builder.py)**: Loads pretrained checkpoints, adds image tokens to the tokenizer, and initializes the vision tower.
- **[`Eagle/evaluate_lmms_eval.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/evaluate_lmms_eval.py)**: CLI entry point that parses arguments (including YAML configs) and invokes the evaluation loop.
- **[`Eagle/lmms_eval/evaluator.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/lmms_eval/evaluator.py)**: Contains `simple_evaluate`, the core evaluation loop that runs inference and aggregates metrics across tasks.
- **`Eagle/lmms_eval/tasks/`**: Directory containing individual benchmark definitions (e.g., `vqav2`, `mmbench`, `mme`), each with data loading utilities and metric functions.
- **[`Eagle/eagle/mm_utils.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/eagle/mm_utils.py)**: Utility functions for inserting image tokens into prompts and handling image processor transformations.

## Summary

- **Eagle integrates natively with lmms-eval** through a registered model wrapper in [`Eagle/lmms_eval/models/eagle.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/lmms_eval/models/eagle.py) that handles checkpoint loading and vision tower initialization.
- **The evaluation workflow** involves model registration, task initialization, and the `simple_evaluate` inference loop that preprocesses images, encodes prompts, and computes task-specific metrics.
- **You can run evaluations** via command-line arguments or YAML configuration files, supporting both single benchmarks and multi-benchmark suites with optional Weights & Biases logging.
- **Key files** include the model builder ([`builder.py`](https://github.com/NVlabs/Eagle/blob/main/builder.py)), evaluation entry point ([`evaluate_lmms_eval.py`](https://github.com/NVlabs/Eagle/blob/main/evaluate_lmms_eval.py)), and task definitions under `lmms_eval/tasks/`.

## Frequently Asked Questions

### What checkpoints are compatible with the Eagle lmms-eval wrapper?

Any Eagle checkpoint hosted on Hugging Face Hub or stored locally works with the wrapper. Pass the checkpoint identifier via the `pretrained` argument in `--model_args`, such as `pretrained=NVEagle/Eagle-X5-7B` or `pretrained=/path/to/local/checkpoint`. The builder automatically detects and loads associated vision tower weights.

### How do I add a custom VLM benchmark to the evaluation harness?

Create a new directory under `Eagle/lmms_eval/tasks/` containing a task definition file that specifies the dataset path, prompt template, and metric functions. The `initialize_tasks` function in [`Eagle/lmms_eval/__init__.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/lmms_eval/__init__.py) automatically discovers new task modules, making them available via the `--tasks` flag without modifying core evaluation code.

### Can I run evaluation on CPU or multiple GPUs?

Yes. Specify `device=cpu` in `--model_args` for CPU inference, or `device=cuda` for GPU acceleration. The lmms-eval harness handles device placement through the model wrapper. For multi-GPU evaluation, the harness supports data parallelism, though you should verify that your specific Eagle checkpoint fits within the memory constraints of your target device.

### Where are the evaluation results and sample logs stored?

Results are written to the directory specified by `--output_path`, containing JSON files with aggregated metrics and individual sample predictions. When using `--log_samples`, the harness saves detailed inputs, model outputs, and ground-truth labels for each example, enabling detailed error analysis and result verification.