# How to Add a Custom Evaluation Task to Eagle's LMMS-Eval Framework

> Easily add custom evaluation tasks to Eagle's LMMS-Eval framework. Create YAML and Python files in a new folder for automatic discovery without core code changes.

- Repository: [NVIDIA Research Projects/Eagle](https://github.com/NVlabs/Eagle)
- Tags: how-to-guide
- Published: 2026-06-28

---

**You can add a custom evaluation task to Eagle's LMMS-Eval framework by creating a YAML configuration file and a Python utils module in a new folder under `Eagle/lmms_eval/tasks/`, which the auto-discovery registry automatically loads at runtime without requiring core code modifications.**

Eagle extends its benchmarking capabilities through the LMMS-Eval harness, which uses a declarative YAML-based configuration system paired with Python utility modules. Understanding how to add a custom evaluation task to Eagle's lmms_eval framework allows you to integrate proprietary or novel datasets into the existing evaluation pipeline while leveraging the built-in registry and aggregation logic.

## Understanding the Task Registry Architecture

Eagle's evaluation system relies on automatic task discovery mechanisms implemented in [`Eagle/lmms_eval/tasks/__init__.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/lmms_eval/tasks/__init__.py). The `include_path()` function traverses the `tasks/` directory tree and invokes `include_task_folder()` for each subdirectory containing YAML configurations.

When the framework encounters a YAML file, it loads the configuration via `utils.load_yaml_config()` and calls `register_configurable_task()` from [`Eagle/lmms_eval/api/registry.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/lmms_eval/api/registry.py). This factory function creates a dynamic subclass of **ConfigurableTask** using the parsed **TaskConfig**, then stores the instance in `TASK_REGISTRY`. 

During evaluation, `get_task(task_name, model_name)` retrieves the registered task instance, which the harness uses to fetch data, generate predictions, and compute metrics. This architecture means new tasks are plug-and-play components that require no manual editing of core evaluation logic.

## Step-by-Step Implementation Guide

### Create the Task Directory Structure

Create a new folder for your benchmark inside the tasks directory:

```bash
mkdir -p Eagle/lmms_eval/tasks/mytask
touch Eagle/lmms_eval/tasks/mytask/mytask.yaml
touch Eagle/lmms_eval/tasks/mytask/utils.py

```

The folder name becomes the task identifier, and the registry will automatically detect any `*.yaml` files inside this tree when Python imports the `lmms_eval` module.

### Configure the YAML Definition

The YAML file defines the dataset source, output type, and function mappings. The `task:` key is automatically inferred from the filename, so you only need to specify the dataset configuration and hook functions:

```yaml

# Eagle/lmms_eval/tasks/mytask/mytask.yaml

dataset_path: my-dataset/CustomSet
dataset_kwargs:
  token: True
output_type: generate_until
doc_to_visual: !function utils.my_doc_to_visual
doc_to_text: !function utils.my_doc_to_text
doc_to_target: "answer"
generation_kwargs:
  max_new_tokens: 32
metadata:
  - version: 0.0
model_specific_prompt_kwargs:
  default:
    pre_prompt: "Question: "
    post_prompt: "\nAnswer:"

```

The `!function` tags reference callables in your [`utils.py`](https://github.com/NVlabs/Eagle/blob/main/utils.py) module, allowing the framework to dynamically load custom logic for visual processing and text generation.

### Implement the Utility Functions

Create [`utils.py`](https://github.com/NVlabs/Eagle/blob/main/utils.py) with the required transformation functions. At minimum, you must implement `doc_to_visual` and `doc_to_text`:

```python

# Eagle/lmms_eval/tasks/mytask/utils.py

import logging
from PIL import Image

eval_logger = logging.getLogger("lmms-eval")

def my_doc_to_visual(doc):
    """Return a list of PIL images for the given datum."""
    # Assume `doc["image_path"]` points to a local file.

    img = Image.open(doc["image_path"]).convert("RGB")
    return [img]

def my_doc_to_text(doc, model_specific_prompt_kwargs=None):
    """Create the prompt string sent to the LMM."""
    if model_specific_prompt_kwargs is None:
        model_specific_prompt_kwargs = {}
    pre = model_specific_prompt_kwargs.get("pre_prompt", "")
    post = model_specific_prompt_kwargs.get("post_prompt", "")
    return f"{pre}{doc['question']}{post}"

```

If your task requires custom accuracy calculations or metric aggregation, add a `process_results` function and reference it in the YAML with `process_results: !function utils.my_process_results`. Examine existing implementations like [`Eagle/lmms_eval/tasks/vqav2/utils.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/lmms_eval/tasks/vqav2/utils.py) for reference patterns.

### Register Task Groups (Optional)

To organize tasks into benchmark suites, create a group definition file:

```yaml

# Eagle/lmms_eval/tasks/mygroup.yaml

task: mytask
group: mybenchmark

```

When the registry walks the folder structure, it automatically adds `mytask` to the group `mybenchmark`. This allows you to run the entire collection via `--tasks mybenchmark` rather than listing individual tasks.

## Executing Your Custom Task

Once the files are in place, the task is immediately available through the registry. Run evaluation from the command line:

```bash
python Eagle/evaluate_lmms_eval.py \
  --model eagle \
  --tasks mytask \
  --output_dir ./eval_results

```

Alternatively, invoke the task programmatically:

```python
from Eagle.lmms_eval.tasks import get_task

task = get_task("mytask", model_name="eagle")
result = task.run()

```

The framework automatically discovers the YAML configuration, imports the utils module, and instantiates the ConfigurableTask subclass with your custom logic.

## Key Source Files Reference

Understanding these core files helps when debugging or extending task functionality:

- **[`Eagle/lmms_eval/tasks/__init__.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/lmms_eval/tasks/__init__.py)**: Contains `include_path()` and `include_task_folder()` which implement the directory walking and auto-registration logic.
- **[`Eagle/lmms_eval/api/registry.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/lmms_eval/api/registry.py)**: Houses `register_configurable_task()`, `TASK_REGISTRY`, and the `get_task()` factory used by the evaluation harness.
- **[`Eagle/lmms_eval/tasks/vqav2/utils.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/lmms_eval/tasks/vqav2/utils.py)**: Reference implementation showing standard patterns for `doc_to_visual`, `doc_to_text`, and result processing.
- **[`Eagle/evaluate_lmms_eval.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/evaluate_lmms_eval.py)**: The main entry point that loads tasks from the registry and orchestrates evaluation runs.

## Summary

- **Auto-registration**: The framework discovers new tasks by walking `Eagle/lmms_eval/tasks/` and automatically registers YAML configurations via `register_configurable_task()`.
- **Minimal requirements**: You only need a YAML config and a [`utils.py`](https://github.com/NVlabs/Eagle/blob/main/utils.py) module with `doc_to_visual` and `doc_to_text` functions.
- **No core edits**: Tasks are plug-and-play components; the registry handles instantiation without modifying [`evaluate_lmms_eval.py`](https://github.com/NVlabs/Eagle/blob/main/evaluate_lmms_eval.py) or other core files.
- **Group support**: Optional group YAML files allow you to organize multiple tasks into benchmark suites for batch evaluation.

## Frequently Asked Questions

### Do I need to modify the core Eagle code to add a custom evaluation task?

No. The registry architecture in [`Eagle/lmms_eval/tasks/__init__.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/lmms_eval/tasks/__init__.py) automatically discovers and registers tasks based on YAML files found in the `tasks/` directory. Simply creating a new folder with a YAML config and utils module is sufficient; the framework imports and registers the task at runtime without requiring changes to [`evaluate_lmms_eval.py`](https://github.com/NVlabs/Eagle/blob/main/evaluate_lmms_eval.py) or other core modules.

### What are the required functions in the utils.py module?

You must implement `doc_to_visual(doc)` which returns a list of PIL Images, and `doc_to_text(doc, model_specific_prompt_kwargs)` which returns the prompt string. Optionally, you can implement `process_results(doc, results)` for custom metric calculation if the default aggregation logic does not meet your requirements. Reference existing tasks like `vqav2` for implementation patterns.

### How does the registry handle YAML configuration loading?

The `include_path()` function walks the task directory tree, identifies YAML files, and calls `utils.load_yaml_config()` to parse the configuration. It then invokes `register_configurable_task()` from [`Eagle/lmms_eval/api/registry.py`](https://github.com/NVlabs/Eagle/blob/main/Eagle/lmms_eval/api/registry.py), which creates a subclass of `ConfigurableTask` with the parsed configuration and stores it in `TASK_REGISTRY` under the task name derived from the filename.

### Can I use custom metrics and result processing functions?

Yes. Define a custom processing function in your [`utils.py`](https://github.com/NVlabs/Eagle/blob/main/utils.py) file (e.g., `my_process_results`) and reference it in your YAML configuration using `process_results: !function utils.my_process_results`. This function receives the document and model results, allowing you to implement task-specific accuracy calculations or metric aggregations before the harness computes final scores.