How to Add a Custom Evaluation Task to Eagle's LMMS-Eval Framework

You can add a custom evaluation task to Eagle's LMMS-Eval framework by creating a YAML configuration file and a Python utils module in a new folder under Eagle/lmms_eval/tasks/, which the auto-discovery registry automatically loads at runtime without requiring core code modifications.

Eagle extends its benchmarking capabilities through the LMMS-Eval harness, which uses a declarative YAML-based configuration system paired with Python utility modules. Understanding how to add a custom evaluation task to Eagle's lmms_eval framework allows you to integrate proprietary or novel datasets into the existing evaluation pipeline while leveraging the built-in registry and aggregation logic.

Understanding the Task Registry Architecture

Eagle's evaluation system relies on automatic task discovery mechanisms implemented in Eagle/lmms_eval/tasks/__init__.py. The include_path() function traverses the tasks/ directory tree and invokes include_task_folder() for each subdirectory containing YAML configurations.

When the framework encounters a YAML file, it loads the configuration via utils.load_yaml_config() and calls register_configurable_task() from Eagle/lmms_eval/api/registry.py. This factory function creates a dynamic subclass of ConfigurableTask using the parsed TaskConfig, then stores the instance in TASK_REGISTRY.

During evaluation, get_task(task_name, model_name) retrieves the registered task instance, which the harness uses to fetch data, generate predictions, and compute metrics. This architecture means new tasks are plug-and-play components that require no manual editing of core evaluation logic.

Step-by-Step Implementation Guide

Create the Task Directory Structure

Create a new folder for your benchmark inside the tasks directory:

mkdir -p Eagle/lmms_eval/tasks/mytask
touch Eagle/lmms_eval/tasks/mytask/mytask.yaml
touch Eagle/lmms_eval/tasks/mytask/utils.py

The folder name becomes the task identifier, and the registry will automatically detect any *.yaml files inside this tree when Python imports the lmms_eval module.

Configure the YAML Definition

The YAML file defines the dataset source, output type, and function mappings. The task: key is automatically inferred from the filename, so you only need to specify the dataset configuration and hook functions:


# Eagle/lmms_eval/tasks/mytask/mytask.yaml

dataset_path: my-dataset/CustomSet
dataset_kwargs:
  token: True
output_type: generate_until
doc_to_visual: !function utils.my_doc_to_visual
doc_to_text: !function utils.my_doc_to_text
doc_to_target: "answer"
generation_kwargs:
  max_new_tokens: 32
metadata:
  - version: 0.0
model_specific_prompt_kwargs:
  default:
    pre_prompt: "Question: "
    post_prompt: "\nAnswer:"

The !function tags reference callables in your utils.py module, allowing the framework to dynamically load custom logic for visual processing and text generation.

Implement the Utility Functions

Create utils.py with the required transformation functions. At minimum, you must implement doc_to_visual and doc_to_text:


# Eagle/lmms_eval/tasks/mytask/utils.py

import logging
from PIL import Image

eval_logger = logging.getLogger("lmms-eval")

def my_doc_to_visual(doc):
    """Return a list of PIL images for the given datum."""
    # Assume `doc["image_path"]` points to a local file.

    img = Image.open(doc["image_path"]).convert("RGB")
    return [img]

def my_doc_to_text(doc, model_specific_prompt_kwargs=None):
    """Create the prompt string sent to the LMM."""
    if model_specific_prompt_kwargs is None:
        model_specific_prompt_kwargs = {}
    pre = model_specific_prompt_kwargs.get("pre_prompt", "")
    post = model_specific_prompt_kwargs.get("post_prompt", "")
    return f"{pre}{doc['question']}{post}"

If your task requires custom accuracy calculations or metric aggregation, add a process_results function and reference it in the YAML with process_results: !function utils.my_process_results. Examine existing implementations like Eagle/lmms_eval/tasks/vqav2/utils.py for reference patterns.

Register Task Groups (Optional)

To organize tasks into benchmark suites, create a group definition file:


# Eagle/lmms_eval/tasks/mygroup.yaml

task: mytask
group: mybenchmark

When the registry walks the folder structure, it automatically adds mytask to the group mybenchmark. This allows you to run the entire collection via --tasks mybenchmark rather than listing individual tasks.

Executing Your Custom Task

Once the files are in place, the task is immediately available through the registry. Run evaluation from the command line:

python Eagle/evaluate_lmms_eval.py \
  --model eagle \
  --tasks mytask \
  --output_dir ./eval_results

Alternatively, invoke the task programmatically:

from Eagle.lmms_eval.tasks import get_task

task = get_task("mytask", model_name="eagle")
result = task.run()

The framework automatically discovers the YAML configuration, imports the utils module, and instantiates the ConfigurableTask subclass with your custom logic.

Key Source Files Reference

Understanding these core files helps when debugging or extending task functionality:

Summary

  • Auto-registration: The framework discovers new tasks by walking Eagle/lmms_eval/tasks/ and automatically registers YAML configurations via register_configurable_task().
  • Minimal requirements: You only need a YAML config and a utils.py module with doc_to_visual and doc_to_text functions.
  • No core edits: Tasks are plug-and-play components; the registry handles instantiation without modifying evaluate_lmms_eval.py or other core files.
  • Group support: Optional group YAML files allow you to organize multiple tasks into benchmark suites for batch evaluation.

Frequently Asked Questions

Do I need to modify the core Eagle code to add a custom evaluation task?

No. The registry architecture in Eagle/lmms_eval/tasks/__init__.py automatically discovers and registers tasks based on YAML files found in the tasks/ directory. Simply creating a new folder with a YAML config and utils module is sufficient; the framework imports and registers the task at runtime without requiring changes to evaluate_lmms_eval.py or other core modules.

What are the required functions in the utils.py module?

You must implement doc_to_visual(doc) which returns a list of PIL Images, and doc_to_text(doc, model_specific_prompt_kwargs) which returns the prompt string. Optionally, you can implement process_results(doc, results) for custom metric calculation if the default aggregation logic does not meet your requirements. Reference existing tasks like vqav2 for implementation patterns.

How does the registry handle YAML configuration loading?

The include_path() function walks the task directory tree, identifies YAML files, and calls utils.load_yaml_config() to parse the configuration. It then invokes register_configurable_task() from Eagle/lmms_eval/api/registry.py, which creates a subclass of ConfigurableTask with the parsed configuration and stores it in TASK_REGISTRY under the task name derived from the filename.

Can I use custom metrics and result processing functions?

Yes. Define a custom processing function in your utils.py file (e.g., my_process_results) and reference it in your YAML configuration using process_results: !function utils.my_process_results. This function receives the document and model results, allowing you to implement task-specific accuracy calculations or metric aggregations before the harness computes final scores.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →