How to Write and Integrate Custom evaluate.py Scripts for Task Evaluation in SIA

To integrate custom evaluation logic into SIA, create an evaluate.py script that accepts a --gen-dir argument, computes metrics from generated artifacts, and writes a results.json file before exiting with status 0.

SIA (System for Intelligent Agents) is an open-source framework from hexo-ai for running and evaluating AI agent tasks. When you need to assess task performance beyond default metrics, you can write custom evaluate.py scripts that analyze generation outputs and report domain-specific scores. This guide walks through the implementation requirements, integration steps, and orchestration flow based on the actual source code.

What SIA Expects from a Custom evaluate.py Script

SIA’s orchestrator automatically discovers and executes task-specific evaluation scripts after each generation phase. The system looks for evaluate.py in specific locations and invokes it as a subprocess with standardized arguments.

Script Location and Discovery

The orchestrator locates your evaluation script using the find_evaluate_script function in sia/layout.py (line 63). SIA searches for the file in the following order:

  1. data/public/evaluate.py inside the task directory (preferred location)
  2. evaluate.py at the task root (fallback)

If the script exists, the orchestrator proceeds with execution; otherwise, the evaluation phase is skipped.

Command-Line Interface

When invoked, the orchestrator constructs a command equivalent to:

python evaluate.py --gen-dir /path/to/generation_directory

Your script must accept the --gen-dir argument. This path points to the directory containing all artifacts produced by the target agent during the generation phase, such as target_agent.py outputs, logs, and intermediate files.

Output Requirements

After computing metrics, your script must write a JSON file named results.json directly into the generation directory (the path provided by --gen-dir). The orchestrator reads this file using the constant Names.RESULTS_JSON defined in sia/layout.py (line 20) to report final outcomes.

The Five Requirements for Custom Evaluation Scripts

Based on the implementation in sia/orchestrator.py, every evaluate.py script must satisfy these criteria:

  1. Accept --gen-dir – Parse the command-line argument to locate generated artifacts.
  2. Read generated outputs – Open and process any files created by the target agent relative to the provided directory.
  3. Compute metrics – Calculate accuracy, F1, BLEU, custom rewards, or other relevant scores.
  4. Write results.json – Serialize results to a JSON file in the generation directory using the exact filename results.json.
  5. Exit with status 0 – Return a zero exit code on success; any non-zero exit is treated as an error and logged in the orchestration output.

Complete evaluate.py Template

Below is a minimal, production-ready template that demonstrates parsing arguments, handling errors, and writing the required output file. Place this code in data/public/evaluate.py within your task directory:

import argparse
import json
import os
import pathlib


def load_generated_output(gen_dir: pathlib.Path) -> str:
    """
    Example helper: read the output produced by the target agent.
    Adjust this to match the format your task generates.
    """
    # Assuming the target agent writes a file called `output.txt`

    output_file = gen_dir / "output.txt"
    if not output_file.is_file():
        raise FileNotFoundError(f"{output_file} not found")
    return output_file.read_text(encoding="utf-8")


def compute_metric(output: str) -> dict:
    """
    Replace this with your real scoring logic.
    Here we just return a dummy accuracy.
    """
    # Example: compute a fake accuracy based on output length

    accuracy = min(len(output) / 100.0, 1.0)
    return {"accuracy": round(accuracy, 3)}


def main() -> None:
    parser = argparse.ArgumentParser(
        description="Task-specific evaluation script for SIA."
    )
    parser.add_argument(
        "--gen-dir",
        required=True,
        help="Path to the generation directory produced by the target agent.",
    )
    args = parser.parse_args()

    gen_path = pathlib.Path(args.gen_dir).resolve()

    # 1. Load generated data

    try:
        agent_output = load_generated_output(gen_path)
    except Exception as exc:
        # Propagate a non-zero exit status so the orchestrator records an error.

        raise SystemExit(f"Failed to read generated output: {exc}") from exc

    # 2. Compute metrics

    results = compute_metric(agent_output)

    # 3. Write results.json (the orchestrator expects this exact filename)

    results_path = gen_path / "results.json"
    results_path.write_text(json.dumps(results, indent=2), encoding="utf-8")

    # 4. Successful exit – orchestrator will report status "success".

    print(f"Evaluation completed. Results written to {results_path}")


if __name__ == "__main__":
    main()

Integration Steps

Follow this workflow to deploy your custom evaluation script:

  1. Create the script – Add the file at data/public/evaluate.py inside your task directory (or at the task root if you prefer the fallback location).

  2. Make it executable (optional but recommended):

    chmod +x data/public/evaluate.py
  3. Test locally – Run the script manually to verify it produces the expected output:

    python data/public/evaluate.py --gen-dir /path/to/gen_1

    Confirm that results.json appears in the specified generation directory and contains valid JSON.

  4. Run the full SIA workflow – The orchestrator will automatically detect and execute the script during the evaluation phase according to the logic in sia/orchestrator.py.

How the Orchestrator Executes Your Script

The run_evaluation function in sia/orchestrator.py (line 85) handles the execution flow. It builds the command [python_exec, evaluate_script, "--gen-dir", gen_directory], captures stdout and stderr, and monitors the subprocess return code.

If the script exits with status 0, SIA reads the results.json file and reports the evaluation as successful. If the script exits with a non-zero status or fails to produce results.json, the orchestrator logs the error and marks the evaluation as failed. This subprocess-based approach ensures that your evaluation logic runs in isolation while still integrating seamlessly with the SIA pipeline.

Summary

  • Place evaluate.py in data/public/ (preferred) or the task root to enable automatic discovery via find_evaluate_script in sia/layout.py.
  • Accept --gen-dir as a command-line argument to receive the path to generation artifacts.
  • Write results.json to the generation directory so the orchestrator can parse the outcome using Names.RESULTS_JSON.
  • Exit with code 0 on success; non-zero exits trigger error reporting in the orchestration logs.
  • Reference files in sia/layout.py and sia/orchestrator.py to understand the discovery and execution mechanisms.

Frequently Asked Questions

Where should I place my evaluate.py file?

SIA searches for evaluate.py in two locations: first in data/public/evaluate.py within your task directory, then as a fallback in the task root. The data/public/ location is preferred because it keeps evaluation logic separate from private task assets and aligns with the find_evaluate_script implementation in sia/layout.py.

What arguments does SIA pass to evaluate.py?

The orchestrator passes exactly one required argument: --gen-dir, which contains the absolute path to the directory where the target agent wrote its outputs. Your script must parse this using argparse or similar to locate the files it needs to score.

How does the orchestrator know if evaluation succeeded?

The run_evaluation function in sia/orchestrator.py checks for two success indicators: a process exit code of 0 and the presence of a results.json file in the generation directory. If either condition fails, SIA reports an evaluation error in the orchestration logs.

Can I use external Python packages in my evaluation script?

Yes, you can import any third-party libraries installed in your Python environment, such as numpy, scikit-learn, or sacrebleu. Ensure dependencies are installed in the same environment where SIA runs, as the orchestrator invokes evaluate.py using the system Python executable configured for the workflow.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →