# How to Write and Integrate Custom evaluate.py Scripts for Task Evaluation in SIA

> Learn to write and integrate custom evaluate.py scripts for task evaluation in SIA. Compute metrics and output results.json for seamless integration before exiting with status 0.

- Repository: [Hexo Labs/sia](https://github.com/hexo-ai/sia)
- Tags: how-to-guide
- Published: 2026-06-12

---

**To integrate custom evaluation logic into SIA, create an [`evaluate.py`](https://github.com/hexo-ai/sia/blob/main/evaluate.py) script that accepts a `--gen-dir` argument, computes metrics from generated artifacts, and writes a [`results.json`](https://github.com/hexo-ai/sia/blob/main/results.json) file before exiting with status 0.**

SIA (System for Intelligent Agents) is an open-source framework from hexo-ai for running and evaluating AI agent tasks. When you need to assess task performance beyond default metrics, you can write custom [`evaluate.py`](https://github.com/hexo-ai/sia/blob/main/evaluate.py) scripts that analyze generation outputs and report domain-specific scores. This guide walks through the implementation requirements, integration steps, and orchestration flow based on the actual source code.

## What SIA Expects from a Custom evaluate.py Script

SIA’s orchestrator automatically discovers and executes task-specific evaluation scripts after each generation phase. The system looks for [`evaluate.py`](https://github.com/hexo-ai/sia/blob/main/evaluate.py) in specific locations and invokes it as a subprocess with standardized arguments.

### Script Location and Discovery

The orchestrator locates your evaluation script using the `find_evaluate_script` function in [`sia/layout.py`](https://github.com/hexo-ai/sia/blob/main/sia/layout.py) (line 63). SIA searches for the file in the following order:

1. **[`data/public/evaluate.py`](https://github.com/hexo-ai/sia/blob/main/data/public/evaluate.py)** inside the task directory (preferred location)
2. **[`evaluate.py`](https://github.com/hexo-ai/sia/blob/main/evaluate.py)** at the task root (fallback)

If the script exists, the orchestrator proceeds with execution; otherwise, the evaluation phase is skipped.

### Command-Line Interface

When invoked, the orchestrator constructs a command equivalent to:

```bash
python evaluate.py --gen-dir /path/to/generation_directory

```

Your script **must accept the `--gen-dir` argument**. This path points to the directory containing all artifacts produced by the target agent during the generation phase, such as [`target_agent.py`](https://github.com/hexo-ai/sia/blob/main/target_agent.py) outputs, logs, and intermediate files.

### Output Requirements

After computing metrics, your script must write a JSON file named **[`results.json`](https://github.com/hexo-ai/sia/blob/main/results.json)** directly into the generation directory (the path provided by `--gen-dir`). The orchestrator reads this file using the constant `Names.RESULTS_JSON` defined in [`sia/layout.py`](https://github.com/hexo-ai/sia/blob/main/sia/layout.py) (line 20) to report final outcomes.

## The Five Requirements for Custom Evaluation Scripts

Based on the implementation in [`sia/orchestrator.py`](https://github.com/hexo-ai/sia/blob/main/sia/orchestrator.py), every [`evaluate.py`](https://github.com/hexo-ai/sia/blob/main/evaluate.py) script must satisfy these criteria:

1. **Accept `--gen-dir`** – Parse the command-line argument to locate generated artifacts.
2. **Read generated outputs** – Open and process any files created by the target agent relative to the provided directory.
3. **Compute metrics** – Calculate accuracy, F1, BLEU, custom rewards, or other relevant scores.
4. **Write [`results.json`](https://github.com/hexo-ai/sia/blob/main/results.json)** – Serialize results to a JSON file in the generation directory using the exact filename [`results.json`](https://github.com/hexo-ai/sia/blob/main/results.json).
5. **Exit with status 0** – Return a zero exit code on success; any non-zero exit is treated as an error and logged in the orchestration output.

## Complete evaluate.py Template

Below is a minimal, production-ready template that demonstrates parsing arguments, handling errors, and writing the required output file. Place this code in [`data/public/evaluate.py`](https://github.com/hexo-ai/sia/blob/main/data/public/evaluate.py) within your task directory:

```python
import argparse
import json
import os
import pathlib


def load_generated_output(gen_dir: pathlib.Path) -> str:
    """
    Example helper: read the output produced by the target agent.
    Adjust this to match the format your task generates.
    """
    # Assuming the target agent writes a file called `output.txt`

    output_file = gen_dir / "output.txt"
    if not output_file.is_file():
        raise FileNotFoundError(f"{output_file} not found")
    return output_file.read_text(encoding="utf-8")


def compute_metric(output: str) -> dict:
    """
    Replace this with your real scoring logic.
    Here we just return a dummy accuracy.
    """
    # Example: compute a fake accuracy based on output length

    accuracy = min(len(output) / 100.0, 1.0)
    return {"accuracy": round(accuracy, 3)}


def main() -> None:
    parser = argparse.ArgumentParser(
        description="Task-specific evaluation script for SIA."
    )
    parser.add_argument(
        "--gen-dir",
        required=True,
        help="Path to the generation directory produced by the target agent.",
    )
    args = parser.parse_args()

    gen_path = pathlib.Path(args.gen_dir).resolve()

    # 1. Load generated data

    try:
        agent_output = load_generated_output(gen_path)
    except Exception as exc:
        # Propagate a non-zero exit status so the orchestrator records an error.

        raise SystemExit(f"Failed to read generated output: {exc}") from exc

    # 2. Compute metrics

    results = compute_metric(agent_output)

    # 3. Write results.json (the orchestrator expects this exact filename)

    results_path = gen_path / "results.json"
    results_path.write_text(json.dumps(results, indent=2), encoding="utf-8")

    # 4. Successful exit – orchestrator will report status "success".

    print(f"Evaluation completed. Results written to {results_path}")


if __name__ == "__main__":
    main()

```

## Integration Steps

Follow this workflow to deploy your custom evaluation script:

1. **Create the script** – Add the file at [`data/public/evaluate.py`](https://github.com/hexo-ai/sia/blob/main/data/public/evaluate.py) inside your task directory (or at the task root if you prefer the fallback location).

2. **Make it executable** (optional but recommended):
   ```bash
   chmod +x data/public/evaluate.py
   ```

3. **Test locally** – Run the script manually to verify it produces the expected output:
   ```bash
   python data/public/evaluate.py --gen-dir /path/to/gen_1
   ```

   Confirm that [`results.json`](https://github.com/hexo-ai/sia/blob/main/results.json) appears in the specified generation directory and contains valid JSON.

4. **Run the full SIA workflow** – The orchestrator will automatically detect and execute the script during the evaluation phase according to the logic in [`sia/orchestrator.py`](https://github.com/hexo-ai/sia/blob/main/sia/orchestrator.py).

## How the Orchestrator Executes Your Script

The `run_evaluation` function in [`sia/orchestrator.py`](https://github.com/hexo-ai/sia/blob/main/sia/orchestrator.py) (line 85) handles the execution flow. It builds the command `[python_exec, evaluate_script, "--gen-dir", gen_directory]`, captures stdout and stderr, and monitors the subprocess return code.

If the script exits with status 0, SIA reads the [`results.json`](https://github.com/hexo-ai/sia/blob/main/results.json) file and reports the evaluation as successful. If the script exits with a non-zero status or fails to produce [`results.json`](https://github.com/hexo-ai/sia/blob/main/results.json), the orchestrator logs the error and marks the evaluation as failed. This subprocess-based approach ensures that your evaluation logic runs in isolation while still integrating seamlessly with the SIA pipeline.

## Summary

- **Place [`evaluate.py`](https://github.com/hexo-ai/sia/blob/main/evaluate.py)** in `data/public/` (preferred) or the task root to enable automatic discovery via `find_evaluate_script` in [`sia/layout.py`](https://github.com/hexo-ai/sia/blob/main/sia/layout.py).
- **Accept `--gen-dir`** as a command-line argument to receive the path to generation artifacts.
- **Write [`results.json`](https://github.com/hexo-ai/sia/blob/main/results.json)** to the generation directory so the orchestrator can parse the outcome using `Names.RESULTS_JSON`.
- **Exit with code 0** on success; non-zero exits trigger error reporting in the orchestration logs.
- **Reference files** in [`sia/layout.py`](https://github.com/hexo-ai/sia/blob/main/sia/layout.py) and [`sia/orchestrator.py`](https://github.com/hexo-ai/sia/blob/main/sia/orchestrator.py) to understand the discovery and execution mechanisms.

## Frequently Asked Questions

### Where should I place my evaluate.py file?

SIA searches for [`evaluate.py`](https://github.com/hexo-ai/sia/blob/main/evaluate.py) in two locations: first in [`data/public/evaluate.py`](https://github.com/hexo-ai/sia/blob/main/data/public/evaluate.py) within your task directory, then as a fallback in the task root. The `data/public/` location is preferred because it keeps evaluation logic separate from private task assets and aligns with the `find_evaluate_script` implementation in [`sia/layout.py`](https://github.com/hexo-ai/sia/blob/main/sia/layout.py).

### What arguments does SIA pass to evaluate.py?

The orchestrator passes exactly one required argument: `--gen-dir`, which contains the absolute path to the directory where the target agent wrote its outputs. Your script must parse this using `argparse` or similar to locate the files it needs to score.

### How does the orchestrator know if evaluation succeeded?

The `run_evaluation` function in [`sia/orchestrator.py`](https://github.com/hexo-ai/sia/blob/main/sia/orchestrator.py) checks for two success indicators: a process exit code of 0 and the presence of a [`results.json`](https://github.com/hexo-ai/sia/blob/main/results.json) file in the generation directory. If either condition fails, SIA reports an evaluation error in the orchestration logs.

### Can I use external Python packages in my evaluation script?

Yes, you can import any third-party libraries installed in your Python environment, such as `numpy`, `scikit-learn`, or `sacrebleu`. Ensure dependencies are installed in the same environment where SIA runs, as the orchestrator invokes [`evaluate.py`](https://github.com/hexo-ai/sia/blob/main/evaluate.py) using the system Python executable configured for the workflow.