How to Create Custom Tasks in SIA: Required Directory Structure Guide
To create custom tasks in SIA, you must organize your task directory with data/public/, data/private/, and reference/ subdirectories, including a mandatory task.md file and optionally an evaluate.py script that implements the evaluate(gen_dir: str) -> dict interface.
SIA (Self-Improving Agents) from the hexo-ai/sia repository executes an automated improvement loop on task directories that follow a strict schema. When you create custom tasks with the required directory structure, the orchestrator can locate the task description, evaluation data, and reference agent template needed to bootstrap the meta-agent.
Required Directory Structure
SIA expects a specific layout relative to your task root directory. The sia/orchestrator.py file validates this structure before initiating the self-improvement loop.
Public Data Folder
The data/public/ directory contains files the target agent may read during generation. This includes the mandatory task.md file describing the problem, any training datasets, and optional test data. According to the source code in sia/orchestrator.py, the meta-agent parses task.md to generate the initial target-agent implementation.
Private Evaluation Data
The data/private/ directory holds held-out data used exclusively by the evaluator. Files placed here are never exposed to the LLM during the generation phase. This separation ensures unbiased evaluation of the generated agents.
Reference Agent Template
The reference/ directory must contain the reference agent (reference_target_agent.py) that serves as the seed for the meta-agent. You can copy the template from sia/tasks/_shared/reference_target_agent.py into this folder. Optionally, include a SAMPLE_TASK_DESCRIPTIONS.md file to provide the meta-agent with additional context about similar tasks.
Evaluation Script
While optional, you can include an evaluate.py file in data/public/ that implements the function evaluate(gen_dir: str) -> dict. The orchestrator imports and executes this function after each generation, returning metrics that feed into the feedback agent's improvement prompts.
Creating the Task Skeleton
Follow these steps to initialize a compliant task directory. This structure aligns with the specifications documented in the repository's walkthrough and README.
# Create the required directory hierarchy
mkdir -p my-task/{data/public,data/private,reference}
# Add the mandatory task description
cp task_description.md my-task/data/public/task.md
# Add public datasets the agent can access
cp train.json my-task/data/public/
cp questions.json my-task/data/public/
# Add private ground-truth data for evaluation
cp answers.json my-task/data/private/
# Copy the reference agent template from the SIA package
cp sia/tasks/_shared/reference_target_agent.py my-task/reference/
# Optional: Add sample task descriptions to guide the meta-agent
cat > my-task/reference/SAMPLE_TASK_DESCRIPTIONS.md <<'EOF'
# Example Tasks
- Summarize technical documentation
- Classify sentiment in product reviews
EOF
Implementing the Evaluation Function
Create a custom evaluate.py in data/public/ to score generated outputs against your private data. The function must accept a string path to the generation directory and return a dictionary of metrics.
# my-task/data/public/evaluate.py
import json
import pathlib
def evaluate(gen_dir: str) -> dict:
"""Score predictions against ground truth.
Args:
gen_dir: Path to the directory containing generated files
Returns:
dict: Evaluation metrics passed to the feedback agent
"""
# Load predictions written by the target agent
pred_path = pathlib.Path(gen_dir) / "predictions.json"
preds = json.loads(pred_path.read_text())
# Load ground truth from private data
truth_path = pathlib.Path(__file__).parents[1] / "private" / "answers.json"
truths = json.loads(truth_path.read_text())
# Calculate accuracy metric
correct = sum(p == t for p, t in zip(preds, truths))
return {"accuracy": correct / len(truths)}
Place this file in my-task/data/public/. The orchestrator automatically detects and executes it after each generation cycle, feeding the returned metrics into the feedback loop.
Running Your Custom Task
Once the directory structure matches the required layout, invoke SIA with the path to your task directory:
sia run --task_dir ./my-task --max_gen 5 --run_id 1
The orchestrator creates per-generation run folders (runs/run_1/gen_1/, etc.) containing the generated target_agent.py, execution logs, and improvement notes. The meta-agent iteratively refines the target agent based on the evaluation scores until reaching the specified --max_gen limit.
Summary
- Directory layout: Create
data/public/,data/private/, andreference/subdirectories in your task root. - Mandatory files: Include
task.mdindata/public/describing the problem, and copyreference_target_agent.pyfromsia/tasks/_shared/into thereference/folder. - Evaluation: Optionally implement
evaluate(gen_dir: str) -> dictindata/public/evaluate.pyto provide automated scoring for the feedback agent. - Execution: Run
sia run --task_dir <path>to start the self-improving loop on your custom task.
Frequently Asked Questions
What happens if I don't include an evaluate.py file?
If evaluate.py is missing from data/public/, the orchestrator will still run the self-improvement loop, but the feedback agent will receive less specific guidance about the target agent's performance. While the meta-agent can suggest improvements based on execution logs, providing an evaluation function enables data-driven optimization using concrete metrics from your private dataset.
Can I modify the reference_target_agent.py template?
Yes, you can customize the reference_target_agent.py in your task's reference/ directory. The orchestrator uses this file as the seed for the first generation. However, ensure the template maintains the expected interface (such as a solve() or run() method) that the generated code will implement, as the meta-agent assumes this structure when creating subsequent iterations.
How does SIA handle the private data during execution?
Files in data/private/ are never exposed to the LLM during the generation phase. The sia/orchestrator.py only passes paths to data/public/ when setting up the target agent's environment. The private data is accessed solely by your evaluate.py function after the generation completes, ensuring a clean separation between training/evaluation data and the agent's working context.
What is the minimum directory structure required to run a task?
The absolute minimum requires three components: data/public/task.md containing the task description, reference/reference_target_agent.py copied from sia/tasks/_shared/reference_target_agent.py, and an empty data/private/ directory (even if unused). Without these elements, the orchestrator will raise errors during the initialization phase.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →