How to Use the SIA Web Dashboard to Visualize Run Results

The SIA web dashboard is a FastAPI-based UI that automatically launches during sia run to visualize multi-generation agent evolution through a vanilla JavaScript interface, or can be started manually via sia web to browse historical runs stored in the runs/ directory.

The SIA web dashboard provides real-time visualization of agent evolution experiments in the hexo-ai/sia repository. This lightweight interface renders code, prompts, evaluation metrics, and execution trajectories from any SIA run without requiring external JavaScript frameworks. It serves as both a live monitoring tool during active runs and a historical browser for completed experiments.

Architecture Overview

The dashboard consists of three decoupled layers that communicate through JSON REST endpoints.

Data Layer

In sia/web/runs.py, pure Python functions scan the runs/ directory tree and parse JSON/text artifacts. These functions expose data as Pydantic models (RunSummary, RunDetail, etc.) without importing FastAPI, making them easy to test in isolation. The layer reads agent code, prompts, evaluation logs, and trajectories directly from the filesystem.

Web Server

The sia/web/server.py module constructs a FastAPI application and registers REST endpoints including /api/runs, /api/runs/{run_name}, and /api/runs/{run_name}/gens/{gen_name}/artifact/{label}. It provides two entry points: serve() for blocking execution and serve_in_background() for daemon thread deployment.

Frontend

A single-page application in sia/web/static/index.html uses vanilla JavaScript to fetch the JSON endpoints. It builds a sidebar of runs, renders generation accuracy charts, and provides tabbed views for code, prompts, and trajectories. The frontend includes custom markdown rendering and Python syntax highlighting without external dependencies.

Launching the Dashboard

You can start the SIA web dashboard automatically during experiments, manually from the CLI, or programmatically from Python code.

Automatic Launch During Runs

By default, sia run launches the dashboard in a background daemon thread on http://127.0.0.1:8000. The --no-web flag disables this behavior.


# Start a run with live dashboard (default behavior)

sia run --task gpqa --max_gen 5 --run_id 1

Your default browser opens automatically, or you can manually navigate to http://127.0.0.1:8000.

Manual CLI Start

If you stopped the background thread or want to analyze existing runs later, install the optional web dependencies and start the server:

pip install 'sia-agent[web]'
sia web --runs-dir ./runs --port 8080

Programmatic Use

Import serve_in_background from sia.web.server to launch the dashboard while orchestrating runs in Python:

from sia.web.server import serve_in_background

# Launch on custom host/port

thread = serve_in_background(
    host="0.0.0.0", 
    port=9000, 
    runs_dir="./runs"
)

# Thread runs as daemon until process exits

The dashboard organizes information hierarchically: runs in the sidebar, generations in the main view, and artifacts in tabbed panes.

Selecting Runs

The left sidebar lists all directories matching run_<id> in the runs/ folder, displaying a badge with the best accuracy achieved. Clicking a run fetches its RunSummary via /api/runs/{run_name} and updates the header with meta-agent, target-agent, and task profile chips.

Browsing Generations

Below the accuracy line chart, a row of pills represents each generation (gen_<n>). Selecting a pill switches the active generation while preserving your current tab selection. This allows you to compare code or prompts across generations without losing context.

OpenHands Session Integration

When a generation produces an openhands_trajectory directory, a badge appears in the interface. You can list sessions and inspect events through the /api/runs/{run_name}/gens/{gen_name}/openhands/ endpoints.

Understanding the Tab Views

Each generation presents six tabs that fetch specific artifacts from the data layer.

Overview

Displays headline evaluation numbers from EvalSummary and per-domain accuracy tables. This tab provides the quantitative summary of the generation's performance.

Code

Fetches target_agent.py via /api/runs/{run_name}/gens/{gen_name}/artifact/target_agent and displays it with custom Python syntax highlighting. This shows the evolved agent code for the selected generation.

Prompt

Renders either meta_agent_prompt.txt (for generation 1) or feedback_agent_prompt.txt (for generation ≥ 2) as markdown. This reveals the instructions given to the agent during that generation.

Improvement

Shows the improvement.md file generated by the feedback agent, explaining what changes were made between the previous generation and the current one.

Trajectories

Allows selection of a question ID to view the full OpenAI or OpenHands chat turns stored in agent_execution/execution_q<id>.json. This provides step-by-step execution traces for debugging agent behavior.

Logs

Provides a dropdown to view target_agent_stdout.log or evaluation.log for runtime output and evaluation details.

Embedding the API

Because the UI consumes a JSON REST API, you can embed SIA visualizations in external tools or scripts by calling the same endpoints:

// Fetch a generation's agent code from the browser console
fetch("/api/runs/run_1/gens/gen_2/artifact/target_agent")
  .then(r => r.text())
  .then(code => console.log(code));

Summary

  • The SIA web dashboard consists of a data layer (sia/web/runs.py), FastAPI server (sia/web/server.py), and vanilla JavaScript frontend (sia/web/static/index.html).
  • Install the web extra with pip install 'sia-agent[web]' to enable the sia web command.
  • The dashboard launches automatically during sia run on http://127.0.0.1:8000, or can be started manually via CLI or Python API.
  • The interface provides six tabs (Overview, Code, Prompt, Improvement, Trajectories, Logs) to inspect every artifact produced during agent evolution.
  • All functionality is exposed through REST endpoints, allowing external tools to consume run data programmatically.

Frequently Asked Questions

How do I view runs after the SIA process has completed?

Start the dashboard manually using sia web --runs-dir ./runs --port 8080. This serves the existing runs/ directory without executing new experiments, allowing you to browse historical results indefinitely.

Can I customize the dashboard host and port?

Yes. When using the Python API, pass host and port parameters to serve_in_background() or serve(). From the CLI, use the --port flag with sia web. Note that sia run automatically uses 127.0.0.1:8000 unless disabled with --no-web.

Why does the dashboard use vanilla JavaScript instead of a framework like React?

The frontend in sia/web/static/index.html intentionally avoids external JavaScript dependencies to minimize maintenance overhead and ensure the tool works offline. It implements custom markdown rendering and Python syntax highlighting natively for fast loading and zero build steps.

What data format does the dashboard use for trajectories?

Trajectories are stored as JSON files in agent_execution/execution_q<id>.json within each generation directory. The dashboard parses these OpenAI/OpenHands chat logs and renders them as interactive conversation threads in the Trajectories tab.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →