Webwright Workspace Structure Explained: plan.md, final_script.py, and finalruns/

A Webwright workspace is a self-contained directory governed by the LocalWorkspaceEnvironment class that stores the task plan (plan.md), the canonical Playwright script (final_script.py), and isolated execution artifacts in sequentially numbered finalruns/run_<id>/ folders.

Microsoft Webwright agents rely on a strict filesystem contract to convert natural language tasks into verified browser automations. Understanding the Webwright workspace structure is essential for tracing how the environment implementation in src/webwright/environments/local_workspace.py orchestrates the planning, authoring, and validation phases.

Top-Level Directory Layout

When the prepare() method initializes a workspace, it creates the following hierarchy:

WORKSPACE_DIR/
├── plan.md                     # Markdown checklist of “critical points” (CPs)

├── final_script.py            # Canonical script that solves the task

├── task.json                   # Serialized task metadata written by the agent

├── self_reflect_config.json    # Prompt templates for self-reflection (optional)

├── finalruns/                  # One sub-folder per clean run of final_script.py

│   ├── run_001/
│   │   ├── final_script.py      # Copy of the canonical script for this run

│   │   ├── final_script_log.txt # Action log produced by the run

│   │   ├── screenshots/         # PNG/JPG screenshots, one per CP

│   │   │   ├── final_execution_1_apply_constraint.png
│   │   │   ├── final_execution_2_sort.png
│   │   │   └── …
│   │   └── self_reflect_result.json  # Verdict from the self-reflection tool

│   ├── run_002/   (… same structure as run_001 …)
│   └── …
└── … other auxiliary files (e.g. .tmp, logs, steps) created by the environment

The agent first writes the plan, then authors a reusable script, and finally executes it within isolated run folders to generate verifiable evidence for every critical point.

plan.md – The Critical Points Checklist

The plan.md file sits at the workspace root and serves as the ground-truth specification for task completion.

  • Location: Root of the workspace directory.

  • Content: Contains a # Critical Points section where each CP is a markdown task list item (- [ ]).

  • CLI-Tool Mode: When operating in CLI-tool mode, the file also includes a # Parameters table listing every CP that should become a function argument or CLI flag.

This format is strictly defined in the workflow documentation at skills/webwright/reference/workflow.md.

final_script.py – The Canonical Automation Script

final_script.py is the authoritative Playwright script that solves the task, but it is never executed directly from the workspace root.

  • Configuration: The filename is fixed by the final_script_name: final_script.py setting in src/webwright/config/base.yaml.
  • Immutability: During execution, the environment copies the script into each finalruns/run_<id>/ folder rather than editing it in place.
  • CLI-Tool Requirements: In CLI-tool mode, the script must expose a single reusable function whose signature mirrors the # Parameters table from plan.md, wrapped with an argparse-based entry point. The required shape is documented in skills/webwright/reference/cli_tool_mode.md.

finalruns/ – Per-Run Execution Artifacts

The finalruns/ directory contains isolated execution environments where the canonical script is tested and validated.

Folder Naming Convention: Subdirectories follow the pattern run_001, run_002, etc., where the integer must always be larger than any existing run folder.

Contents of Each Run Folder (enforced by LocalWorkspaceEnvironment._capture_observation in src/webwright/environments/local_workspace.py):

  • final_script.py: A snapshot of the canonical script for reproducibility.
  • final_script_log.txt: A line-by-line action log printed during execution.
  • screenshots/: One PNG or JPG per critical point, named final_execution_<step>_<action>.png.
  • self_reflect_result.json: The JSON verdict from the self-reflection tool, where the predicted_label must be 1 for the run to be considered successful.

Supporting Files and Configuration

Beyond the core workflow files, the workspace contains several auxiliary artifacts:

File Purpose Source Reference
task.json Serialized description of the original task written by the agent src/webwright/environments/local_workspace.py
self_reflect_config.json Prompt templates for the self-reflection tool (generated once) src/webwright/config/base.yaml
.tmp/ Temporary directory for intermediate files src/webwright/environments/local_workspace.py
steps/ & logs/ Bash step files and raw console logs captured during agent actions src/webwright/environments/local_workspace.py

Configuration files such as task_showcase.yaml and crafted_cli.yaml in src/webwright/config/ repeatedly reference this layout to enforce the workspace contract.

Creating and Populating a Workspace

Initializing the Environment

import json
from pathlib import Path
from webwright.environments.local_workspace import LocalWorkspaceEnvironment

# Initialise the environment – this creates the workspace directory.

env = LocalWorkspaceEnvironment()
env.prepare(task_name="Find cheapest flight", start_url="https://example.com")

workspace = Path(env._workspace_dir())

# Write a minimal plan.md.

plan_md = """# Task

Find the cheapest round-trip flight from JFK to LAX on 2026-07-01.

# Critical Points

- [ ] CP1: Set origin airport to JFK
- [ ] CP2: Set destination airport to LAX
- [ ] CP3: Select departure date 2026-07-01
- [ ] CP4: Apply "cheapest" sort order
"""

(plan_md_path := workspace / "plan.md").write_text(plan_md)
print(f"Created {plan_md_path}")

The prepare() method automatically populates task.json and creates auxiliary directories such as screenshots/, steps/, and logs/.

Authoring the Canonical Script


# workspace/final_script.py

import argparse
from playwright.sync_api import sync_playwright

def run_flight_search(origin: str = "JFK", destination: str = "LAX", depart_date: str = "2026-07-01"):
    """
    Args:
        origin: Airport code for departure.
        destination: Airport code for arrival.
        depart_date: Date in YYYY-MM-DD format.
    """
    with sync_playwright() as p:
        browser = p.chromium.launch(headless=True)
        page = browser.new_page()
        page.goto("https://example.com/flights")
        # … Fill the form, click search, take screenshots …

        page.screenshot(path="final_execution_1_set_origin.png")
        # (More actions, more screenshots per CP)

        print("DONE")
        browser.close()

if __name__ == "__main__":
    parser = argparse.ArgumentParser()
    parser.add_argument("--origin", default="JFK")
    parser.add_argument("--destination", default="LAX")
    parser.add_argument("--depart-date", default="2026-07-01")
    args = parser.parse_args()
    run_flight_search(args.origin, args.destination, args.depart_date)

When the agent triggers execution, the environment copies this file into finalruns/run_001/ before launching the browser.

Executing and Verifying Runs


# Assume you are in the workspace root.

RUN_DIR=$(ls -d finalruns/run_* | sort -V | tail -n1)   # pick the latest run id

python "$RUN_DIR/final_script.py"                     # runs with default args

cat "$RUN_DIR/final_script_log.txt"                  # view the action log

ls "$RUN_DIR/screenshots"                             # verify one screenshot per CP

The workflow marks the task as complete only when every CP in plan.md is ticked (- [x]) and the self-reflection tool returns predicted_label: 1 in self_reflect_result.json.

Summary

  • The Webwright workspace structure is enforced by LocalWorkspaceEnvironment in src/webwright/environments/local_workspace.py and follows a strict contract for reproducible automations.
  • plan.md acts as a markdown checklist of critical points (CPs) that define task success criteria.
  • final_script.py is the immutable canonical script that gets copied into versioned finalruns/run_<id>/ folders for execution, never modified in place.
  • finalruns/ contains isolated execution contexts with screenshots, logs, and self_reflect_result.json verdicts, where the self-reflection tool validates that predicted_label equals 1.
  • Configuration files in src/webwright/config/ and reference documentation in skills/webwright/reference/ codify these relationships for CLI-tool and showcase modes.

Frequently Asked Questions

What role does plan.md play in the Webwright workspace?

plan.md serves as the executable specification for the task. It contains a # Critical Points section with markdown task list items (- [ ]) that the agent must tick off with concrete evidence. In CLI-tool mode, it also defines a # Parameters table that maps critical points to function arguments, as specified in skills/webwright/reference/workflow.md.

Why is final_script.py copied into finalruns/ instead of running from the root?

The LocalWorkspaceEnvironment copies final_script.py into each finalruns/run_<id>/ folder to ensure reproducibility and isolation. Each run captures its own snapshot of the script, preventing cross-run contamination and preserving the exact code that generated specific screenshots and logs. This immutability constraint is enforced by the environment's execution logic in src/webwright/environments/local_workspace.py.

How does the self-reflection tool use the finalruns/ directory?

After a script execution completes, the self-reflection tool (implemented in src/webwright/tools/self_reflection.py) consumes the screenshots/ and final_script_log.txt within a specific finalruns/run_<id>/ folder to produce a self_reflect_result.json verdict. The workspace is considered validated only when this JSON contains "predicted_label": 1, indicating that the visual evidence matches the requirements listed in plan.md.

Where is the workspace structure defined in the Microsoft Webwright source code?

The directory layout and file naming conventions are primarily defined in src/webwright/environments/local_workspace.py, specifically within the prepare() method and the _capture_observation() method. Default filenames and configurations are set in src/webwright/config/base.yaml, while the workflow semantics are documented in skills/webwright/reference/workflow.md and skills/webwright/reference/cli_tool_mode.md.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →