How the Webwright Agent Loop Works: Message Flow and Action Execution Explained

Webwright’s agent loop is a flat “prompt → observe → execute → repeat” cycle implemented in src/webwright/agents/default.py that drives an LLM-backed agent to control a Playwright browser through generated Python code.

The microsoft/Webwright project provides a lightweight, transparent framework for browser automation using large language models. At its heart lies the Webwright agent loop, implemented in the DefaultAgent class, which orchestrates a continuous cycle between the LLM, a browser environment, and a structured message history that drives multi-step task completion.

Core Architecture: The DefaultAgent Class

The agent loop logic resides in src/webwright/agents/default.py within the DefaultAgent class. When invoked via the CLI in src/webwright/run/cli.py or programmatically, the run() method initializes state and enters a persistent cycle that alternates between querying the language model and executing browser actions.

Initialization Phase

When DefaultAgent.run(task, …) is called, the agent prepares the conversation context by:

  1. Storing the task instruction and template variables.
  2. Clearing internal message lists, call counters (self.n_calls), and format-error counters.
  3. Rendering and injecting two initial prompts using Jinja templates:
self.add_messages(
    self.model.format_message(role="system", content=self._render_template(self.config.system_template)),
    self.model.format_message(role="user",   content=self._render_template(self.config.instance_template)),
)

This initialization establishes the system behavior and task-specific instructions before the main loop begins (see lines 46‑49 in default.py).

The Main Agent Loop

The core mechanism is a while True loop inside run() that continues until an exit condition is met:

while True:
    self.step()
    # … handle interrupt, save, maybe compact history …

    if self.messages[-1].get("role") == "exit":
        break

Each iteration of the Webwright agent loop consists of two distinct phases: the Query phase and the Execute phase.

Query Phase: Sending Messages to the LLM

The query() method (lines 95‑99) handles communication with the language model backend:

  • It validates that the step limit (self.config.step_limit) has not been exceeded.
  • It transmits the current transcript (self.messages) to the model via self.model.query(self.messages).
  • It increments self.n_calls and appends the model’s response to the conversation history.

This phase transforms the current state of observations into a model-generated plan encoded in the response’s extra metadata.

Action Execution Phase

The execute_actions() method (lines 103‑115) processes the model’s response and drives browser interaction. The flow follows three strict steps:

1. Completion Verification

If the model signals completion via extra["done"], the agent verifies any configured gating requirements (such as self-reflection checks) and returns an exit message:

if extra.get("done"):
    # … optional self‑reflection gate …

    return self.add_messages(self.model.format_message(role="exit", …))

2. Browser Action Execution

For each action in extra["actions"], the environment’s execute method runs the specified Python code:

outputs = [self.env.execute(action) for action in extra.get("actions", [])]

In the default configuration, self.env is an instance of LocalBrowserEnvironment (src/webwright/environments/local_browser.py). Each action contains a python_code string that executes inside an async Playwright session with access to page, context, browser, playwright, and the original task dictionary (see _execute_async in that file).

3. Observation Capture

After execution, the environment returns a structured observation dictionary containing:

  • Current URL and page title
  • ARIA accessibility snapshot
  • Screenshot file path
  • Original Python code, stdout/stderr, and browser console logs

The agent then converts this raw observation into model-friendly messages using self.model.format_observation_messages(), embedding JSON data and screenshot references into the transcript for the next loop iteration.

Observation Feedback and Template Management

After appending observation messages, the agent may re-attach the instance template or the current plan.md content (lines 28‑36 in default.py). This ensures the model maintains task context across multiple browser interactions without losing the original objective.

Debug Artifacts and History Management

Webwright provides comprehensive visibility into the agent loop through debug artifacts:

  • Step artifacts: After each iteration, _write_debug_step_artifact (lines 98‑139) writes a JSON file and Markdown summary containing the model’s thought process, executed code, and observation data.
  • History compaction: When summary_every_n_steps is configured, the agent periodically asks the LLM to generate a compact summary of previous interactions, replacing the full transcript with [system, summary] to maintain context window limits (lines 103‑119).

Termination Conditions

The loop terminates when:

  • The model emits a message with role="exit" (via done=True or agent raising LimitsExceeded)
  • The run() method returns a final payload containing the submission text, exit status, and model usage statistics

Running Webwright: CLI and Programmatic Examples

Command Line Execution

python -m webwright.run.cli \
    -c base.yaml -c model_openai.yaml \
    -t "Search for flights from SEA to JFK on 2026‑08‑15 to 2026‑08‑20" \
    --start-url https://www.google.com/flights \
    --task-id demo_openai \
    -o outputs/default

The CLI constructs a DefaultAgent, initializes the LocalBrowserEnvironment, and invokes agent.run(task).

Programmatic Usage

from webwright import Webwright, Model, Environment
from webwright.run.cli import load_config

# Initialize model and environment

model = Model.from_config("model_openai.yaml")
env = Environment.from_config("local_browser.yaml")

# Create agent

agent = Webwright.Agent(model=model, env=env)

# Execute task

result = agent.run(
    task="Book a hotel in Paris for 2 guests between 2026-09-01 and 2026-09-05"
)
print("Final answer:", result.get("final_response"))

Inspecting Debug Output

Each step generates detailed artifacts in outputs/default/debug/steps/:

{
  "step": 3,
  "thought": "I need to click the date picker …",
  "python_code": "await page.click('#date-picker')",
  "done": false,
  "outputs": [
    {
      "observation": {
        "url": "...",
        "screenshot_path": "outputs/default/screenshots/step_0003.png",
        "aria_snapshot": "...",
        "python_output": "",
        "console_output": ""
      }
    }
  ]
}

Summary

  • The Webwright agent loop in src/webwright/agents/default.py implements a simple but powerful "prompt → observe → execute" cycle.
  • Message flow moves from system/instance prompts through LLM generation to browser observation and back.
  • Action execution relies on the LocalBrowserEnvironment in src/webwright/environments/local_browser.py to run Python code snippets controlling Playwright.
  • The loop maintains transparency through debug artifacts in JSON/Markdown and manages long contexts via optional summarization.
  • Termination occurs when the model signals done=True or step limits are exceeded.

Frequently Asked Questions

What is the Webwright agent loop?

The Webwright agent loop is the core control flow in src/webwright/agents/default.py that alternates between querying a language model and executing browser actions. It maintains a transcript of messages that evolves with each observation, allowing the LLM to react to changing page states until the task is complete or limits are reached.

How does Webwright execute browser actions?

Webwright executes actions by passing Python code strings from the LLM to LocalBrowserEnvironment.execute() in src/webwright/environments/local_browser.py. These snippets run inside an async Playwright session with direct access to page, context, and browser objects, then return structured observations including screenshots and ARIA snapshots.

Where is the agent loop configured?

Configuration occurs through YAML files in src/webwright/config/ that define system prompts, instance templates, model credentials (OpenAI, Anthropic, OpenRouter), and browser settings. The CLI entry point in src/webwright/run/cli.py loads these configurations to instantiate the DefaultAgent with the appropriate templates and environment.

How does Webwright handle long conversation histories?

When summary_every_n_steps is configured, the agent triggers a compaction phase where the LLM summarizes the current transcript. The agent then replaces the message history with a fresh system message containing only that summary, effectively resetting the context window while preserving task-relevant information (see lines 103‑119 in default.py).

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →