How to Integrate Webwright with OpenAI Codex as a Plugin: Complete Developer Guide

To integrate Webwright with OpenAI Codex, configure the .codex-plugin/plugin.json manifest to expose the browser automation skills directory, allowing Codex to discover and invoke the run command that orchestrates Playwright execution through the OpenAI Responses API backend.

The microsoft/Webwright repository ships with a Codex-compatible plugin that transforms OpenAI Codex into a state-of-the-art browser agent. By integrating Webwright with OpenAI Codex as a plugin, you enable the model to drive local Playwright browsers through natural language requests, capture screenshots, and self-verify results without leaving the conversational interface.

Understanding the Plugin Architecture

The integration relies on three core components: a JSON manifest that describes the plugin capabilities, markdown-based skill definitions that Codex parses into function schemas, and a Python runtime that bridges the LLM with Playwright execution.

The Plugin Manifest (.codex-plugin/plugin.json)

Codex discovers Webwright through the manifest located at .codex-plugin/plugin.json in the repository root. This file declares the plugin as a "productivity" tool with read and write capabilities, pointing Codex to the skills directory where command definitions reside.

{
  "name": "webwright",
  "version": "0.1.0",
  "description": "Turn your coding agent into a SOTA browser agent …",
  "interface": {
    "displayName": "Webwright",
    "shortDescription": "SOTA browser agent driven by your coding agent",
    "capabilities": ["Read", "Write"],
    "websiteURL": "https://github.com/microsoft/Webwright"
  },
  "skills": "./skills/"
}

The interface.capabilities array tells Codex that Webwright can both retrieve web content and persist files (screenshots, logs), while the skills field directs the plugin loader to parse the markdown catalog at skills/webwright/.

Skill Definitions and the Run Command

Each skill is defined in human-readable markdown files that Codex converts into callable function schemas. The primary entry point is the run command, defined in skills/webwright/commands/run.md, which specifies the contract for executing one-shot web tasks. When Codex loads the plugin, it indexes these commands—such as run and craft—making them available as tools the model can invoke during conversations.

The OpenAI Model Backend

Webwright uses src/webwright/models/openai_model.py to handle all LLM interactions during plugin execution. The OpenAIModel class extends the base model interface and specifically wraps the OpenAI Responses API. Key methods include:

  • _serialize_response_input: Converts message histories into the Responses API format required by OpenAI
  • _extract_response_text: Parses the JSON-schema-constrained payload to retrieve the LLM's textual output
  • _usage_metrics_from_response_payload: Extracts token consumption data for monitoring and billing

This backend is what the plugin runtime calls when generating Playwright scripts or verifying execution results.

Step-by-Step Integration Process

Integrating Webwright requires configuring the manifest, ensuring the runtime environment is available, and establishing the communication flow between Codex and the local browser environment.

  1. Clone and prepare the repository: Ensure microsoft/Webwright is available on your host where the plugin will execute.

  2. Verify the manifest path: Confirm that .codex-plugin/plugin.json is accessible at the repository root, as Codex will fetch this to discover capabilities.

  3. Configure model settings: Adjust src/webwright/config/model_openai.yaml to specify your OpenAI model endpoint, API keys, and token limits.

  4. Enable plugin loading in Codex: Configure your Codex client to load the Webwright manifest URL, allowing the model to index the skill commands.

  5. Initialize execution environment: Ensure Playwright browsers (Firefox/Chromium) are installed, as src/webwright/environments/local_browser.py will spawn sandboxed browser sessions during task execution.

Code Example: Triggering Webwright from Codex

To activate the plugin from an OpenAI Codex session, send a JSON payload that references the manifest and declares the run function. Below is a minimal request structure compatible with the OpenAI SDK or direct HTTP calls:

{
  "model": "gpt-4o",
  "messages": [
    {
      "role": "system",
      "content": "You have access to the Webwright plugin."
    },
    {
      "role": "user",
      "content": "Find the current price of the Apple iPhone 15 on amazon.com and give me the exact figure."
    }
  ],
  "plugins": [
    {
      "type": "manifest",
      "url": "https://github.com/microsoft/Webwright/blob/main/.codex-plugin/plugin.json"
    }
  ],
  "functions": [
    {
      "name": "run",
      "description": "Execute a web-task using the Webwright workflow.",
      "parameters": {
        "type": "object",
        "properties": {
          "task": {
            "type": "string",
            "description": "Natural-language description of the web task."
          }
        },
        "required": ["task"]
      }
    }
  ],
  "function_call": "auto"
}

When Codex processes this request, it identifies the run function from the parsed run.md skill definition, extracts the user's natural language query, and invokes the Webwright runtime with the task parameter.

Runtime Execution Flow

Once Codex calls the run function, src/webwright/run/cli.py orchestrates the end-to-end workflow:

  1. Workspace creation: The CLI generates a temporary directory at final_runs/run_<id>/ to isolate the task execution.

  2. Plan generation: Using OpenAIModel, the system writes a plan.md file based on the model's high-level outline for completing the task.

  3. Script instrumentation: The runtime generates final_script.py, an instrumented Playwright script that implements the plan.

  4. Browser execution: The script executes against the local browser environment defined in src/webwright/environments/local_browser.py, driving Firefox or Chromium through the planned steps.

  5. Artifact collection: Screenshots and execution logs are captured to final_runs/run_<id>/, including final_script_log.txt and PNG files.

  6. Self-verification: The src/webwright/agents/default.py agent feeds artifacts back to the OpenAI model to verify the extracted data before returning the final answer to Codex.

Configuration and Environment Setup

The default agent behavior is controlled by src/webwright/agents/default.py, which coordinates between the model backend and the browser environment. For OpenAI-specific settings, modify src/webwright/config/model_openai.yaml to adjust:

  • Model name (e.g., gpt-4o, gpt-4o-mini)
  • Maximum token limits
  • Endpoint URLs
  • Temperature and sampling parameters

Ensure the host running the plugin has network access to both the OpenAI API and the target websites, as the OpenAIModel class makes synchronous API calls during plan generation and verification phases.

Summary

Frequently Asked Questions

What file does OpenAI Codex read to discover Webwright's capabilities?

Codex reads .codex-plugin/plugin.json from the repository root. This JSON manifest declares the plugin name, version, capabilities (Read/Write), and the relative path to the skills directory where command definitions are stored.

How does Webwright handle LLM interactions during plugin execution?

All LLM interactions route through src/webwright/models/openai_model.py. The OpenAIModel class serializes conversation history into the OpenAI Responses API format, extracts structured output from JSON-schema-constrained responses, and tracks token usage via the _usage_metrics_from_response_payload method.

Where does Webwright store screenshots and execution logs during a Codex plugin run?

The runtime creates isolated workspaces under final_runs/run_<id>/ for each task. This directory contains plan.md, final_script.py, final_script_log.txt, and any screenshots captured during Playwright execution, which can be retrieved after the run completes.

Can I use a different OpenAI model than the default for plugin tasks?

Yes. Modify src/webwright/config/model_openai.yaml to change the model name, token limits, and endpoint parameters. The OpenAIModel class reads this configuration at runtime, allowing you to specify alternative models like gpt-4o-mini or custom deployments without changing the core plugin code.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →