How to Integrate Webwright with OpenAI Codex as a Plugin: Complete Developer Guide
To integrate Webwright with OpenAI Codex, configure the .codex-plugin/plugin.json manifest to expose the browser automation skills directory, allowing Codex to discover and invoke the run command that orchestrates Playwright execution through the OpenAI Responses API backend.
The microsoft/Webwright repository ships with a Codex-compatible plugin that transforms OpenAI Codex into a state-of-the-art browser agent. By integrating Webwright with OpenAI Codex as a plugin, you enable the model to drive local Playwright browsers through natural language requests, capture screenshots, and self-verify results without leaving the conversational interface.
Understanding the Plugin Architecture
The integration relies on three core components: a JSON manifest that describes the plugin capabilities, markdown-based skill definitions that Codex parses into function schemas, and a Python runtime that bridges the LLM with Playwright execution.
The Plugin Manifest (.codex-plugin/plugin.json)
Codex discovers Webwright through the manifest located at .codex-plugin/plugin.json in the repository root. This file declares the plugin as a "productivity" tool with read and write capabilities, pointing Codex to the skills directory where command definitions reside.
{
"name": "webwright",
"version": "0.1.0",
"description": "Turn your coding agent into a SOTA browser agent …",
"interface": {
"displayName": "Webwright",
"shortDescription": "SOTA browser agent driven by your coding agent",
"capabilities": ["Read", "Write"],
"websiteURL": "https://github.com/microsoft/Webwright"
},
"skills": "./skills/"
}
The interface.capabilities array tells Codex that Webwright can both retrieve web content and persist files (screenshots, logs), while the skills field directs the plugin loader to parse the markdown catalog at skills/webwright/.
Skill Definitions and the Run Command
Each skill is defined in human-readable markdown files that Codex converts into callable function schemas. The primary entry point is the run command, defined in skills/webwright/commands/run.md, which specifies the contract for executing one-shot web tasks. When Codex loads the plugin, it indexes these commands—such as run and craft—making them available as tools the model can invoke during conversations.
The OpenAI Model Backend
Webwright uses src/webwright/models/openai_model.py to handle all LLM interactions during plugin execution. The OpenAIModel class extends the base model interface and specifically wraps the OpenAI Responses API. Key methods include:
_serialize_response_input: Converts message histories into the Responses API format required by OpenAI_extract_response_text: Parses the JSON-schema-constrained payload to retrieve the LLM's textual output_usage_metrics_from_response_payload: Extracts token consumption data for monitoring and billing
This backend is what the plugin runtime calls when generating Playwright scripts or verifying execution results.
Step-by-Step Integration Process
Integrating Webwright requires configuring the manifest, ensuring the runtime environment is available, and establishing the communication flow between Codex and the local browser environment.
-
Clone and prepare the repository: Ensure
microsoft/Webwrightis available on your host where the plugin will execute. -
Verify the manifest path: Confirm that
.codex-plugin/plugin.jsonis accessible at the repository root, as Codex will fetch this to discover capabilities. -
Configure model settings: Adjust
src/webwright/config/model_openai.yamlto specify your OpenAI model endpoint, API keys, and token limits. -
Enable plugin loading in Codex: Configure your Codex client to load the Webwright manifest URL, allowing the model to index the skill commands.
-
Initialize execution environment: Ensure Playwright browsers (Firefox/Chromium) are installed, as
src/webwright/environments/local_browser.pywill spawn sandboxed browser sessions during task execution.
Code Example: Triggering Webwright from Codex
To activate the plugin from an OpenAI Codex session, send a JSON payload that references the manifest and declares the run function. Below is a minimal request structure compatible with the OpenAI SDK or direct HTTP calls:
{
"model": "gpt-4o",
"messages": [
{
"role": "system",
"content": "You have access to the Webwright plugin."
},
{
"role": "user",
"content": "Find the current price of the Apple iPhone 15 on amazon.com and give me the exact figure."
}
],
"plugins": [
{
"type": "manifest",
"url": "https://github.com/microsoft/Webwright/blob/main/.codex-plugin/plugin.json"
}
],
"functions": [
{
"name": "run",
"description": "Execute a web-task using the Webwright workflow.",
"parameters": {
"type": "object",
"properties": {
"task": {
"type": "string",
"description": "Natural-language description of the web task."
}
},
"required": ["task"]
}
}
],
"function_call": "auto"
}
When Codex processes this request, it identifies the run function from the parsed run.md skill definition, extracts the user's natural language query, and invokes the Webwright runtime with the task parameter.
Runtime Execution Flow
Once Codex calls the run function, src/webwright/run/cli.py orchestrates the end-to-end workflow:
-
Workspace creation: The CLI generates a temporary directory at
final_runs/run_<id>/to isolate the task execution. -
Plan generation: Using
OpenAIModel, the system writes aplan.mdfile based on the model's high-level outline for completing the task. -
Script instrumentation: The runtime generates
final_script.py, an instrumented Playwright script that implements the plan. -
Browser execution: The script executes against the local browser environment defined in
src/webwright/environments/local_browser.py, driving Firefox or Chromium through the planned steps. -
Artifact collection: Screenshots and execution logs are captured to
final_runs/run_<id>/, includingfinal_script_log.txtand PNG files. -
Self-verification: The
src/webwright/agents/default.pyagent feeds artifacts back to the OpenAI model to verify the extracted data before returning the final answer to Codex.
Configuration and Environment Setup
The default agent behavior is controlled by src/webwright/agents/default.py, which coordinates between the model backend and the browser environment. For OpenAI-specific settings, modify src/webwright/config/model_openai.yaml to adjust:
- Model name (e.g.,
gpt-4o,gpt-4o-mini) - Maximum token limits
- Endpoint URLs
- Temperature and sampling parameters
Ensure the host running the plugin has network access to both the OpenAI API and the target websites, as the OpenAIModel class makes synchronous API calls during plan generation and verification phases.
Summary
- Load the manifest: Point Codex to
.codex-plugin/plugin.jsonto discover Webwright's capabilities and skill commands. - Invoke skills naturally: Once loaded, Codex can call the
runfunction (defined inskills/webwright/commands/run.md) to execute browser tasks. - Leverage the OpenAI backend: The
OpenAIModelclass insrc/webwright/models/openai_model.pyhandles Responses API serialization and token tracking. - Orchestrate via CLI:
src/webwright/run/cli.pymanages workspace creation, Playwright execution, and artifact collection infinal_runs/run_<id>/directories. - Verify results: The default agent in
src/webwright/agents/default.pyenables self-correction by feeding screenshots and logs back to the model before returning final answers.
Frequently Asked Questions
What file does OpenAI Codex read to discover Webwright's capabilities?
Codex reads .codex-plugin/plugin.json from the repository root. This JSON manifest declares the plugin name, version, capabilities (Read/Write), and the relative path to the skills directory where command definitions are stored.
How does Webwright handle LLM interactions during plugin execution?
All LLM interactions route through src/webwright/models/openai_model.py. The OpenAIModel class serializes conversation history into the OpenAI Responses API format, extracts structured output from JSON-schema-constrained responses, and tracks token usage via the _usage_metrics_from_response_payload method.
Where does Webwright store screenshots and execution logs during a Codex plugin run?
The runtime creates isolated workspaces under final_runs/run_<id>/ for each task. This directory contains plan.md, final_script.py, final_script_log.txt, and any screenshots captured during Playwright execution, which can be retrieved after the run completes.
Can I use a different OpenAI model than the default for plugin tasks?
Yes. Modify src/webwright/config/model_openai.yaml to change the model name, token limits, and endpoint parameters. The OpenAIModel class reads this configuration at runtime, allowing you to specify alternative models like gpt-4o-mini or custom deployments without changing the core plugin code.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →