How Webwright Handles History Compaction Using LLM Summarization After N Steps

Webwright triggers history compaction every N agent calls by prompting the LLM to summarize the conversation, then replaces the full transcript with only the system prompt and that summary to bound token usage.

Webwright, Microsoft's open-source web automation framework, implements an intelligent history compaction using LLM summarization mechanism to prevent token overflow during long-running agent sessions. This feature periodically condenses the conversation history after a configurable number of steps, ensuring that the model context window remains efficient without losing critical task context. The implementation relies on three core components defined in the source: the summary_every_n_steps configuration, the summary_user_prompt template, and the _compact_history() method.

What Triggers History Compaction?

Webwright's compaction logic operates automatically during the agent's execution loop, monitoring call counts against a user-defined threshold.

The Configuration Parameter

The trigger frequency is controlled by summary_every_n_steps, defined in [src/webwright/config/base.yaml](https://github.com/microsoft/Webwright/blob/main/src/webwright/config/base.yaml#L90-L92). By default, this value is set to 20, meaning compaction occurs after every 20 agent calls. This configuration is inherited by specialized profiles such as local_browser.yaml, ensuring consistent behavior across different agent types unless explicitly overridden.

The Trigger Condition

During execution, the agent's run() method maintains a counter (self.n_calls) that increments with each LLM interaction. When summary_every_n_steps is greater than zero and the counter is a multiple of that value, the system invokes the compaction routine:

if (
    self.config.summary_every_n_steps > 0
    and self.n_calls > 0
    and self.n_calls % self.config.summary_every_n_steps == 0
):
    self._compact_history()

This check appears in [src/webwright/agents/default.py](https://github.com/microsoft/Webwright/blob/main/src/webwright/agents/default.py) within the main execution loop (lines 74-79), ensuring compaction happens predictably without interrupting active task processing.

How the Compaction Process Works

The _compact_history() method (lines 303-340 in default.py) executes a five-phase pipeline to compress the message history while preserving essential context.

1. Locating the System Message

The method first identifies and preserves the initial system message, which contains the agent's core instructions and personality. This message is the only component from the original transcript that survives the compaction process unchanged.

2. Constructing the Summarization Request

Webwright creates a specialized user-role message containing the summary_user_prompt and a metadata flag marking it as a compaction request:

summary_request = self.model.format_message(
    role="user",
    content=self.config.summary_user_prompt,
    extra={"interrupt_type": "HistoryCompactionRequest"},
)
summary_messages = list(self.messages) + [summary_request]

The summary_user_prompt (defined in lines 16-29 of default.py) instructs the LLM to produce a structured summary that includes the original task goal, critical constraints, workspace details, key findings, current state, and proposed next action, explicitly forbidding code generation or termination signals.

3. Model Invocation and Error Handling

The agent sends the complete message history plus the summarization request to the LLM:

response = self.model.query(summary_messages)

If the query fails or returns an error, the method returns early without modifying the message history (lines 23-24), ensuring that transient model failures do not corrupt the agent state or terminate the run.

4. Extracting the Summary Content

The method extracts the summary text from the model response's content field. If this field is empty, it falls back to any final_response present in the response's extra payload. As a safety mechanism, if both sources are empty, the system uses the string "(empty summary)" to prevent null values from entering the conversation history (lines 25-29).

5. Reconstructing the Message Buffer

Finally, Webwright constructs a new user message containing the generated summary wrapped in clear demarcation headers, then resets the internal message list to contain only the system message and this new summary:

summary_message = self.model.format_message(
    role="user",
    content=(
        "## Compacted History Summary\n"

        f"(context was compacted after step {self.n_calls}; earlier turns have been replaced "
        "by the summary below)\n\n"
        f"{summary_text}\n\n## End of Compacted Summary"

    ),
    extra={"interrupt_type": "HistoryCompactionSummary"},
)
self.messages = [system_message, summary_message]

This operation dramatically reduces token usage by replacing potentially hundreds of detailed interaction turns with a single concise summary. Immediately following compaction, the run() loop calls save() again (line 80) to persist the trimmed transcript to the configured output path.

Configuring History Compaction for Your Use Case

Webwright exposes the compaction behavior through its YAML configuration system, allowing customization without modifying source code.

Adjusting the Step Frequency

To change how often compaction occurs, override the summary_every_n_steps value in your configuration file or at runtime:

from webwright import load_config, WebWright

cfg = load_config("src/webwright/config/base.yaml")
cfg.summary_every_n_steps = 10  # Compact every 10 steps instead of 20

agent = WebWright(model=my_llm, env=my_env, config=cfg)

Setting this value to 0 disables automatic compaction entirely, though this risks exceeding model context limits during extended sessions.

Customizing the Summary Prompt

You can tailor the summarization instructions to emphasize domain-specific details or reduce output length. The summary_user_prompt accepts any valid instruction string:

custom_prompt = """
Summarize the previous interaction for a token-limited model. Include:
- The original goal
- URLs visited
- Last successful action
- Outstanding blockers
Write in plain prose. Do not include code or set done=true.
"""
cfg.summary_user_prompt = custom_prompt

This customization is merged into the configuration using the recursive_merge utility from [src/webwright/utils/serialize.py](https://github.com/microsoft/Webwright/blob/main/src/webwright/utils/serialize.py), ensuring that partial prompt overrides integrate correctly with base templates.

Summary

  • Webwright implements history compaction via the _compact_history() method in src/webwright/agents/default.py, triggered every N steps defined by summary_every_n_steps.
  • The default trigger threshold is 20 steps, configurable in src/webwright/config/base.yaml.
  • The process preserves the system message, requests a structured summary using summary_user_prompt, and replaces the full transcript with [system, summary].
  • Failed summarization attempts fail gracefully without corrupting the agent state.
  • Compaction behavior is fully customizable through YAML configuration or direct runtime parameter modification.

Frequently Asked Questions

What happens if the LLM fails to generate a summary during compaction?

According to the implementation in src/webwright/agents/default.py (lines 23-24), if the model.query() call raises an exception or returns an error response, the _compact_history() method returns early without modifying the message buffer. This ensures that transient LLM failures do not crash the agent or result in data loss, allowing the conversation to continue with the full history intact.

How does Webwright ensure critical context isn't lost when compacting history?

The summary_user_prompt (lines 16-29 in default.py) explicitly instructs the model to retain essential metadata including the original task goal, critical constraints, workspace details, key findings, current state, and next proposed action. Additionally, the system message containing core agent instructions is always preserved separately and never included in the summarization process, ensuring behavioral consistency even after compaction.

Can I disable history compaction entirely?

Yes. Setting summary_every_n_steps to 0 in your configuration file disables the automatic compaction trigger. When this value is zero or negative, the modulo check in the run() loop evaluates to false, preventing _compact_history() from ever executing. Note that disabling compaction may lead to context window exhaustion during long-running tasks.

Where does Webwright store the compacted conversation history?

Immediately after the _compact_history() method completes and updates self.messages, the run() loop invokes save() (line 80) to persist the newly compacted transcript—containing only the system message and summary—to the configured output path. This ensures that checkpoints, debugging tools, and downstream consumers always see the same trimmed context that the LLM receives for subsequent steps.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →