# How Webwright Handles History Compaction Using LLM Summarization After N Steps

> Learn how Webwright uses LLM summarization for history compaction. Discover how it bounds token usage by summarizing conversations every N steps with the system prompt.

- Repository: [Microsoft/Webwright](https://github.com/microsoft/Webwright)
- Tags: internals
- Published: 2026-06-25

---

**Webwright triggers history compaction every N agent calls by prompting the LLM to summarize the conversation, then replaces the full transcript with only the system prompt and that summary to bound token usage.**

Webwright, Microsoft's open-source web automation framework, implements an intelligent **history compaction using LLM summarization** mechanism to prevent token overflow during long-running agent sessions. This feature periodically condenses the conversation history after a configurable number of steps, ensuring that the model context window remains efficient without losing critical task context. The implementation relies on three core components defined in the source: the `summary_every_n_steps` configuration, the `summary_user_prompt` template, and the `_compact_history()` method.

## What Triggers History Compaction?

Webwright's compaction logic operates automatically during the agent's execution loop, monitoring call counts against a user-defined threshold.

### The Configuration Parameter

The trigger frequency is controlled by **`summary_every_n_steps`**, defined in [[`src/webwright/config/base.yaml`](https://github.com/microsoft/Webwright/blob/main/src/webwright/config/base.yaml)](https://github.com/microsoft/Webwright/blob/main/src/webwright/config/base.yaml#L90-L92). By default, this value is set to `20`, meaning compaction occurs after every 20 agent calls. This configuration is inherited by specialized profiles such as [`local_browser.yaml`](https://github.com/microsoft/Webwright/blob/main/local_browser.yaml), ensuring consistent behavior across different agent types unless explicitly overridden.

### The Trigger Condition

During execution, the agent's `run()` method maintains a counter (`self.n_calls`) that increments with each LLM interaction. When `summary_every_n_steps` is greater than zero and the counter is a multiple of that value, the system invokes the compaction routine:

```python
if (
    self.config.summary_every_n_steps > 0
    and self.n_calls > 0
    and self.n_calls % self.config.summary_every_n_steps == 0
):
    self._compact_history()

```

This check appears in [[`src/webwright/agents/default.py`](https://github.com/microsoft/Webwright/blob/main/src/webwright/agents/default.py)](https://github.com/microsoft/Webwright/blob/main/src/webwright/agents/default.py) within the main execution loop (lines 74-79), ensuring compaction happens predictably without interrupting active task processing.

## How the Compaction Process Works

The `_compact_history()` method (lines 303-340 in [`default.py`](https://github.com/microsoft/Webwright/blob/main/default.py)) executes a five-phase pipeline to compress the message history while preserving essential context.

### 1. Locating the System Message

The method first identifies and preserves the initial **system message**, which contains the agent's core instructions and personality. This message is the only component from the original transcript that survives the compaction process unchanged.

### 2. Constructing the Summarization Request

Webwright creates a specialized user-role message containing the **`summary_user_prompt`** and a metadata flag marking it as a compaction request:

```python
summary_request = self.model.format_message(
    role="user",
    content=self.config.summary_user_prompt,
    extra={"interrupt_type": "HistoryCompactionRequest"},
)
summary_messages = list(self.messages) + [summary_request]

```

The **`summary_user_prompt`** (defined in lines 16-29 of [`default.py`](https://github.com/microsoft/Webwright/blob/main/default.py)) instructs the LLM to produce a structured summary that includes the original task goal, critical constraints, workspace details, key findings, current state, and proposed next action, explicitly forbidding code generation or termination signals.

### 3. Model Invocation and Error Handling

The agent sends the complete message history plus the summarization request to the LLM:

```python
response = self.model.query(summary_messages)

```

If the query fails or returns an error, the method returns early without modifying the message history (lines 23-24), ensuring that transient model failures do not corrupt the agent state or terminate the run.

### 4. Extracting the Summary Content

The method extracts the summary text from the model response's `content` field. If this field is empty, it falls back to any `final_response` present in the response's `extra` payload. As a safety mechanism, if both sources are empty, the system uses the string `"(empty summary)"` to prevent null values from entering the conversation history (lines 25-29).

### 5. Reconstructing the Message Buffer

Finally, Webwright constructs a new user message containing the generated summary wrapped in clear demarcation headers, then resets the internal message list to contain only the system message and this new summary:

```python
summary_message = self.model.format_message(
    role="user",
    content=(
        "## Compacted History Summary\n"

        f"(context was compacted after step {self.n_calls}; earlier turns have been replaced "
        "by the summary below)\n\n"
        f"{summary_text}\n\n## End of Compacted Summary"

    ),
    extra={"interrupt_type": "HistoryCompactionSummary"},
)
self.messages = [system_message, summary_message]

```

This operation dramatically reduces token usage by replacing potentially hundreds of detailed interaction turns with a single concise summary. Immediately following compaction, the `run()` loop calls `save()` again (line 80) to persist the trimmed transcript to the configured output path.

## Configuring History Compaction for Your Use Case

Webwright exposes the compaction behavior through its YAML configuration system, allowing customization without modifying source code.

### Adjusting the Step Frequency

To change how often compaction occurs, override the `summary_every_n_steps` value in your configuration file or at runtime:

```python
from webwright import load_config, WebWright

cfg = load_config("src/webwright/config/base.yaml")
cfg.summary_every_n_steps = 10  # Compact every 10 steps instead of 20

agent = WebWright(model=my_llm, env=my_env, config=cfg)

```

Setting this value to `0` disables automatic compaction entirely, though this risks exceeding model context limits during extended sessions.

### Customizing the Summary Prompt

You can tailor the summarization instructions to emphasize domain-specific details or reduce output length. The `summary_user_prompt` accepts any valid instruction string:

```python
custom_prompt = """
Summarize the previous interaction for a token-limited model. Include:
- The original goal
- URLs visited
- Last successful action
- Outstanding blockers
Write in plain prose. Do not include code or set done=true.
"""
cfg.summary_user_prompt = custom_prompt

```

This customization is merged into the configuration using the `recursive_merge` utility from [[`src/webwright/utils/serialize.py`](https://github.com/microsoft/Webwright/blob/main/src/webwright/utils/serialize.py)](https://github.com/microsoft/Webwright/blob/main/src/webwright/utils/serialize.py), ensuring that partial prompt overrides integrate correctly with base templates.

## Summary

- **Webwright** implements history compaction via the **`_compact_history()`** method in [`src/webwright/agents/default.py`](https://github.com/microsoft/Webwright/blob/main/src/webwright/agents/default.py), triggered every N steps defined by `summary_every_n_steps`.
- The default trigger threshold is **20 steps**, configurable in [`src/webwright/config/base.yaml`](https://github.com/microsoft/Webwright/blob/main/src/webwright/config/base.yaml).
- The process preserves the **system message**, requests a structured summary using **`summary_user_prompt`**, and replaces the full transcript with `[system, summary]`.
- Failed summarization attempts **fail gracefully** without corrupting the agent state.
- Compaction behavior is fully customizable through YAML configuration or direct runtime parameter modification.

## Frequently Asked Questions

### What happens if the LLM fails to generate a summary during compaction?

According to the implementation in [`src/webwright/agents/default.py`](https://github.com/microsoft/Webwright/blob/main/src/webwright/agents/default.py) (lines 23-24), if the `model.query()` call raises an exception or returns an error response, the `_compact_history()` method returns early without modifying the message buffer. This ensures that transient LLM failures do not crash the agent or result in data loss, allowing the conversation to continue with the full history intact.

### How does Webwright ensure critical context isn't lost when compacting history?

The **`summary_user_prompt`** (lines 16-29 in [`default.py`](https://github.com/microsoft/Webwright/blob/main/default.py)) explicitly instructs the model to retain essential metadata including the original task goal, critical constraints, workspace details, key findings, current state, and next proposed action. Additionally, the **system message** containing core agent instructions is always preserved separately and never included in the summarization process, ensuring behavioral consistency even after compaction.

### Can I disable history compaction entirely?

Yes. Setting **`summary_every_n_steps`** to `0` in your configuration file disables the automatic compaction trigger. When this value is zero or negative, the modulo check in the `run()` loop evaluates to false, preventing `_compact_history()` from ever executing. Note that disabling compaction may lead to context window exhaustion during long-running tasks.

### Where does Webwright store the compacted conversation history?

Immediately after the `_compact_history()` method completes and updates `self.messages`, the `run()` loop invokes **`save()`** (line 80) to persist the newly compacted transcript—containing only the system message and summary—to the configured output path. This ensures that checkpoints, debugging tools, and downstream consumers always see the same trimmed context that the LLM receives for subsequent steps.