# How the Self‑Reflection Tool Validates Task Completion Using `predicted_label`

> Learn how the Webwright Self-Reflection tool validates task completion using the predicted_label signal. Discover how it determines success or failure for downstream agents.

- Repository: [Microsoft/Webwright](https://github.com/microsoft/Webwright)
- Tags: deep-dive
- Published: 2026-06-25

---

**The Self‑Reflection tool parses the LLM’s final `Status:` line to set `predicted_label` to `1` (success) or `0` (failure), which downstream agents use as the definitive signal for task completion.**

The Microsoft **Webwright** repository provides an autonomous agent framework where the `self_reflection` tool automates task validation without human intervention. By analyzing screenshots through a two‑stage LLM judge workflow, the tool converts the model’s natural‑language verdict into a machine‑readable binary flag.

## Two‑Stage Screenshot Judge Workflow

The validation logic lives in [`src/webwright/tools/self_reflection.py`](https://github.com/microsoft/Webwright/blob/main/src/webwright/tools/self_reflection.py) and follows a strict separation between evidence gathering and final classification.

### Stage 1 – Per‑Image Scoring

For every screenshot provided, the tool sends a request to the configured LLM with a system prompt and a user prompt containing the image. The model returns a structured assessment containing:

- **Score** – An integer from 1 to 5 indicating quality or completeness.
- **Reasoning** – A short textual explanation justifying the score.

These individual assessments are collected but do not yet determine the final outcome.

### Stage 2 – Final Verdict

All per‑image reasonings are inserted into a caller‑provided template using the `{image_reasonings}` placeholder. The tool then sends a second request to the LLM with **all** screenshots attached simultaneously. The model must conclude its response with a strict metadata line:

```

Status: success

```

or

```

Status: failure

```

This line acts as the single source of truth for the entire task.

## Parsing the Status Line into `predicted_label`

After receiving the final response, the tool performs deterministic parsing using the regular expression pattern `^Status:\s*(\w+)` to extract the status word. The mapping to `predicted_label` follows a rigid schema documented in the tool’s docstring:

- **`success`** → `predicted_label` is set to **`1`**
- **`failure`** → `predicted_label` is set to **`0`**
- **Unparsable or missing line** → `predicted_label` is set to **`null`**

This logic ensures that natural‑language output from the LLM is transformed into a consistent JSON field that automated agents can evaluate without ambiguity.

## Exit Code Derivation and Validation Logic

The tool’s exit code is derived directly from `predicted_label` to integrate cleanly with shell pipelines and CI/CD systems:

- Exit code **`0`** (PASS) when `predicted_label == 1`
- Exit code **`1`** (FAIL) when `predicted_label == 0` or `null`

Downstream controllers invoke the CLI and inspect the JSON output. A `predicted_label` of `1` authorizes the agent to proceed, while any other value triggers error handling or retry logic. According to the source documentation, the output JSON structure contains the final `predicted_label`, the model’s `final_response`, and per‑image metadata, allowing full auditability of the judgment.

## Key Implementation Files

The validation pipeline spans several modules in the `microsoft/Webwright` codebase:

- **[`src/webwright/tools/self_reflection.py`](https://github.com/microsoft/Webwright/blob/main/src/webwright/tools/self_reflection.py)** – Implements the two‑stage judge, regex parsing (`^Status:\s*(\w+)`), and the `predicted_label` assignment logic.
- **[`src/webwright/models/base.py`](https://github.com/microsoft/Webwright/blob/main/src/webwright/models/base.py)** – Provides the `text_part` helper used to format messages sent to the LLM during both scoring stages.
- **[`src/webwright/tools/_model_config.py`](https://github.com/microsoft/Webwright/blob/main/src/webwright/tools/_model_config.py)** – Loads runtime configuration including model selection and API credentials required to execute the reflection workflow.
- **[`tests/unit/test_tool_model_routing.py`](https://github.com/microsoft/Webwright/blob/main/tests/unit/test_tool_model_routing.py)** – Contains unit tests that exercise the tool’s model routing and indirectly verify `predicted_label` handling under various response scenarios.

## Practical Usage Example

You can invoke the Self‑Reflection tool from the command line and programmatically evaluate the result using the `predicted_label` field.

Run the tool with a JSON configuration file:

```bash
python -m webwright.tools.self_reflection \
    --config my_self_reflect_config.json \
    --output result.json

```

Inspect the generated JSON to verify the task status:

```json
{
  "images": ["screenshot1.png", "screenshot2.png"],
  "per_image": [
    {"score": 5, "reasoning": "All elements visible"},
    {"score": 4, "reasoning": "Minor layout shift"}
  ],
  "final_response": "The task completed successfully.\n\nStatus: success",
  "predicted_label": 1
}

```

Use the label in a downstream Python script to gate execution:

```python
import json
import sys

with open('result.json') as f:
    data = json.load(f)

if data['predicted_label'] == 1:
    print("✅ Task validated successfully")
    sys.exit(0)
else:
    print("❌ Task validation failed")
    sys.exit(1)

```

## Summary

- The Self‑Reflection tool uses a **two‑stage LLM judge** (per‑image scoring followed by a final verdict) to assess task completion.
- **`predicted_label`** is derived by parsing the final `Status:` line with the regex `^Status:\s*(\w+)` and mapping `success`/`failure` to `1`/`0`.
- An **exit code of 0** indicates PASS (`predicted_label == 1`), while any other value results in exit code 1.
- The implementation resides primarily in [`src/webwright/tools/self_reflection.py`](https://github.com/microsoft/Webwright/blob/main/src/webwright/tools/self_reflection.py), utilizing helpers from [`src/webwright/models/base.py`](https://github.com/microsoft/Webwright/blob/main/src/webwright/models/base.py) and configuration from [`src/webwright/tools/_model_config.py`](https://github.com/microsoft/Webwright/blob/main/src/webwright/tools/_model_config.py).

## Frequently Asked Questions

### What does `predicted_label` represent in the Self‑Reflection tool?

`predicted_label` is a machine‑readable integer flag where `1` indicates the LLM judged the task as successfully completed, `0` indicates failure, and `null` means the `Status:` line could not be parsed from the model’s response.

### How is the `predicted_label` value generated from the LLM output?

The tool extracts the last line matching the pattern `^Status:\s*(\w+)`. If the captured word is `success`, the label becomes `1`; if `failure`, it becomes `0`. If no match is found, the label defaults to `null`.

### What happens if the LLM forgets to include the `Status:` line?

When the `Status:` line is missing or malformed, `predicted_label` is set to `null`, and the tool exits with code 1, signaling an inconclusive result that should be treated as a failure by downstream automation.

### Can I use the Self‑Reflection tool without modifying the source code?

Yes. The tool exposes a CLI interface in [`src/webwright/tools/self_reflection.py`](https://github.com/microsoft/Webwright/blob/main/src/webwright/tools/self_reflection.py) that accepts a JSON configuration file and writes results to a specified output path, allowing integration into existing CI/CD pipelines without code changes.