How the Self‑Reflection Tool Validates Task Completion Using `predicted_label`
The Self‑Reflection tool parses the LLM’s final Status: line to set predicted_label to 1 (success) or 0 (failure), which downstream agents use as the definitive signal for task completion.
The Microsoft Webwright repository provides an autonomous agent framework where the self_reflection tool automates task validation without human intervention. By analyzing screenshots through a two‑stage LLM judge workflow, the tool converts the model’s natural‑language verdict into a machine‑readable binary flag.
Two‑Stage Screenshot Judge Workflow
The validation logic lives in src/webwright/tools/self_reflection.py and follows a strict separation between evidence gathering and final classification.
Stage 1 – Per‑Image Scoring
For every screenshot provided, the tool sends a request to the configured LLM with a system prompt and a user prompt containing the image. The model returns a structured assessment containing:
- Score – An integer from 1 to 5 indicating quality or completeness.
- Reasoning – A short textual explanation justifying the score.
These individual assessments are collected but do not yet determine the final outcome.
Stage 2 – Final Verdict
All per‑image reasonings are inserted into a caller‑provided template using the {image_reasonings} placeholder. The tool then sends a second request to the LLM with all screenshots attached simultaneously. The model must conclude its response with a strict metadata line:
Status: success
or
Status: failure
This line acts as the single source of truth for the entire task.
Parsing the Status Line into predicted_label
After receiving the final response, the tool performs deterministic parsing using the regular expression pattern ^Status:\s*(\w+) to extract the status word. The mapping to predicted_label follows a rigid schema documented in the tool’s docstring:
success→predicted_labelis set to1failure→predicted_labelis set to0- Unparsable or missing line →
predicted_labelis set tonull
This logic ensures that natural‑language output from the LLM is transformed into a consistent JSON field that automated agents can evaluate without ambiguity.
Exit Code Derivation and Validation Logic
The tool’s exit code is derived directly from predicted_label to integrate cleanly with shell pipelines and CI/CD systems:
- Exit code
0(PASS) whenpredicted_label == 1 - Exit code
1(FAIL) whenpredicted_label == 0ornull
Downstream controllers invoke the CLI and inspect the JSON output. A predicted_label of 1 authorizes the agent to proceed, while any other value triggers error handling or retry logic. According to the source documentation, the output JSON structure contains the final predicted_label, the model’s final_response, and per‑image metadata, allowing full auditability of the judgment.
Key Implementation Files
The validation pipeline spans several modules in the microsoft/Webwright codebase:
src/webwright/tools/self_reflection.py– Implements the two‑stage judge, regex parsing (^Status:\s*(\w+)), and thepredicted_labelassignment logic.src/webwright/models/base.py– Provides thetext_parthelper used to format messages sent to the LLM during both scoring stages.src/webwright/tools/_model_config.py– Loads runtime configuration including model selection and API credentials required to execute the reflection workflow.tests/unit/test_tool_model_routing.py– Contains unit tests that exercise the tool’s model routing and indirectly verifypredicted_labelhandling under various response scenarios.
Practical Usage Example
You can invoke the Self‑Reflection tool from the command line and programmatically evaluate the result using the predicted_label field.
Run the tool with a JSON configuration file:
python -m webwright.tools.self_reflection \
--config my_self_reflect_config.json \
--output result.json
Inspect the generated JSON to verify the task status:
{
"images": ["screenshot1.png", "screenshot2.png"],
"per_image": [
{"score": 5, "reasoning": "All elements visible"},
{"score": 4, "reasoning": "Minor layout shift"}
],
"final_response": "The task completed successfully.\n\nStatus: success",
"predicted_label": 1
}
Use the label in a downstream Python script to gate execution:
import json
import sys
with open('result.json') as f:
data = json.load(f)
if data['predicted_label'] == 1:
print("✅ Task validated successfully")
sys.exit(0)
else:
print("❌ Task validation failed")
sys.exit(1)
Summary
- The Self‑Reflection tool uses a two‑stage LLM judge (per‑image scoring followed by a final verdict) to assess task completion.
predicted_labelis derived by parsing the finalStatus:line with the regex^Status:\s*(\w+)and mappingsuccess/failureto1/0.- An exit code of 0 indicates PASS (
predicted_label == 1), while any other value results in exit code 1. - The implementation resides primarily in
src/webwright/tools/self_reflection.py, utilizing helpers fromsrc/webwright/models/base.pyand configuration fromsrc/webwright/tools/_model_config.py.
Frequently Asked Questions
What does predicted_label represent in the Self‑Reflection tool?
predicted_label is a machine‑readable integer flag where 1 indicates the LLM judged the task as successfully completed, 0 indicates failure, and null means the Status: line could not be parsed from the model’s response.
How is the predicted_label value generated from the LLM output?
The tool extracts the last line matching the pattern ^Status:\s*(\w+). If the captured word is success, the label becomes 1; if failure, it becomes 0. If no match is found, the label defaults to null.
What happens if the LLM forgets to include the Status: line?
When the Status: line is missing or malformed, predicted_label is set to null, and the tool exits with code 1, signaling an inconclusive result that should be treated as a failure by downstream automation.
Can I use the Self‑Reflection tool without modifying the source code?
Yes. The tool exposes a CLI interface in src/webwright/tools/self_reflection.py that accepts a JSON configuration file and writes results to a specified output path, allowing integration into existing CI/CD pipelines without code changes.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →