# How the Evidence Collector Validates Visual QA with Screenshot Requirements in Agency Agents

> Learn how the Evidence Collector validates visual QA with screenshot requirements by capturing Playwright screenshots, comparing them to specs, and auto-failing builds on visual discrepancies.

- Repository: [Michael Sitarzewski/agency-agents](https://github.com/msitarzewski/agency-agents)
- Tags: deep-dive
- Published: 2026-03-09

---

**The Evidence Collector validates visual QA by executing a three-stage workflow that captures Playwright screenshots, compares them against specification quotes, and auto-fails builds when visual evidence contradicts claims.**

The **Evidence Collector** (also referred to as *EvidenceQA*) is a specialized testing agent within the `msitarzewski/agency-agents` repository designed to enforce visual compliance through concrete screenshot evidence. Unlike subjective QA reports, this agent validates visual QA with screenshot requirements by demanding pixel-perfect proof for every UI claim, ensuring that specifications are met with deterministic artifacts rather than human interpretation.

## The Three-Stage Validation Workflow

The validation process is defined in [`testing/testing-evidence-collector.md`](https://github.com/msitarzewski/agency-agents/blob/main/testing/testing-evidence-collector.md) and operates through three tightly-coupled stages that transform shell commands into structured compliance reports.

### Stage 1: Reality-Check Commands

The agent initiates validation by executing shell commands that generate a comprehensive suite of Playwright screenshots and raw artifacts. This stage runs [`./qa-playwright-capture.sh`](https://github.com/msitarzewski/agency-agents/blob/main/./qa-playwright-capture.sh), which opens the local site, captures full-page screenshots across multiple viewports, toggles dark mode, and stores outputs under `public/qa-screenshots`.

The script also generates a [`test-results.json`](https://github.com/msitarzewski/agency-agents/blob/main/test-results.json) file that records device compatibility matrices and interaction outcomes, providing machine-readable evidence of the capture session.

### Stage 2: Visual Evidence Analysis

Once screenshots are captured, the Evidence Collector manually inspects each image, matching visual output against the exact wording of the original specification. This stage produces a markdown-styled "What I Actually See" section that lists every screenshot file—such as `responsive-desktop.png` and `dark-mode-01.png`—and flags mismatches between the spec quote and the visual reality.

The agent populates a compliance matrix using ✅/❌ indicators to show whether each specification requirement is visually satisfied by the corresponding screenshot evidence.

### Stage 3: Interactive-Element Tests

The final stage validates dynamic UI components through focused interaction checks. Using Playwright, the agent tests accordions, forms, navigation menus, mobile responsiveness, and theme toggles. Each interaction generates paired screenshots—such as `accordion-0-before.png` versus `accordion-0-after.png`—that demonstrate state changes.

These paired images are embedded in structured markdown test-result blocks that include PASS/FAIL status and detailed issue descriptions, creating auditable proof of interactive functionality.

## Detailed Validation Steps

The Evidence Collector follows a deterministic procedure defined in its agent configuration to ensure consistent validation across runs.

### Execute the Playwright Capture Script

The validation begins by running the capture script against the local development server:

```bash
./qa-playwright-capture.sh http://localhost:8000 public/qa-screenshots

```

This command generates full-page screenshots, dark-mode variants, and responsive viewport captures while writing compatibility data to [`test-results.json`](https://github.com/msitarzewski/agency-agents/blob/main/test-results.json).

### Verify Premium Feature Claims

Before visual comparison, the agent performs a sanity check to ensure that claimed "premium" features exist in the actual markup:

```bash
ls -la resources/views/ || ls -la *.html
grep -r "luxury\|premium\|glass\|morphism" . \
   --include="*.html" --include="*.css" --include="*.blade.php" \
   || echo "NO PREMIUM FEATURES FOUND"

```

This prevents false validation of features that do not exist in the source code.

### Review Test Results

The agent examines the generated JSON file to extract device compatibility and error states:

```bash
cat public/qa-screenshots/test-results.json

```

This file contains structured data about which viewports were tested and whether any interactions failed during the capture phase.

### Generate the Evidence Report

Finally, the agent compiles findings into a markdown report using a predefined template (lines 119-172 in [`testing/testing-evidence-collector.md`](https://github.com/msitarzewski/agency-agents/blob/main/testing/testing-evidence-collector.md)):

```markdown

# QA Evidence-Based Report

## 🔍 Reality Check Results

**Commands Executed**: ./qa-playwright-capture.sh …
**Screenshot Evidence**: responsive-desktop.png, dark-mode-01.png
**Specification Quote**: "Hero section must have a gradient overlay."

## 📸 Visual Evidence Analysis

- ✅ Spec says: "gradient overlay" → Screenshot shows: "gradient applied correctly"
- ❌ Spec says: "luxury glass button" → Screenshot shows: "plain rectangular button"

```

## Automatic Failure Triggers

The Evidence Collector enforces strict compliance through automatic failure conditions defined in its configuration. The build fails immediately if any of the following conditions are met:

- **No screenshots produced** – If the capture script fails to generate images in `public/qa-screenshots/`, validation cannot proceed.
- **Visual contradiction** – When screenshots directly contradict claimed UI features, such as missing "luxury" styling that was asserted in the specification.
- **Insufficient issue detection** – The agent expects to find 3-5 realistic issues on a first implementation; reports with fewer issues trigger a failure as they suggest insufficient scrutiny.

These triggers ensure that the Evidence Collector maintains rigorous standards and prevents false positives in visual QA validation.

## Key Files and Their Roles

| File | Role |
|------|------|
| [`testing/testing-evidence-collector.md`](https://github.com/msitarzewski/agency-agents/blob/main/testing/testing-evidence-collector.md) | The full agent definition, including the mandatory three-stage process, screenshot commands, report template, and automatic fail rules |
| [`testing/qa-playwright-capture.sh`](https://github.com/msitarzewski/agency-agents/blob/main/testing/qa-playwright-capture.sh) | External Playwright script that generates screenshots and [`test-results.json`](https://github.com/msitarzewski/agency-agents/blob/main/test-results.json) |
| `public/qa-screenshots/` | Output directory containing all captured images and the JSON results file |
| [`README.md`](https://github.com/msitarzewski/agency-agents/blob/main/README.md) | Repository overview with instructions for running the Evidence Collector agent |

## Summary

- The **Evidence Collector** validates visual QA through a deterministic three-stage workflow: reality-check commands, visual evidence analysis, and interactive-element testing.
- **Screenshot requirements** are enforced by the [`qa-playwright-capture.sh`](https://github.com/msitarzewski/agency-agents/blob/main/qa-playwright-capture.sh) script, which generates pixel-perfect evidence across multiple viewports and color modes.
- **Automatic failure triggers** prevent false validations by requiring actual screenshot files, visual-spec alignment, and detection of realistic implementation issues.
- All validation logic is defined in [`testing/testing-evidence-collector.md`](https://github.com/msitarzewski/agency-agents/blob/main/testing/testing-evidence-collector.md), making the process transparent and reproducible for any UI testing scenario.

## Frequently Asked Questions

### How does the Evidence Collector handle missing screenshot files?

If the [`qa-playwright-capture.sh`](https://github.com/msitarzewski/agency-agents/blob/main/qa-playwright-capture.sh) script fails to produce images in the `public/qa-screenshots/` directory, the Evidence Collector triggers an automatic failure. This ensures that validation cannot proceed without concrete visual evidence, preventing subjective or assumption-based QA reports.

### What specific screenshot variants does the validation process generate?

The capture script generates multiple screenshot variants including full-page captures, responsive viewport shots (desktop and mobile), dark-mode variants, and paired before/after images for interactive elements. These are stored as PNG files with descriptive names like `responsive-desktop.png`, `dark-mode-01.png`, and `accordion-0-before.png`.

### How does the Evidence Collector verify premium UI features?

Before comparing screenshots against specifications, the agent runs grep commands to search for keywords like "luxury," "premium," "glass," and "morphism" in HTML, CSS, and Blade template files. This sanity check ensures that claimed premium features actually exist in the source code before visual validation proceeds.

### What happens when screenshots contradict the specification?

When visual evidence contradicts the specification—such as a screenshot showing a plain button when the spec requires a "luxury glass button"—the Evidence Collector flags this with a ❌ in the compliance matrix and triggers a build failure. The agent documents the exact mismatch in the markdown report, quoting both the specification text and the visual reality observed in the screenshot.