How the Testing-Reality-Checker Determines Production Readiness in Agency-Agents

The testing-reality-checker defaults to a "NEEDS WORK" state and only flips to "READY" after passing a strict three-phase validation that requires overwhelming, evidence-based proof across reality-check commands, QA cross-validation, and end-to-end system testing.

The testing-reality-checker in the msitarzewski/agency-agents repository acts as the final hard gate before deployment. Unlike typical checklists that assume success until failure is found, this agent starts from a position of skepticism and demands concrete proof that the system meets production standards. Its workflow is defined in testing/testing-reality-checker.md and orchestrated through specialized/agents-orchestrator.md.

The Three-Phase Validation Workflow

The testing-reality-checker executes a sequential validation pipeline that moves from raw evidence gathering to holistic system verification.

Phase 1: Reality-Check Commands and Evidence Gathering

The agent begins by validating what actually exists in the codebase rather than trusting documentation. It executes specific shell commands to gather raw artefacts:

  • List built files: ls -la resources/views/ || ls -la *.html confirms the concrete UI stack (Laravel Blade, plain HTML, etc.)
  • Search for premium UI claims: grep -r "luxury\|premium\|glass\|morphism" detects marketing assertions that must be substantiated
  • Capture screenshots: ./qa-playwright-capture.sh http://localhost:8000 public/qa-screenshots generates device-specific PNGs and a test-results.json payload containing load times, errors, and interaction status
  • Review artefacts: ls -la public/qa-screenshots/ && cat public/qa-screenshots/test-results.json provides the agent with concrete UI and performance metrics

These commands ensure the agent works from observable reality rather than developer claims, as defined in lines 39-55 of testing/testing-reality-checker.md.

Phase 2: QA Cross-Validation

The agent cross-checks the gathered evidence against the earlier QA agent's report. It reads the QA findings and compares the screenshots and test-results.json data with the QA-reported issues. Any mismatch—such as a bug still visible in the screenshots—triggers an automatic "FAIL" flag for that item.

This step ensures that previous automated QA cannot hide defects once the system is fully assembled, as implemented in lines 56-61.

Phase 3: End-to-End System Validation

Finally, the checker evaluates complete user journeys across devices. It examines desktop, tablet, and mobile screenshots for layout consistency. It compares before-and-after screenshots of interaction sequences (navigation clicks, form fills, accordion toggles). It validates performance numbers from test-results.json, requiring load times under 3 seconds and error-free operation.

If every journey passes, the agent records a PASS; otherwise it logs a FAIL with explicit evidence links, as detailed in lines 62-67.

Decision Logic and Production Readiness Criteria

After completing the three phases, the testing-reality-checker builds a Production-Readiness Report using the markdown template beginning at line 40. The report contains an Overall Quality Rating (C+, B-, etc.) derived from passed versus failed checks.

The Production Readiness field defaults to "NEEDS WORK" and only changes to "READY" when all of the following conditions are met:

  1. No automatic fail triggers are present (no fantasy claims, missing screenshots, or performance exceeding 3 seconds), as defined in lines 22-27
  2. Every spec-vs-implementation comparison shows PASS with concrete screenshot evidence
  3. All QA-identified issues are resolved, verified by the cross-validation step
  4. The overall rating meets or exceeds the predefined threshold (typically B- or better)

If any condition fails, the final status remains "NEEDS WORK" or "FAILED" for critical blockers. The report's "Deployment Readiness Assessment" line makes this explicit, as shown in lines 77-80.

Integration in the CI/CD Pipeline

The specialized/agents-orchestrator.md automatically spawns the testing-reality-checker as the final step of the autonomous development workflow at line 103. Only when the reality-checker returns "READY" does the orchestrator proceed to release or deployment. This integration makes the check a hard quality gate for the entire system, ensuring that no code reaches production without passing the evidence-based validation.

Code Examples

Below are concise snippets showing how to trigger the check and interpret its results.


# Spawn the reality-checker via the orchestrator (recommended)

./run-orchestrator.sh project-specs/my-app-setup.md

# Directly invoke the checker for ad-hoc validation

cd testing
./testing-reality-checker.sh   # wraps the three steps described above

# Example: programmatically parse the final report for the readiness flag

import toml, pathlib, re

report_path = pathlib.Path("public/qa-screenshots/report.md")
text = report_path.read_text()

# Extract the "Production Readiness" line

ready = re.search(r"Production Readiness:\s+(\w+)", text)
status = ready.group(1) if ready else "UNKNOWN"
print(f"⚙️  Production readiness: {status}")
<!-- Sample excerpt from the generated report -->

## 📈 Success Metrics for Next Iteration

**What Needs Improvement**: …

## 🔄 Deployment Readiness Assessment

**Status**: **NEEDS WORK** (default unless overwhelming evidence supports ready)

Summary

  • The testing-reality-checker defaults to "NEEDS WORK" and requires overwhelming evidence to declare "READY".
  • It executes a three-phase validation: reality-check commands, QA cross-validation, and end-to-end system testing.
  • Automatic fail triggers (fantasy claims, missing screenshots, slow performance) immediately block production readiness.
  • The checker serves as a hard quality gate in the CI/CD pipeline, with the orchestrator only proceeding to deployment upon receiving a "READY" status.
  • All evidence is captured in public/qa-screenshots/test-results.json and cross-referenced against the QA agent's findings.

Frequently Asked Questions

What triggers an automatic "FAIL" in the testing-reality-checker?

The testing-reality-checker flags an automatic FAIL when it detects fantasy claims (unsubstantiated marketing language like "luxury" or "premium" without evidence), missing screenshots in the public/qa-screenshots/ directory, or performance metrics exceeding the 3-second load-time threshold defined in test-results.json. These triggers are defined in lines 22-27 of testing/testing-reality-checker.md.

How does the reality-checker verify that QA issues are actually resolved?

During Phase 2 (QA Cross-Validation), the agent reads the previous QA agent's report and compares it against the current screenshots and test-results.json data. If the QA report identified a bug that still appears in the fresh screenshots, or if the performance metrics still show errors, the checker flags a FAIL for that specific item. This ensures that earlier QA approvals cannot mask persistent defects.

Can the testing-reality-checker be run independently of the orchestrator?

Yes, while the specialized/agents-orchestrator.md automatically spawns the reality-checker as the final pipeline step at line 103, you can invoke it directly for ad-hoc validation. Navigate to the testing directory and execute ./testing-reality-checker.sh, which wraps the three validation phases (reality-check commands, QA cross-validation, and end-to-end system testing) described in the workflow documentation.

What quality rating must a project achieve to be considered "READY"?

The project must achieve an Overall Quality Rating of B- or better, though this threshold can be configured. More importantly, the status only flips from the default "NEEDS WORK" to "READY" when all automatic fail triggers are absent, every spec-vs-implementation comparison shows PASS with screenshot evidence, and all QA-identified issues are resolved. The rating is derived from the count of passed versus failed checks in the final report.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →