How the verify-work Workflow Handles Manual User Acceptance Testing in gsd-build

The verify-work workflow in gsd-build/get-shit-done orchestrates manual user acceptance testing by creating a stateful UAT.md file, guiding users through testable deliverables one checkpoint at a time, automatically inferring issue severity from free-form responses, and piping diagnosed gaps directly into an automated closure planning pipeline.

The verify-work workflow serves as the primary entry point for manual user acceptance testing (UAT) within the gsd-build/get-shit-done repository. Unlike traditional manual testing that relies on external ticketing systems, this workflow embeds the entire UAT session into a persistent markdown file that survives crashes and context switches. By structuring user interactions around deliverables extracted from phase summaries, the workflow ensures that only user-observable changes undergo verification.

Workflow Architecture Overview

The workflow operates as a state machine persisted in .planning/phases/<phase>/<phase>-UAT.md. According to the source code in get-shit-done/workflows/verify-work.md, the process begins with gsd-tools init verify-work, which either loads existing session metadata or discovers unfinished UAT files across all phases. This design ensures that a /clear command or unexpected termination never loses testing progress.

The architecture follows a gap-driven feedback loop: every issue captured during testing becomes a structured gap entry in the UAT file's YAML front matter, which the diagnose-issues workflow later enriches with root-cause analysis before plan-phase --gaps generates fix plans.

Step-by-Step UAT Execution

Session Initialization and State Recovery

When a user invokes the workflow, the initialize step (lines 23-31 in verify-work.md) checks for command-line arguments. If a phase number is provided, it loads the phase metadata including planner/checker models and directory paths. If no arguments are supplied, the check_active_session step (lines 33-74) scans .planning/phases/*-UAT.md to detect any unfinished sessions.

If multiple active sessions exist, the workflow presents them in a table and prompts the user to select one or provide a new phase number. This ensures that testing can resume exactly where it left off without data loss.

Test Discovery from Phase Summaries

Once a session is established, the workflow must identify what is actually testable. The find_summaries and extract_tests steps (lines 80-101) recursively read every *-SUMMARY.md file within the phase directory, parsing the Accomplishments and User-facing changes sections.

Only items listed as user-facing changes become test candidates. This filtering ensures that internal refactoring or infrastructure changes do not clutter the UAT checklist, keeping the manual testing focused strictly on observable behavior.

Interactive Verification Loop

The workflow creates the UAT file using the template at get-shit-done/templates/UAT.md, which structures the document with front-matter, a Current Test placeholder, a list of Tests with result: [pending] status, a Summary block, and an empty Gaps section.

During the present_test step (lines 72-92), the workflow displays the current test inside a boxed checkpoint (╔…╚…) and prompts the user for one of three responses:

  • "pass", "yes", or empty input → marks result: pass
  • "skip" → marks result: skipped with optional reason
  • Any other text → treated as an issue description

The process_response step (lines 96-130) handles the logic. For issues, it stores the text verbatim and infers severity using the severity_inference table (lines 43-52) without asking the user. The inference rules scan for keywords:

  • Blocker: "crash", "error", "exception", "fails"
  • Major: "doesn't work", "wrong", "missing", "can't"
  • Minor: "slow", "weird", "minor", "small"
  • Cosmetic: "color", "spacing", "alignment", "visual"

Each gap is appended to the Gaps section as structured YAML, ready for gsd:plan-phase --gaps.

Automated Gap Diagnosis and Closure Planning

When all tests are processed, the complete_session step (lines 93-112) commits the UAT file with a message like test(04): complete UAT – 5 passed, 1 issues. If gaps exist, the workflow immediately triggers diagnose-issues, which spawns parallel debug agents to populate each gap with root_cause, artifacts, missing, and debug_session references.

Following diagnosis, the workflow executes plan_gap_closure → verify_gap_plans → an iterative revision loop (max 3 cycles) that validates fix plans against the plan-checker before presenting the final "Fixes Ready" banner. The user is then instructed to run /gsd:execute-phase {phase} --gaps-only to apply the verified fixes.

Key Implementation Details

Severity Inference Engine

The workflow eliminates subjective severity classification by implementing a lexical analysis engine in the severity_inference table. When a user describes an issue in free-form text, the engine scans for specific keyword clusters to assign severity automatically:

def infer_severity(text):
    lowered = text.lower()
    if any(w in lowered for w in ["crash","error","exception","fails"]):
        return "blocker"
    if any(w in lowered for w in ["doesn't work","wrong","missing","can't"]):
        return "major"
    if any(w in lowered for w in ["slow","weird","minor","small"]):
        return "minor"
    if any(w in lowered for w in ["color","spacing","alignment","visual"]):
        return "cosmetic"
    return "major"   # safe default

This approach ensures consistent severity assignment across different users and testing sessions without requiring additional input steps.

Stateful UAT.md Structure

The UAT file acts as a database and conversation log simultaneously. Located at .planning/phases/<phase>/<phase>-UAT.md, it contains YAML front-matter tracking status, phase, started, and updated timestamps, followed by markdown sections for Current Test, Tests (with checkboxes and result states), Summary (pass/fail counts), and Gaps (structured YAML entries).

This file-based state machine allows the workflow to resume after interruptions by simply re-reading the UAT file and locating the first test with result: [pending].

Practical Usage Examples

Initializing a Verification Session

To start UAT for a specific phase, use the gsd-tools CLI wrapper:


# Start verification for phase 04

INIT=$(node ~/.claude/get-shit-done/bin/gsd-tools.cjs init verify-work "04")

# Extract metadata for subsequent steps

phase_dir=$(echo "$INIT" | jq -r .phase_dir)
phase_number=$(echo "$INIT" | jq -r .phase_number)

This initialization loads the phase metadata and determines whether to resume an existing session or create a new one.

Detecting Active Sessions

When no phase argument is provided, scan for unfinished UAT files:


# List all active UAT sessions

find .planning/phases -name "*-UAT.md" -type f | head -5

The workflow uses this detection logic in the check_active_session step to present users with resumable testing sessions.

Committing Completed UAT

Upon completion, the workflow commits the UAT file with a standardized message:

node ~/.claude/get-shit-done/bin/gsd-tools.cjs commit \
  "test(${phase_number}): complete UAT - ${passed} passed, ${issues} issues" \
  --files "$UAT_PATH"

This commit triggers the transition to gap diagnosis if issues were recorded.

Summary

  • The verify-work workflow in gsd-build/get-shit-done transforms manual UAT into a structured, file-based state machine that persists across interruptions.
  • Test discovery automatically extracts verifiable items from *-SUMMARY.md files, filtering for user-facing changes only.
  • Severity inference eliminates manual classification by analyzing free-form issue descriptions for keywords like "crash" (blocker) or "color" (cosmetic).
  • Gap-driven automation pipes recorded issues through diagnose-issues, plan-phase --gaps, and plan-checker to generate verified fix plans without human intervention.
  • The workflow ultimately hands off to /gsd:execute-phase {phase} --gaps-only for automated remediation of accepted issues.

Frequently Asked Questions

How does the verify-work workflow resume a crashed or interrupted testing session?

The workflow maintains all session state in a markdown file located at .planning/phases/<phase>/<phase>-UAT.md. When restarted without arguments, the check_active_session step scans for existing *-UAT.md files and lists them for the user to select. If a specific phase is provided, the workflow reads the existing UAT file, locates the first test with result: [pending], and resumes the present_test loop from that checkpoint.

What criteria does the workflow use to determine which deliverables require manual testing?

The workflow only tests items explicitly marked as user-facing changes. During the extract_tests step, it parses every *-SUMMARY.md file in the phase directory and extracts entries from the Accomplishments and User-facing changes sections. Internal refactoring, infrastructure updates, or technical debt items that do not appear in these sections are automatically excluded from the UAT checklist, ensuring testers focus strictly on observable functionality.

How does the workflow handle issue severity without asking users to classify problems?

The process_response step treats any free-form text that is not "pass" or "skip" as an issue description. It then applies the severity_inference lexical analysis rules to assign severity automatically. The engine scans for specific keyword clusters: "crash", "error", "exception", or "fails" trigger blocker; "doesn't work", "wrong", "missing", or "can't" trigger major; "slow", "weird", "minor", or "small" trigger minor; and "color", "spacing", "alignment", or "visual" trigger cosmetic. This eliminates subjective bias and accelerates the testing workflow.

What happens to recorded gaps after the UAT session is marked complete?

Upon completion, the complete_session step commits the UAT file and checks for any entries in the Gaps section. If gaps exist, it immediately triggers the diagnose-issues workflow, which spawns parallel debug agents to populate each gap with root_cause, artifacts, missing, and debug_session fields. The workflow then executes plan_gap_closure to generate fix plans, validates them through the plan-checker in an iterative loop (maximum three cycles), and finally presents a "Fixes Ready" banner. The user is then instructed to execute /gsd:execute-phase {phase} --gaps-only to apply the verified fixes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →