Verification Loop Protocol in Oh-My-Codex: Fresh Evidence Requirements Explained

The verification loop protocol in oh-my-codex mandates that tasks cannot be marked complete until a fresh cycle of verification evidence—such as test output, build success, or lint results—confirms zero failures, as enforced by the Ralph persistence loop.

The oh-my-codex repository implements a rigorous completion standard through the Ralph skill, ensuring that automated tasks meet quality gates before finalization. This verification loop protocol requiring fresh evidence prevents premature task completion by forcing a re-verification cycle whenever initial checks fail.

How the Ralph Verification Loop Protocol Works

The protocol operates as a six-step persistence loop defined in skills/ralph/SKILL.md. Each iteration ensures that completion claims are backed by objective, freshly generated evidence rather than cached or assumed results.

Step 1: Entering the Ralph Loop

Users invoke the protocol by triggering the $ralph mode, documented in AGENTS.md at line 66. This keyword activates the Ralph skill specification, initiating a supervised session where the engine manages parallel agents through implementation, testing, and verification phases.

Steps 2-5: Execution Phases

Parallel agents execute the primary task work, including code implementation, testing, builds, and linting. These steps establish the baseline changes but do not yet authorize completion.

Step 6: Fresh Evidence Verification

The critical verification step (step 6) requires fresh evidence before accepting any completion claim. According to skills/ralph/SKILL.md at line 71, the loop executes four mandatory sub-steps:

  • a. Identify the success command – The loop determines which command proves task completion (e.g., npm test, go test, cargo test, or make build).
  • b. Run verification – The identified command executes, often with run_in_background: true for long-running processes.
  • c. Read the output – The engine parses output to confirm zero failures and successful execution.
  • d. Check pending TODOs – The scan ensures no remaining "in-progress" or "TODO" items exist in the codebase.

The skill file explicitly states: "6. Verify completion with fresh evidence: a. Identify what command proves the task is complete b. Run verification (test, build, lint) c. Read the output – confirm it actually passed d. Check: zero pending/in_progress TODO items"

Architect Review and Final Checklist

After fresh evidence passes, the loop performs a tiered architect review (standard for ≥5 files or 100 lines, thorough for ≥20 files or security changes). The engine then runs a final checklist defined at line 76 of skills/ralph/SKILL.md that must satisfy all conditions:

  • Fresh test run output shows all tests pass
  • Fresh build output shows success
  • lsp_diagnostics shows 0 errors on affected files

Only when all checklist items pass does Ralph execute /cancel to clean state and declare completion. Otherwise, the loop iterates, fixing issues and re-verifying with new evidence.

Triggering the Protocol with the $ralph Keyword

You activate the verification loop protocol requiring fresh evidence through the $ralph keyword or the omx CLI:


# Start a Ralph session for guaranteed completion

omx ralph "Add a new CLI command to list active projects"

When using the keyword directly:

$ralph "Refactor the authentication module to support OAuth2"

The engine loads skills/ralph/SKILL.md, dispatches parallel agents for steps 1-5, executes step 6 for fresh evidence, and runs the final checklist before exiting.

Evidence Types and Validation Criteria

The protocol accepts three primary categories of fresh evidence:

  1. Test execution output – Must explicitly show "0 failed" or equivalent success metrics
  2. Build artifacts – Must demonstrate clean compilation without errors
  3. Linting diagnostics – Must report zero errors on affected files via lsp_diagnostics

The loop validates this evidence by parsing command output in real-time, rejecting any completion claims based on stale or cached results.

Source Implementation and Key Files

The verification loop protocol is implemented across these critical files:

  • skills/ralph/SKILL.md – Defines the complete Ralph workflow including the "Verify completion with fresh evidence" step and final checklist requirements
  • AGENTS.md (line 66) – Maps the $ralph keyword to the verification loop trigger
  • README.md (line 141) – Documents the user-facing $ralph command reference
  • templates/AGENTS.md (line 134) – Provides keyword-to-skill mapping for template generation

These files collectively enforce that no task completes without fresh verification evidence.

Summary

  • The Ralph persistence loop enforces a verification loop protocol requiring fresh evidence before task completion in oh-my-codex.
  • Step 6 of the Ralph workflow mandates four specific verification actions: identifying success commands, running verification, reading output for zero failures, and checking for pending TODOs.
  • The final checklist requires fresh test output, fresh build success, and zero LSP diagnostic errors before allowing loop exit.
  • The protocol is triggered via the $ralph keyword, documented in AGENTS.md and implemented in skills/ralph/SKILL.md.
  • Failed verification causes the loop to iterate rather than complete, ensuring all evidence is current and valid.

Frequently Asked Questions

What triggers the verification loop protocol in oh-my-codex?

The protocol triggers when a user invokes the $ralph keyword or runs omx ralph [task]. According to AGENTS.md, this keyword activates the Ralph skill specification, which implements the six-step persistence loop including mandatory fresh evidence verification.

What constitutes "fresh evidence" in the Ralph loop?

Fresh evidence refers to newly generated verification output from the current iteration, not cached or previous results. As defined in skills/ralph/SKILL.md, this includes current test run outputs showing zero failures, fresh build success messages, and real-time lsp_diagnostics reporting zero errors on affected files.

Where is the verification loop protocol defined?

The protocol is formally defined in skills/ralph/SKILL.md at lines 71-76, which specify the "Verify completion with fresh evidence" step and the final completion checklist. The trigger mechanism is documented in AGENTS.md at line 66, while user-facing documentation appears in README.md.

How does the protocol handle failing verification?

If the fresh evidence check fails—whether due to test failures, build errors, remaining TODOs, or LSP diagnostics—the Ralph loop does not exit. Instead, it iterates back through the workflow, dispatches agents to fix the identified issues, and re-runs the verification step with new evidence until all checklist criteria pass.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →