Verification Loop Protocol in Oh-My-Codex: Fresh Evidence Requirements Explained
The verification loop protocol in oh-my-codex mandates that tasks cannot be marked complete until a fresh cycle of verification evidence—such as test output, build success, or lint results—confirms zero failures, as enforced by the Ralph persistence loop.
The oh-my-codex repository implements a rigorous completion standard through the Ralph skill, ensuring that automated tasks meet quality gates before finalization. This verification loop protocol requiring fresh evidence prevents premature task completion by forcing a re-verification cycle whenever initial checks fail.
How the Ralph Verification Loop Protocol Works
The protocol operates as a six-step persistence loop defined in skills/ralph/SKILL.md. Each iteration ensures that completion claims are backed by objective, freshly generated evidence rather than cached or assumed results.
Step 1: Entering the Ralph Loop
Users invoke the protocol by triggering the $ralph mode, documented in AGENTS.md at line 66. This keyword activates the Ralph skill specification, initiating a supervised session where the engine manages parallel agents through implementation, testing, and verification phases.
Steps 2-5: Execution Phases
Parallel agents execute the primary task work, including code implementation, testing, builds, and linting. These steps establish the baseline changes but do not yet authorize completion.
Step 6: Fresh Evidence Verification
The critical verification step (step 6) requires fresh evidence before accepting any completion claim. According to skills/ralph/SKILL.md at line 71, the loop executes four mandatory sub-steps:
- a. Identify the success command – The loop determines which command proves task completion (e.g.,
npm test,go test,cargo test, ormake build). - b. Run verification – The identified command executes, often with
run_in_background: truefor long-running processes. - c. Read the output – The engine parses output to confirm zero failures and successful execution.
- d. Check pending TODOs – The scan ensures no remaining "in-progress" or "TODO" items exist in the codebase.
The skill file explicitly states: "6. Verify completion with fresh evidence: a. Identify what command proves the task is complete b. Run verification (test, build, lint) c. Read the output – confirm it actually passed d. Check: zero pending/in_progress TODO items"
Architect Review and Final Checklist
After fresh evidence passes, the loop performs a tiered architect review (standard for ≥5 files or 100 lines, thorough for ≥20 files or security changes). The engine then runs a final checklist defined at line 76 of skills/ralph/SKILL.md that must satisfy all conditions:
- Fresh test run output shows all tests pass
- Fresh build output shows success
lsp_diagnosticsshows 0 errors on affected files
Only when all checklist items pass does Ralph execute /cancel to clean state and declare completion. Otherwise, the loop iterates, fixing issues and re-verifying with new evidence.
Triggering the Protocol with the $ralph Keyword
You activate the verification loop protocol requiring fresh evidence through the $ralph keyword or the omx CLI:
# Start a Ralph session for guaranteed completion
omx ralph "Add a new CLI command to list active projects"
When using the keyword directly:
$ralph "Refactor the authentication module to support OAuth2"
The engine loads skills/ralph/SKILL.md, dispatches parallel agents for steps 1-5, executes step 6 for fresh evidence, and runs the final checklist before exiting.
Evidence Types and Validation Criteria
The protocol accepts three primary categories of fresh evidence:
- Test execution output – Must explicitly show "0 failed" or equivalent success metrics
- Build artifacts – Must demonstrate clean compilation without errors
- Linting diagnostics – Must report zero errors on affected files via
lsp_diagnostics
The loop validates this evidence by parsing command output in real-time, rejecting any completion claims based on stale or cached results.
Source Implementation and Key Files
The verification loop protocol is implemented across these critical files:
skills/ralph/SKILL.md– Defines the complete Ralph workflow including the "Verify completion with fresh evidence" step and final checklist requirementsAGENTS.md(line 66) – Maps the$ralphkeyword to the verification loop triggerREADME.md(line 141) – Documents the user-facing$ralphcommand referencetemplates/AGENTS.md(line 134) – Provides keyword-to-skill mapping for template generation
These files collectively enforce that no task completes without fresh verification evidence.
Summary
- The Ralph persistence loop enforces a verification loop protocol requiring fresh evidence before task completion in oh-my-codex.
- Step 6 of the Ralph workflow mandates four specific verification actions: identifying success commands, running verification, reading output for zero failures, and checking for pending TODOs.
- The final checklist requires fresh test output, fresh build success, and zero LSP diagnostic errors before allowing loop exit.
- The protocol is triggered via the
$ralphkeyword, documented inAGENTS.mdand implemented inskills/ralph/SKILL.md. - Failed verification causes the loop to iterate rather than complete, ensuring all evidence is current and valid.
Frequently Asked Questions
What triggers the verification loop protocol in oh-my-codex?
The protocol triggers when a user invokes the $ralph keyword or runs omx ralph [task]. According to AGENTS.md, this keyword activates the Ralph skill specification, which implements the six-step persistence loop including mandatory fresh evidence verification.
What constitutes "fresh evidence" in the Ralph loop?
Fresh evidence refers to newly generated verification output from the current iteration, not cached or previous results. As defined in skills/ralph/SKILL.md, this includes current test run outputs showing zero failures, fresh build success messages, and real-time lsp_diagnostics reporting zero errors on affected files.
Where is the verification loop protocol defined?
The protocol is formally defined in skills/ralph/SKILL.md at lines 71-76, which specify the "Verify completion with fresh evidence" step and the final completion checklist. The trigger mechanism is documented in AGENTS.md at line 66, while user-facing documentation appears in README.md.
How does the protocol handle failing verification?
If the fresh evidence check fails—whether due to test failures, build errors, remaining TODOs, or LSP diagnostics—the Ralph loop does not exit. Instead, it iterates back through the workflow, dispatches agents to fix the identified issues, and re-runs the verification step with new evidence until all checklist criteria pass.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →