Creating Effective Verification Checks for Multi-Step Plans: A Goal-Driven Execution Guide
Apply Goal-Driven Execution by defining concrete success criteria upfront, writing failing verification tests before implementation, and executing regression guards after each surgical change to ensure reliable multi-step plan validation.
The forrestchang/andrej-karpathy-skills repository provides a battle-tested framework for LLM-assisted development that prioritizes systematic verification over assumption. This knowledge base teaches Claude and developers to transform ambiguous requirements into automated, testable validation loops that catch failures before they propagate through complex workflows.
The Six-Step Verification Framework
As documented in CLAUDE.md at line 71 and README.md at line 128, Goal-Driven Execution forms the architectural backbone of reliable multi-step verification. The framework mandates that every plan consist of concrete, testable goals rather than vague instructions.
Step 1: Define Success Criteria
Convert ambiguous requirements into specific, measurable outcomes. According to EXAMPLES.md at line 74, verifiable goals enable automatic validation where vague descriptions fail. Success criteria must be concrete enough to evaluate programmatically.
Step 2: Write Verification Tests First
Create a failing test that reproduces the expected outcome or problem before writing implementation code. This safety net ensures that the LLM or developer must make the test pass before proceeding, as outlined in EXAMPLES.md at line 74 under the "Vague vs. Verifiable" section.
Step 3: Implement Surgical Changes
Add only the minimal code necessary to satisfy the verification test. The repository emphasizes Surgical Changes in EXAMPLES.md at line 70 to reduce accidental side effects and maintain clean, reviewable diffs.
Step 4: Execute Immediate Verification
Run the specific test suite to confirm the change works. The Goal-Driven Execution Loop in CLAUDE.md at line 71 mandates immediate feedback—if the test fails, iterate without additional modifications.
Step 5: Run Regression Guards
Execute the full existing test suite to ensure no other functionality broke. This step prevents regressions that the new test does not cover, as specified in the same execution loop in CLAUDE.md.
Step 6: Document Verification Criteria
Write a brief verification checklist in the PR description or code comments. The skills/karpathy-guidelines/SKILL.md file provides a verification checklist template for documenting intent explicitly.
Practical Code Implementation Patterns
The repository demonstrates concrete patterns for embedding verification into multi-step workflows using isolated test execution.
Simple Verification Checklist Pattern
Create a self-contained verification script that can execute from CI jobs or PR comments. This pattern validates response status, payload shape, and existing endpoint integrity:
# verification_checklist.py
def verify_new_endpoint(client):
"""
Verification steps for adding /api/v1/health endpoint.
"""
# 1️⃣ Verify response status
resp = client.get("/api/v1/health")
assert resp.status_code == 200, "Health endpoint should return 200"
# 2️⃣ Verify payload shape
data = resp.json()
assert "status" in data, "Payload must contain 'status' key"
assert data["status"] == "ok", "Health status must be 'ok'"
# 3️⃣ Regression guard – ensure existing endpoints still work
existing = client.get("/api/v1/users")
assert existing.status_code == 200, "Existing /users endpoint broke"
Multi-Step Plan with Explicit Verification Blocks
Structure each phase with dedicated verification blocks that run isolated test files. This architecture stops execution at the first failing verification:
# multi_step_plan.py
def step_1_add_rate_limit(app):
"""Add in‑memory rate limiting to a single route."""
# ✅ Verification for step 1
assert app.has_route("/login")
app.add_middleware(in_memory_rate_limiter, limit=10) # 10 req/min
# Run step-1 test suite
run_tests("tests/rate_limit_step1.py")
def step_2_extract_middleware(app):
"""Make rate limiting reusable across all routes."""
# ✅ Verification for step 2
assert app.middleware_contains(in_memory_rate_limiter)
app.apply_middleware_to_all()
# Run step-2 test suite
run_tests("tests/rate_limit_step2.py")
def step_3_persist_backend(app):
"""Swap in a Redis backend for distributed rate limits."""
# ✅ Verification for step 3
assert app.rate_limiter.backend == "redis"
# Run step-3 test suite
run_tests("tests/rate_limit_step3.py")
Summary
- Goal-Driven Execution requires concrete, testable success criteria before any code changes, as defined in
README.mdandCLAUDE.md. - Write failing tests first to establish a safety net that must pass before proceeding to the next step.
- Surgical changes minimize side effects by implementing only the minimal code necessary to satisfy verification.
- Regression guards prevent collateral damage by running the full test suite after individual step verification.
- Documentation in
skills/karpathy-guidelines/SKILL.mdprovides templates for verification checklists that make intent explicit.
Frequently Asked Questions
What is Goal-Driven Execution in the context of verification?
Goal-Driven Execution is a core principle in CLAUDE.md that mandates breaking any request into a series of verify-check steps. It requires defining concrete, testable goals rather than vague instructions, ensuring every change can be automatically validated through the execution loop documented at line 71.
How does writing failing tests first improve multi-step plan reliability?
Writing verification tests before implementation creates an immediate feedback mechanism. As shown in EXAMPLES.md at line 74, this approach forces the developer or LLM to satisfy specific success criteria before proceeding, preventing the accumulation of untested changes that could compound errors across multiple steps.
Where should verification criteria be documented according to the repository?
Verification criteria should be documented in the PR description using the verification checklist template provided in skills/karpathy-guidelines/SKILL.md. Additionally, inline code comments marked with ✅ verification blocks (as demonstrated in the multi-step plan examples) make validation steps explicit within the source code.
Why are regression guards necessary if individual step tests pass?
Regression guards detect side effects that the new verification test does not cover. According to the Goal-Driven Execution Loop in CLAUDE.md, running the full existing test suite ensures that surgical changes to one component have not inadvertently broken unrelated functionality elsewhere in the codebase.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →