Structuring Multi-Step Tasks with Goal-Driven Execution and Verification: A Complete Guide

Goal-Driven Execution (GDE) breaks complex software requests into minimal, verifiable sub-tasks that iterate until explicit success criteria are met, preventing over-engineering and reducing risk in LLM-assisted coding workflows.

The forrestchang/andrej-karpathy-skills repository codifies this approach through the Karpathy-Guidelines—a set of principles stored in CLAUDE.md and illustrated in EXAMPLES.md that teach large language models (LLMs) to structure multi-step tasks with clear verification checkpoints rather than open-ended "make it work" instructions.

What is Goal-Driven Execution?

Goal-Driven Execution is the first of four principles defined in the repository’s skill definition. It mandates that any development request be decomposed into a sequence of small, testable changes where each step must pass verification before proceeding to the next.

Core Philosophy from CLAUDE.md

According to CLAUDE.md lines 71‑93, GDE replaces vague imperative verbs like "fix," "add," or "make faster" with a structured plan → verify markdown table. Each row contains:

  • A minimal transformation (e.g., "Add validation" or "Write failing test then make it pass")
  • Explicit verification criteria that confirm the transformation worked
  • A hard stop if verification fails, forcing iteration on the current step rather than advancing prematurely

This approach eliminates the "vague request → big bang implementation" anti-pattern that leads to untested complexity.

The Architectural Safety Loop

The repository exposes GDE through three interconnected components:

Component Role in Structuring Multi-Step Tasks Source Reference
CLAUDE.md Declares the philosophy and provides the canonical plan → verify template Lines 71‑93
EXAMPLES.md Demonstrates concrete walkthroughs for authentication fixes, rate-limiting, and sorting bugs Lines 70‑93, 68‑94, 72‑88
SKILL.md Packages the guidelines as a Claude-Code plugin for automatic prompting Entire file

How GDE Works Internally

When processing a request, the LLM follows a six-phase execution model derived from the EXAMPLES.md implementations:

  1. Parse the user request to identify ambiguous verbs and implicit assumptions.
  2. Generate a structured plan using the markdown table format, where each step represents the smallest possible testable change.
  3. Execute the first step—often writing a failing test or stubbing an interface.
  4. Run verification via unit tests, runtime metrics, or manual checks defined in the plan.
  5. Loop conditionally: if verification fails, revise the current step; if it passes, advance to the next goal.
  6. Finalize by running the full regression suite to ensure no collateral damage from the incremental changes.

Because the guidelines reside in plain Markdown at forrestchang/andrej-karpathy-skills/CLAUDE.md, LLMs can retrieve the exact phrasing without executing code, maintaining strict compliance with security constraints.

Practical Implementation Examples

The repository provides three canonical patterns for structuring multi-step tasks with Goal-Driven Execution and verification.

Incremental Feature Development: Adding Rate Limiting

This example from EXAMPLES.md lines 70‑93 demonstrates how to build complex middleware iteratively:

Plan:
1️⃣ Add in-memory rate limiting for `/login`.
   ✅ Verify: Test that the 11th request in a minute returns 429.
2️⃣ Turn the limiter into reusable middleware.
   ✅ Verify: Same test passes for `/register` and `/login`.
3️⃣ Swap in a Redis backend for multi-process safety.
   ✅ Verify: Counter survives process restart.

Key insight: The implementation stops after the minimal viable solution (step 1) and only adds architectural complexity (step 3) after verifying the abstraction layers.

Test-First Bug Fixing: Sorting with Stable Ordering

When resolving a non-deterministic sorting bug, GDE mandates writing the failing test before touching the implementation:


# Step 1: Write a reproducible test for duplicate scores

def test_sort_duplicate_scores():
    scores = [
        {"name": "Alice", "score": 100},
        {"name": "Bob",   "score": 100},
        {"name": "Cara",  "score": 90},
    ]
    result = sort_scores(scores)
    assert result[0]["score"] == 100
    assert result[1]["score"] == 100
    assert result[2]["score"] == 90

# Step 2: Run the test → observe failure (non-deterministic order)

# Step 3: Implement stable sort

def sort_scores(scores):
    """Sort by score descending, then name ascending for ties."""
    return sorted(scores, key=lambda x: (-x["score"], x["name"]))

# Step 4: Re-run test → passes consistently

This mirrors the "test-first verification" flow documented in EXAMPLES.md lines 68‑94.

Security-Critical Changes: Authentication Session Invalidation

For high-risk modifications like invalidating active sessions after password changes, GDE enforces explicit failure verification before applying fixes:

Goal: Invalidate all active sessions after a password change.

Steps:
1️⃣ Write failing test: Change password → old session still works.
   ✅ Verify: Test fails (captures bug).
2️⃣ Add session invalidation logic.
   ✅ Verify: Test now passes.
3️⃣ Add edge-case test for concurrent sessions.
   ✅ Verify: Both old and new sessions behave correctly.

This pattern appears in EXAMPLES.md lines 72‑88 and demonstrates how verification acts as a safety guard when refactoring authentication systems.

Summary

  • Goal-Driven Execution structures multi-step tasks by decomposing vague requests into minimal, verifiable sub-goals stored in markdown tables.
  • Verification gates between steps prevent over-engineering; the LLM cannot advance until the current criterion passes, as defined in CLAUDE.md lines 71‑93.
  • Concrete examples in EXAMPLES.md provide copy-ready templates for rate limiting, sorting bugs, and authentication fixes.
  • Skill packaging via SKILL.md exposes these guidelines as a native Claude-Code plugin, enabling automatic GDE prompting in downstream tools.

Frequently Asked Questions

What makes Goal-Driven Execution different from traditional Agile task breakdown?

Traditional task breakdown often stops at the user-story level, leaving implementation details open to interpretation. Goal-Driven Execution requires explicit verification criteria for every atomic step, such as "Test that the 11th request returns 429" versus "Implement rate limiting." This eliminates ambiguity about when a step is truly complete, reducing the risk of LLM hallucinations or untested assumptions.

How does the repository handle verification when no automated tests exist?

According to EXAMPLES.md, verification can include manual checks, runtime metrics, or stubbed interfaces—not just unit tests. For example, when verifying Redis persistence in the rate-limiting example, the criterion "Counter survives process restart" can be validated manually if the testing framework lacks process isolation tools. The key is that the verification criterion is explicit and reproducible, not necessarily automated.

Can GDE be applied to non-coding tasks like documentation or data analysis?

Yes. The SKILL.md definition treats the plan → verify pattern as domain-agnostic. For documentation, a step might read "Add usage example for sort_scores()" with verification "Example renders correctly in Markdown preview." For data analysis, it could be "Filter null values" with verification "Row count drops from 10,000 to 9,847 as expected." The repository’s structure in forrestchang/andrej-karpathy-skills specifically targets code, but the underlying principle applies to any multi-step deliverable.

Where should teams store their custom GDE plans if not using the Claude-Code plugin?

Teams should maintain a CLAUDE.md or GUIDELINES.md in their repository root, following the same ATX-heading structure as the source repository. The key is placing verification criteria immediately adjacent to each step using the ✅ Verify: syntax shown in EXAMPLES.md lines 70‑93, allowing LLMs to parse the checkpoints without additional tooling.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →