# Structuring Multi-Step Tasks with Goal-Driven Execution and Verification: A Complete Guide

> Master goal-driven execution and verification to structure multi-step tasks. Break down complex requests into verifiable sub-tasks for LLM-assisted coding workflows. Reduce risk and prevent over-engineering.

- Repository: [Jiayuan Zhang/andrej-karpathy-skills](https://github.com/forrestchang/andrej-karpathy-skills)
- Tags: deep-dive
- Published: 2026-04-08

---

**Goal-Driven Execution (GDE) breaks complex software requests into minimal, verifiable sub-tasks that iterate until explicit success criteria are met, preventing over-engineering and reducing risk in LLM-assisted coding workflows.**

The `forrestchang/andrej-karpathy-skills` repository codifies this approach through the **Karpathy-Guidelines**—a set of principles stored in [`CLAUDE.md`](https://github.com/forrestchang/andrej-karpathy-skills/blob/main/CLAUDE.md) and illustrated in [`EXAMPLES.md`](https://github.com/forrestchang/andrej-karpathy-skills/blob/main/EXAMPLES.md) that teach large language models (LLMs) to structure multi-step tasks with clear verification checkpoints rather than open-ended "make it work" instructions.

## What is Goal-Driven Execution?

Goal-Driven Execution is the first of four principles defined in the repository’s skill definition. It mandates that any development request be decomposed into a sequence of small, testable changes where each step must pass verification before proceeding to the next.

### Core Philosophy from CLAUDE.md

According to [`CLAUDE.md`](https://github.com/forrestchang/andrej-karpathy-skills/blob/main/CLAUDE.md) lines 71‑93, GDE replaces vague imperative verbs like "fix," "add," or "make faster" with a structured **plan → verify** markdown table. Each row contains:

- **A minimal transformation** (e.g., "Add validation" or "Write failing test then make it pass")
- **Explicit verification criteria** that confirm the transformation worked
- **A hard stop** if verification fails, forcing iteration on the current step rather than advancing prematurely

This approach eliminates the "vague request → big bang implementation" anti-pattern that leads to untested complexity.

### The Architectural Safety Loop

The repository exposes GDE through three interconnected components:

| Component | Role in Structuring Multi-Step Tasks | Source Reference |
|-----------|--------------------------------------|------------------|
| [`CLAUDE.md`](https://github.com/forrestchang/andrej-karpathy-skills/blob/main/CLAUDE.md) | Declares the philosophy and provides the canonical *plan → verify* template | Lines 71‑93 |
| [`EXAMPLES.md`](https://github.com/forrestchang/andrej-karpathy-skills/blob/main/EXAMPLES.md) | Demonstrates concrete walkthroughs for authentication fixes, rate-limiting, and sorting bugs | Lines 70‑93, 68‑94, 72‑88 |
| [`SKILL.md`](https://github.com/forrestchang/andrej-karpathy-skills/blob/main/SKILL.md) | Packages the guidelines as a Claude-Code plugin for automatic prompting | Entire file |

## How GDE Works Internally

When processing a request, the LLM follows a six-phase execution model derived from the [`EXAMPLES.md`](https://github.com/forrestchang/andrej-karpathy-skills/blob/main/EXAMPLES.md) implementations:

1. **Parse the user request** to identify ambiguous verbs and implicit assumptions.
2. **Generate a structured plan** using the markdown table format, where each step represents the smallest possible testable change.
3. **Execute the first step**—often writing a failing test or stubbing an interface.
4. **Run verification** via unit tests, runtime metrics, or manual checks defined in the plan.
5. **Loop conditionally**: if verification fails, revise the current step; if it passes, advance to the next goal.
6. **Finalize** by running the full regression suite to ensure no collateral damage from the incremental changes.

Because the guidelines reside in plain Markdown at [`forrestchang/andrej-karpathy-skills/CLAUDE.md`](https://github.com/forrestchang/andrej-karpathy-skills/blob/main/forrestchang/andrej-karpathy-skills/CLAUDE.md), LLMs can retrieve the exact phrasing without executing code, maintaining strict compliance with security constraints.

## Practical Implementation Examples

The repository provides three canonical patterns for structuring multi-step tasks with Goal-Driven Execution and verification.

### Incremental Feature Development: Adding Rate Limiting

This example from [`EXAMPLES.md`](https://github.com/forrestchang/andrej-karpathy-skills/blob/main/EXAMPLES.md) lines 70‑93 demonstrates how to build complex middleware iteratively:

```markdown
Plan:
1️⃣ Add in-memory rate limiting for `/login`.
   ✅ Verify: Test that the 11th request in a minute returns 429.
2️⃣ Turn the limiter into reusable middleware.
   ✅ Verify: Same test passes for `/register` and `/login`.
3️⃣ Swap in a Redis backend for multi-process safety.
   ✅ Verify: Counter survives process restart.

```

**Key insight:** The implementation stops after the minimal viable solution (step 1) and only adds architectural complexity (step 3) after verifying the abstraction layers.

### Test-First Bug Fixing: Sorting with Stable Ordering

When resolving a non-deterministic sorting bug, GDE mandates writing the failing test before touching the implementation:

```python

# Step 1: Write a reproducible test for duplicate scores

def test_sort_duplicate_scores():
    scores = [
        {"name": "Alice", "score": 100},
        {"name": "Bob",   "score": 100},
        {"name": "Cara",  "score": 90},
    ]
    result = sort_scores(scores)
    assert result[0]["score"] == 100
    assert result[1]["score"] == 100
    assert result[2]["score"] == 90

# Step 2: Run the test → observe failure (non-deterministic order)

# Step 3: Implement stable sort

def sort_scores(scores):
    """Sort by score descending, then name ascending for ties."""
    return sorted(scores, key=lambda x: (-x["score"], x["name"]))

# Step 4: Re-run test → passes consistently

```

This mirrors the "test-first verification" flow documented in [`EXAMPLES.md`](https://github.com/forrestchang/andrej-karpathy-skills/blob/main/EXAMPLES.md) lines 68‑94.

### Security-Critical Changes: Authentication Session Invalidation

For high-risk modifications like invalidating active sessions after password changes, GDE enforces explicit failure verification before applying fixes:

```markdown
Goal: Invalidate all active sessions after a password change.

Steps:
1️⃣ Write failing test: Change password → old session still works.
   ✅ Verify: Test fails (captures bug).
2️⃣ Add session invalidation logic.
   ✅ Verify: Test now passes.
3️⃣ Add edge-case test for concurrent sessions.
   ✅ Verify: Both old and new sessions behave correctly.

```

This pattern appears in [`EXAMPLES.md`](https://github.com/forrestchang/andrej-karpathy-skills/blob/main/EXAMPLES.md) lines 72‑88 and demonstrates how verification acts as a safety guard when refactoring authentication systems.

## Summary

- **Goal-Driven Execution** structures multi-step tasks by decomposing vague requests into minimal, verifiable sub-goals stored in markdown tables.
- **Verification gates** between steps prevent over-engineering; the LLM cannot advance until the current criterion passes, as defined in [`CLAUDE.md`](https://github.com/forrestchang/andrej-karpathy-skills/blob/main/CLAUDE.md) lines 71‑93.
- **Concrete examples** in [`EXAMPLES.md`](https://github.com/forrestchang/andrej-karpathy-skills/blob/main/EXAMPLES.md) provide copy-ready templates for rate limiting, sorting bugs, and authentication fixes.
- **Skill packaging** via [`SKILL.md`](https://github.com/forrestchang/andrej-karpathy-skills/blob/main/SKILL.md) exposes these guidelines as a native Claude-Code plugin, enabling automatic GDE prompting in downstream tools.

## Frequently Asked Questions

### What makes Goal-Driven Execution different from traditional Agile task breakdown?

Traditional task breakdown often stops at the user-story level, leaving implementation details open to interpretation. **Goal-Driven Execution requires explicit verification criteria for every atomic step**, such as "Test that the 11th request returns 429" versus "Implement rate limiting." This eliminates ambiguity about when a step is truly complete, reducing the risk of LLM hallucinations or untested assumptions.

### How does the repository handle verification when no automated tests exist?

According to [`EXAMPLES.md`](https://github.com/forrestchang/andrej-karpathy-skills/blob/main/EXAMPLES.md), verification can include manual checks, runtime metrics, or stubbed interfaces—not just unit tests. For example, when verifying Redis persistence in the rate-limiting example, the criterion "Counter survives process restart" can be validated manually if the testing framework lacks process isolation tools. The key is that the verification criterion is **explicit and reproducible**, not necessarily automated.

### Can GDE be applied to non-coding tasks like documentation or data analysis?

Yes. The [`SKILL.md`](https://github.com/forrestchang/andrej-karpathy-skills/blob/main/SKILL.md) definition treats the *plan → verify* pattern as domain-agnostic. For documentation, a step might read "Add usage example for `sort_scores()`" with verification "Example renders correctly in Markdown preview." For data analysis, it could be "Filter null values" with verification "Row count drops from 10,000 to 9,847 as expected." The repository’s structure in `forrestchang/andrej-karpathy-skills` specifically targets code, but the underlying principle applies to any multi-step deliverable.

### Where should teams store their custom GDE plans if not using the Claude-Code plugin?

Teams should maintain a [`CLAUDE.md`](https://github.com/forrestchang/andrej-karpathy-skills/blob/main/CLAUDE.md) or [`GUIDELINES.md`](https://github.com/forrestchang/andrej-karpathy-skills/blob/main/GUIDELINES.md) in their repository root, following the same ATX-heading structure as the source repository. The key is placing verification criteria immediately adjacent to each step using the `✅ Verify:` syntax shown in [`EXAMPLES.md`](https://github.com/forrestchang/andrej-karpathy-skills/blob/main/EXAMPLES.md) lines 70‑93, allowing LLMs to parse the checkpoints without additional tooling.