# What Happens When Test‑First Verification Finds Issues in Andrej Karpathy Skills

> Discover what happens when test-first verification finds issues. Learn how developers fix bugs, iterate on code, and prevent regressions in the Andrej Karpathy skills repo.

- Repository: [multica-ai/andrej-karpathy-skills](https://github.com/multica-ai/andrej-karpathy-skills)
- Tags: deep-dive
- Published: 2026-04-19

---

**When test‑first verification finds issues in the multica‑ai/andrej‑karpathy‑skills repository, developers capture the bug in a failing test, iterate on the implementation until the test passes, and verify the entire suite remains green to prevent regressions.**

The multica‑ai/andrej‑karpathy‑skills repository enforces a rigorous **Goal‑Driven Execution** principle that mandates a test‑first workflow whenever issues are discovered. According to the source code in [`README.md`](https://github.com/multica-ai/andrej-karpathy-skills/blob/main/README.md) (lines 77‑88), this approach treats every bug as a requirement for a failing test that must pass before the implementation is considered complete.

## The Goal‑Driven Execution Verification Loop

When test‑first verification finds issues, the repository dictates a four‑step verification loop that turns failures into concrete success criteria. This loop is documented in [`EXAMPLES.md`](https://github.com/multica-ai/andrej-karpathy-skills/blob/main/EXAMPLES.md) (lines 54‑95) and ensures that every fix is provably correct.

### Step 1: Write a Failing Test That Reproduces the Issue

The developer first writes a test that explicitly reproduces the discovered issue. This test serves as the *success criterion* for the bug—if the test passes, the issue is resolved. In [`EXAMPLES.md`](https://github.com/multica-ai/andrej-karpathy-skills/blob/main/EXAMPLES.md) (lines 70‑84), the `test_sort_with_duplicate_scores` function demonstrates this by asserting deterministic ordering for items with identical scores.

### Step 2: Run the Test to Confirm the Failure

The test is executed to confirm it fails, validating that the issue is real and that the test correctly captures the bug. This step prevents false positives and ensures the baseline is honest.

### Step 3: Implement the Minimal Fix

With a reproducible failure in place, the developer implements the smallest possible change that makes the test pass. In the repository’s example, this meant replacing an unstable sort with a **stable sort** that uses a composite key to break ties deterministically.

### Step 4: Re‑run the Full Test Suite

After the fix, the developer runs the original test again—it must now pass—and then executes the entire test suite to ensure no regressions were introduced. If any test fails, the loop repeats until the criteria are satisfied.

## Real‑World Example: Fixing a Nondeterministic Sort in EXAMPLES.md

The repository provides a concrete scenario in [`EXAMPLES.md`](https://github.com/multica-ai/andrej-karpathy-skills/blob/main/EXAMPLES.md) (lines 54‑95) where test‑first verification finds issues with duplicate score handling. The initial implementation produced nondeterministic ordering, which the test surfaced immediately.

### The Failing Test

```python
def test_sort_with_duplicate_scores():
    """
    Test sorting when multiple items have the same score.
    """
    scores = [
        {'name': 'Alice',   'score': 100},
        {'name': 'Bob',     'score': 100},
        {'name': 'Charlie', 'score': 90},
    ]

    result = sort_scores(scores)

    # The bug: order is nondeterministic for duplicates

    # Run this test many times – it should be consistent

    assert result[0]['score'] == 100
    assert result[1]['score'] == 100
    assert result[2]['score'] == 90

```

*The test fails because the original `sort_scores` does not guarantee a deterministic order for equal scores* (see lines 70‑84 of [`EXAMPLES.md`](https://github.com/multica-ai/andrej-karpathy-skills/blob/main/EXAMPLES.md)).

### The Fix After Verification

Once the failure confirmed the bug, the developer implemented a stable sort using a composite key:

```python
def sort_scores(scores):
    """Sort by score descending, then name ascending for ties."""
    return sorted(scores, key=lambda x: (-x['score'], x['name']))

```

*Running the full test suite now shows the new implementation passes the reproducibility check* (see lines 89‑92 of [`EXAMPLES.md`](https://github.com/multica-ai/andrej-karpathy-skills/blob/main/EXAMPLES.md)).

## Key Source Files and Their Roles

| File | Role | Key Lines |
|------|------|-----------|
| [`README.md`](https://github.com/multica-ai/andrej-karpathy-skills/blob/main/README.md) | Defines the **Goal‑Driven Execution** principle that mandates the test‑first workflow | 77‑88 |
| [`EXAMPLES.md`](https://github.com/multica-ai/andrej-karpathy-skills/blob/main/EXAMPLES.md) | Provides concrete demonstrations of the verification loop, including the sorting bug example | 54‑95, 70‑84, 89‑92 |
| [`CLAUDE.md`](https://github.com/multica-ai/andrej-karpathy-skills/blob/main/CLAUDE.md) | Contains the full set of coding guidelines enforced for LLM‑assisted development | Full file |

## Summary

- When **test‑first verification finds issues**, the repository requires developers to write a failing test that reproduces the bug before touching implementation code.
- The **Goal‑Driven Execution** principle in [`README.md`](https://github.com/multica-ai/andrej-karpathy-skills/blob/main/README.md) (lines 77‑88) mandates that every change is justified by a concrete, observable outcome.
- The verification loop consists of writing a failing test, confirming the failure, implementing the minimal fix, and re‑running the full suite to prevent regressions.
- [`EXAMPLES.md`](https://github.com/multica-ai/andrej-karpathy-skills/blob/main/EXAMPLES.md) (lines 54‑95) demonstrates this workflow with a real sorting bug, showing how a stable sort implementation satisfies the test criteria.

## Frequently Asked Questions

### What is the Goal‑Driven Execution principle?

The **Goal‑Driven Execution** principle is a core guideline in the multica‑ai/andrej‑karpathy‑skills repository that forces a test‑first workflow. According to [`README.md`](https://github.com/multica-ai/andrej-karpathy-skills/blob/main/README.md) (lines 77‑88), it requires developers to define success criteria as automated tests before writing implementation code, ensuring every change is measurable and verifiable.

### How does the test‑first workflow prevent regressions?

When test‑first verification finds issues, developers must run the **full test suite** after implementing a fix. This step, documented in [`EXAMPLES.md`](https://github.com/multica-ai/andrej-karpathy-skills/blob/main/EXAMPLES.md) (lines 89‑92), ensures that the new code satisfies the original failing test without breaking existing functionality. If any test fails, the developer iterates until the entire suite passes.

### Where are the coding guidelines documented?

The complete coding guidelines are located in [`CLAUDE.md`](https://github.com/multica-ai/andrej-karpathy-skills/blob/main/CLAUDE.md) at the root of the repository. This file defines the standards for LLM‑assisted development, while the specific **Goal‑Driven Execution** rules are detailed in [`README.md`](https://github.com/multica-ai/andrej-karpathy-skills/blob/main/README.md) (lines 77‑88) and illustrated with concrete examples in [`EXAMPLES.md`](https://github.com/multica-ai/andrej-karpathy-skills/blob/main/EXAMPLES.md) (lines 54‑95).

### What happens if the test still fails after the fix?

If the test continues to fail after an attempted fix, the verification loop repeats. The developer revisits the implementation, adds edge‑case checks, or adjusts the test criteria until the failure is resolved. As shown in the sorting example from [`EXAMPLES.md`](https://github.com/multica-ai/andrej-karpathy-skills/blob/main/EXAMPLES.md), this iterative process continues until the test passes and the full suite remains green.