# How to Create Test Cases That Reproduce Bugs Before Fixing: A Test-First Workflow

> Learn to write failing tests that reproduce bugs before fixing them. Implement minimal code changes and establish a regression safety net with this test-first workflow.

- Repository: [Jiayuan Zhang/andrej-karpathy-skills](https://github.com/forrestchang/andrej-karpathy-skills)
- Tags: best-practices
- Published: 2026-04-08

---

**Write a failing test that isolates the exact bug behavior, verify it fails, implement the minimal code change, and confirm the test passes to establish a regression safety net.**

The *andrej-karpathy-skills* repository by forrestchang codifies a disciplined, LLM-assisted development workflow centered on **Goal-Driven Execution**. This approach mandates that every bug fix begins with creating a deterministic, reproducible test case that captures the failure before any code changes are made.

## The Goal-Driven Execution Principle

At the heart of this methodology lies a single transformative rule defined in [`skills/karpathy-guidelines/SKILL.md`](https://github.com/forrestchang/andrej-karpathy-skills/blob/main/skills/karpathy-guidelines/SKILL.md). The **Goal-Driven Execution** section explicitly maps the vague instruction "Fix the bug" to the concrete action: "Write a test that reproduces it, then make it pass"【https://github.com/forrestchang/andrej-karpathy-skills/blob/main/skills/karpathy-guidelines/SKILL.md#goal-driven-execution】.

This principle is reinforced across the repository. The top-level [`README.md`](https://github.com/forrestchang/andrej-karpathy-skills/blob/main/README.md) reiterates the same transformation【https://github.com/forrestchang/andrej-karpathy-skills/blob/main/README.md#goal-driven-execution】, while [`CLAUDE.md`](https://github.com/forrestchang/andrej-karpathy-skills/blob/main/CLAUDE.md) serves as a policy file you can inject into any project to enforce this test-first discipline automatically【https://github.com/forrestchang/andrej-karpathy-skills/blob/main/CLAUDE.md#goal-driven-execution】.

## Step-by-Step Workflow to Create Bug-Reproducing Tests

Following the architecture outlined in the Karpathy Guidelines, here is the rigorous process to **create test cases that reproduce bugs before fixing** them:

### 1. Isolate and Identify the Failure

Capture the observable bug in isolation. Document the exact exception, incorrect output, or flaky behavior. The test must target this specific failure mode without external dependencies like databases or file systems.

### 2. Write a Minimal, Deterministic Test

Construct a test that fails deterministically every time the bug is present. According to the repository's guidelines, the test should define its own fixture data and assert the **incorrect** behavior currently produced by the buggy code. This ensures the test truly validates the fix later.

### 3. Confirm the Test Fails

Run the test suite to verify the new test fails as expected. This step establishes your safety net—if the test passes before you change any production code, your reproduction is flawed.

### 4. Implement the Minimal Fix

Modify the production code while repeatedly running the test. The [`SKILL.md`](https://github.com/forrestchang/andrej-karpathy-skills/blob/main/SKILL.md) guidelines emphasize making the test pass with the smallest possible change to avoid introducing new regressions.

### 5. Verify and Expand Coverage

Once the test passes, add assertions for edge cases if necessary. Run the full test suite to ensure no existing functionality broke during the fix.

## Real-World Example: Nondeterministic Sorting Bug

The [`EXAMPLES.md`](https://github.com/forrestchang/andrej-karpathy-skills/blob/main/EXAMPLES.md) file provides a concrete demonstration of this pattern in the **Test-First Verification** section (Example 3)【https://github.com/forrestchang/andrej-karpathy-skills/blob/main/EXAMPLES.md#example-3-test-first-verification】. The example addresses a Python sorting function that produces unstable ordering when handling duplicate scores.

First, write the failing test that reproduces the nondeterministic behavior:

```python
def test_sort_with_duplicate_scores():
    """
    The sorting function should produce a stable order when scores are equal.
    """
    scores = [
        {"name": "Alice",   "score": 100},
        {"name": "Bob",     "score": 100},
        {"name": "Charlie", "score": 90},
    ]

    result = sort_scores(scores)

    # The bug: order of equal-score items is nondeterministic

    # Run this test multiple times – it should be consistent

    assert result[0]["score"] == 100
    assert result[1]["score"] == 100
    assert result[2]["score"] == 90

```

Running `pytest` at this stage fails because `sort_scores` either does not exist or returns an unstable order. Next, implement the fix using a stable sort key:

```python
def sort_scores(scores):
    """
    Sort by score descending, then by name ascending for ties.
    """
    return sorted(scores, key=lambda x: (-x["score"], x["name"]))

```

Finally, rerun the test to confirm it passes. The test now verifies that equal scores sort deterministically by name, preventing future regressions where the ordering might become random again.

## Key Repository Files for Test-First Development

 Understanding the provenance of these guidelines helps enforce them in your own projects:

- **[`skills/karpathy-guidelines/SKILL.md`](https://github.com/forrestchang/andrej-karpathy-skills/blob/main/skills/karpathy-guidelines/SKILL.md)** — Contains the core **Goal-Driven Execution** rule that mandates writing reproduction tests before fixes.
- **[`EXAMPLES.md`](https://github.com/forrestchang/andrej-karpathy-skills/blob/main/EXAMPLES.md)** — Provides the concrete Python sorting bug example demonstrating the write-test-first pattern.
- **[`CLAUDE.md`](https://github.com/forrestchang/andrej-karpathy-skills/blob/main/CLAUDE.md)** — A policy file you can add to any AI-assisted project to automatically enforce the test-first workflow.
- **[`README.md`](https://github.com/forrestchang/andrej-karpathy-skills/blob/main/README.md)** — Offers the high-level architectural overview of the four principles, including the test-first transformation.

## Summary

- **Always start with a failing test** that captures the exact bug behavior before touching production code, as mandated in [`skills/karpathy-guidelines/SKILL.md`](https://github.com/forrestchang/andrej-karpathy-skills/blob/main/skills/karpathy-guidelines/SKILL.md).
- **Ensure test isolation** by defining fixtures inline and avoiding external state, ensuring deterministic failure reproduction.
- **Verify the fix** by watching the previously failing test turn green, confirming the change addresses the root cause.
- **Prevent regressions** by keeping the reproduction test in your permanent suite, catching future reintroductions of the same bug automatically.

## Frequently Asked Questions

### Why should I write the test before fixing the bug?

Writing the test first guarantees you understand the failure mode completely and provides an objective verification mechanism. As implemented in the *andrej-karpathy-skills* repository, this practice transforms vague bug reports into concrete, executable specifications that prevent regressions.

### What makes a good bug reproduction test?

A good reproduction test is **minimal**, **deterministic**, and **isolated**. It should fail every time the bug is present and pass every time it is absent, using only in-memory fixtures without external dependencies like databases or network calls.

### How do I handle nondeterministic bugs in tests?

For flaky or timing-dependent bugs, structure your test to assert specific invariants rather than exact sequences when possible. In the [`EXAMPLES.md`](https://github.com/forrestchang/andrej-karpathy-skills/blob/main/EXAMPLES.md) sorting example, the test enforces deterministic ordering by asserting the final position of items, eliminating randomness from the validation.

### Where are the Karpathy Guidelines documented?

The primary guidelines reside in [`skills/karpathy-guidelines/SKILL.md`](https://github.com/forrestchang/andrej-karpathy-skills/blob/main/skills/karpathy-guidelines/SKILL.md) within the forrestchang/andrej-karpathy-skills repository. Supporting documentation appears in [`CLAUDE.md`](https://github.com/forrestchang/andrej-karpathy-skills/blob/main/CLAUDE.md) for AI policy enforcement and [`EXAMPLES.md`](https://github.com/forrestchang/andrej-karpathy-skills/blob/main/EXAMPLES.md) for practical demonstrations.