How to Create Test Cases That Reproduce Bugs Before Fixing: A Test-First Workflow
Write a failing test that isolates the exact bug behavior, verify it fails, implement the minimal code change, and confirm the test passes to establish a regression safety net.
The andrej-karpathy-skills repository by forrestchang codifies a disciplined, LLM-assisted development workflow centered on Goal-Driven Execution. This approach mandates that every bug fix begins with creating a deterministic, reproducible test case that captures the failure before any code changes are made.
The Goal-Driven Execution Principle
At the heart of this methodology lies a single transformative rule defined in skills/karpathy-guidelines/SKILL.md. The Goal-Driven Execution section explicitly maps the vague instruction "Fix the bug" to the concrete action: "Write a test that reproduces it, then make it pass"【https://github.com/forrestchang/andrej-karpathy-skills/blob/main/skills/karpathy-guidelines/SKILL.md#goal-driven-execution】.
This principle is reinforced across the repository. The top-level README.md reiterates the same transformation【https://github.com/forrestchang/andrej-karpathy-skills/blob/main/README.md#goal-driven-execution】, while CLAUDE.md serves as a policy file you can inject into any project to enforce this test-first discipline automatically【https://github.com/forrestchang/andrej-karpathy-skills/blob/main/CLAUDE.md#goal-driven-execution】.
Step-by-Step Workflow to Create Bug-Reproducing Tests
Following the architecture outlined in the Karpathy Guidelines, here is the rigorous process to create test cases that reproduce bugs before fixing them:
1. Isolate and Identify the Failure
Capture the observable bug in isolation. Document the exact exception, incorrect output, or flaky behavior. The test must target this specific failure mode without external dependencies like databases or file systems.
2. Write a Minimal, Deterministic Test
Construct a test that fails deterministically every time the bug is present. According to the repository's guidelines, the test should define its own fixture data and assert the incorrect behavior currently produced by the buggy code. This ensures the test truly validates the fix later.
3. Confirm the Test Fails
Run the test suite to verify the new test fails as expected. This step establishes your safety net—if the test passes before you change any production code, your reproduction is flawed.
4. Implement the Minimal Fix
Modify the production code while repeatedly running the test. The SKILL.md guidelines emphasize making the test pass with the smallest possible change to avoid introducing new regressions.
5. Verify and Expand Coverage
Once the test passes, add assertions for edge cases if necessary. Run the full test suite to ensure no existing functionality broke during the fix.
Real-World Example: Nondeterministic Sorting Bug
The EXAMPLES.md file provides a concrete demonstration of this pattern in the Test-First Verification section (Example 3)【https://github.com/forrestchang/andrej-karpathy-skills/blob/main/EXAMPLES.md#example-3-test-first-verification】. The example addresses a Python sorting function that produces unstable ordering when handling duplicate scores.
First, write the failing test that reproduces the nondeterministic behavior:
def test_sort_with_duplicate_scores():
"""
The sorting function should produce a stable order when scores are equal.
"""
scores = [
{"name": "Alice", "score": 100},
{"name": "Bob", "score": 100},
{"name": "Charlie", "score": 90},
]
result = sort_scores(scores)
# The bug: order of equal-score items is nondeterministic
# Run this test multiple times – it should be consistent
assert result[0]["score"] == 100
assert result[1]["score"] == 100
assert result[2]["score"] == 90
Running pytest at this stage fails because sort_scores either does not exist or returns an unstable order. Next, implement the fix using a stable sort key:
def sort_scores(scores):
"""
Sort by score descending, then by name ascending for ties.
"""
return sorted(scores, key=lambda x: (-x["score"], x["name"]))
Finally, rerun the test to confirm it passes. The test now verifies that equal scores sort deterministically by name, preventing future regressions where the ordering might become random again.
Key Repository Files for Test-First Development
Understanding the provenance of these guidelines helps enforce them in your own projects:
skills/karpathy-guidelines/SKILL.md— Contains the core Goal-Driven Execution rule that mandates writing reproduction tests before fixes.EXAMPLES.md— Provides the concrete Python sorting bug example demonstrating the write-test-first pattern.CLAUDE.md— A policy file you can add to any AI-assisted project to automatically enforce the test-first workflow.README.md— Offers the high-level architectural overview of the four principles, including the test-first transformation.
Summary
- Always start with a failing test that captures the exact bug behavior before touching production code, as mandated in
skills/karpathy-guidelines/SKILL.md. - Ensure test isolation by defining fixtures inline and avoiding external state, ensuring deterministic failure reproduction.
- Verify the fix by watching the previously failing test turn green, confirming the change addresses the root cause.
- Prevent regressions by keeping the reproduction test in your permanent suite, catching future reintroductions of the same bug automatically.
Frequently Asked Questions
Why should I write the test before fixing the bug?
Writing the test first guarantees you understand the failure mode completely and provides an objective verification mechanism. As implemented in the andrej-karpathy-skills repository, this practice transforms vague bug reports into concrete, executable specifications that prevent regressions.
What makes a good bug reproduction test?
A good reproduction test is minimal, deterministic, and isolated. It should fail every time the bug is present and pass every time it is absent, using only in-memory fixtures without external dependencies like databases or network calls.
How do I handle nondeterministic bugs in tests?
For flaky or timing-dependent bugs, structure your test to assert specific invariants rather than exact sequences when possible. In the EXAMPLES.md sorting example, the test enforces deterministic ordering by asserting the final position of items, eliminating randomness from the validation.
Where are the Karpathy Guidelines documented?
The primary guidelines reside in skills/karpathy-guidelines/SKILL.md within the forrestchang/andrej-karpathy-skills repository. Supporting documentation appears in CLAUDE.md for AI policy enforcement and EXAMPLES.md for practical demonstrations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →