What Happens When Test‑First Verification Finds Issues in Andrej Karpathy Skills

When test‑first verification finds issues in the multica‑ai/andrej‑karpathy‑skills repository, developers capture the bug in a failing test, iterate on the implementation until the test passes, and verify the entire suite remains green to prevent regressions.

The multica‑ai/andrej‑karpathy‑skills repository enforces a rigorous Goal‑Driven Execution principle that mandates a test‑first workflow whenever issues are discovered. According to the source code in README.md (lines 77‑88), this approach treats every bug as a requirement for a failing test that must pass before the implementation is considered complete.

The Goal‑Driven Execution Verification Loop

When test‑first verification finds issues, the repository dictates a four‑step verification loop that turns failures into concrete success criteria. This loop is documented in EXAMPLES.md (lines 54‑95) and ensures that every fix is provably correct.

Step 1: Write a Failing Test That Reproduces the Issue

The developer first writes a test that explicitly reproduces the discovered issue. This test serves as the success criterion for the bug—if the test passes, the issue is resolved. In EXAMPLES.md (lines 70‑84), the test_sort_with_duplicate_scores function demonstrates this by asserting deterministic ordering for items with identical scores.

Step 2: Run the Test to Confirm the Failure

The test is executed to confirm it fails, validating that the issue is real and that the test correctly captures the bug. This step prevents false positives and ensures the baseline is honest.

Step 3: Implement the Minimal Fix

With a reproducible failure in place, the developer implements the smallest possible change that makes the test pass. In the repository’s example, this meant replacing an unstable sort with a stable sort that uses a composite key to break ties deterministically.

Step 4: Re‑run the Full Test Suite

After the fix, the developer runs the original test again—it must now pass—and then executes the entire test suite to ensure no regressions were introduced. If any test fails, the loop repeats until the criteria are satisfied.

Real‑World Example: Fixing a Nondeterministic Sort in EXAMPLES.md

The repository provides a concrete scenario in EXAMPLES.md (lines 54‑95) where test‑first verification finds issues with duplicate score handling. The initial implementation produced nondeterministic ordering, which the test surfaced immediately.

The Failing Test

def test_sort_with_duplicate_scores():
    """
    Test sorting when multiple items have the same score.
    """
    scores = [
        {'name': 'Alice',   'score': 100},
        {'name': 'Bob',     'score': 100},
        {'name': 'Charlie', 'score': 90},
    ]

    result = sort_scores(scores)

    # The bug: order is nondeterministic for duplicates

    # Run this test many times – it should be consistent

    assert result[0]['score'] == 100
    assert result[1]['score'] == 100
    assert result[2]['score'] == 90

The test fails because the original sort_scores does not guarantee a deterministic order for equal scores (see lines 70‑84 of EXAMPLES.md).

The Fix After Verification

Once the failure confirmed the bug, the developer implemented a stable sort using a composite key:

def sort_scores(scores):
    """Sort by score descending, then name ascending for ties."""
    return sorted(scores, key=lambda x: (-x['score'], x['name']))

Running the full test suite now shows the new implementation passes the reproducibility check (see lines 89‑92 of EXAMPLES.md).

Key Source Files and Their Roles

File Role Key Lines
README.md Defines the Goal‑Driven Execution principle that mandates the test‑first workflow 77‑88
EXAMPLES.md Provides concrete demonstrations of the verification loop, including the sorting bug example 54‑95, 70‑84, 89‑92
CLAUDE.md Contains the full set of coding guidelines enforced for LLM‑assisted development Full file

Summary

  • When test‑first verification finds issues, the repository requires developers to write a failing test that reproduces the bug before touching implementation code.
  • The Goal‑Driven Execution principle in README.md (lines 77‑88) mandates that every change is justified by a concrete, observable outcome.
  • The verification loop consists of writing a failing test, confirming the failure, implementing the minimal fix, and re‑running the full suite to prevent regressions.
  • EXAMPLES.md (lines 54‑95) demonstrates this workflow with a real sorting bug, showing how a stable sort implementation satisfies the test criteria.

Frequently Asked Questions

What is the Goal‑Driven Execution principle?

The Goal‑Driven Execution principle is a core guideline in the multica‑ai/andrej‑karpathy‑skills repository that forces a test‑first workflow. According to README.md (lines 77‑88), it requires developers to define success criteria as automated tests before writing implementation code, ensuring every change is measurable and verifiable.

How does the test‑first workflow prevent regressions?

When test‑first verification finds issues, developers must run the full test suite after implementing a fix. This step, documented in EXAMPLES.md (lines 89‑92), ensures that the new code satisfies the original failing test without breaking existing functionality. If any test fails, the developer iterates until the entire suite passes.

Where are the coding guidelines documented?

The complete coding guidelines are located in CLAUDE.md at the root of the repository. This file defines the standards for LLM‑assisted development, while the specific Goal‑Driven Execution rules are detailed in README.md (lines 77‑88) and illustrated with concrete examples in EXAMPLES.md (lines 54‑95).

What happens if the test still fails after the fix?

If the test continues to fail after an attempted fix, the verification loop repeats. The developer revisits the implementation, adds edge‑case checks, or adjusts the test criteria until the failure is resolved. As shown in the sorting example from EXAMPLES.md, this iterative process continues until the test passes and the full suite remains green.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →