# Real-World Performance Impact of Test-First Verification

> Discover the real-world performance impact of test-first verification. Reduce latency spikes and CPU usage with this agile approach for measurable improvements.

- Repository: [multica-ai/andrej-karpathy-skills](https://github.com/multica-ai/andrej-karpathy-skills)
- Tags: testing
- Published: 2026-04-19

---

**Test-first verification forces minimal, repeatable changes that prevent speculative optimizations and catch performance regressions early, resulting in measurable reductions in latency spikes and CPU usage.**

The **multica-ai/andrej-karpathy-skills** repository operationalizes Andrej Karpathy's software engineering principles through concrete coding patterns. At its core lies **Goal-Driven Execution**, a methodology that mandates writing a failing test before touching implementation code. This article examines how this test-first approach directly impacts real-world performance characteristics in production systems.

## What Is Test-First Verification?

Test-first verification is the practical implementation of the **Goal-Driven Execution** principle defined in [`skills/karpathy-guidelines/SKILL.md`](https://github.com/multica-ai/andrej-karpathy-skills/blob/main/skills/karpathy-guidelines/SKILL.md). The pattern requires developers to convert every user request into a concrete, verifiable goal—typically a failing test that reproduces a bug or defines desired behavior—before writing implementation code.

According to the repository's guidelines documented in [`CLAUDE.md`](https://github.com/multica-ai/andrej-karpathy-skills/blob/main/CLAUDE.md) and [`README.md`](https://github.com/multica-ai/andrej-karpathy-skills/blob/main/README.md), the cycle follows six strict steps:

1.  **User Story → Goal**: Express the request as a concrete success criterion.
2.  **Write Failing Test**: Create a minimal test reproducing the issue or asserting a performance metric.
3.  **Run Tests → Confirm Failure**: Verify the test isolates the problem.
4.  **Implement Change**: Add only the code needed to make the test pass.
5.  **Run Tests → Verify Pass**: Accept the change only if the test passes.
6.  **Iterate**: Add new goals as separate tests.

## How Test-First Verification Impacts Performance

The repository explicitly links test-first patterns to performance outcomes through three primary mechanisms.

### Preventing Premature Optimization

By forcing minimal, repeatable changes, test-first verification prevents speculative optimizations that degrade runtime performance. When developers must justify every change against a specific failing test, they avoid adding complex caching layers or concurrency mechanisms that introduce hidden latency. The guidelines in [`SKILL.md`](https://github.com/multica-ai/andrej-karpathy-skills/blob/main/SKILL.md) emphasize that only code required to pass the test is implemented, eliminating bloat that impacts CPU usage and memory allocation.

### Enforcing Performance Contracts

Tests can encode explicit performance constraints. The repository demonstrates this pattern in [`EXAMPLES.md`](https://github.com/multica-ai/andrej-karpathy-skills/blob/main/EXAMPLES.md), where timing assertions serve as regression guards. Any implementation change that violates the time budget fails the build immediately, forcing developers to address slowdowns before they reach production.

### Reducing Regression Risk

Automated regression testing catches performance degradations early. When every feature is protected by a test that runs on every CI build, inadvertent slowdowns from seemingly unrelated changes are flagged immediately. This creates a predictable performance baseline that prevents latency spikes from accumulating across sprints.

## Real-World Implementation: The Sorting Bug Example

The repository provides a concrete demonstration in [`EXAMPLES.md`](https://github.com/multica-ai/andrej-karpathy-skills/blob/main/EXAMPLES.md) titled "Sorting breaks when there are duplicate scores." This example illustrates how test-first verification preserves deterministic performance characteristics.

### The Failing Test

First, the developer writes a test exposing non-deterministic ordering:

```python

# test_sort_with_duplicate_scores.py

def test_sort_with_duplicate_scores():
    """Reproduces non‑deterministic ordering when scores are duplicated."""
    scores = [
        {"name": "Alice",   "score": 100},
        {"name": "Bob",     "score": 100},
        {"name": "Charlie", "score": 90},
    ]

    result = sort_scores(scores)

    # The bug: ordering of equal scores is nondeterministic

    assert result[0]["score"] == 100
    assert result[1]["score"] == 100
    assert result[2]["score"] == 90

```

### The Buggy Implementation

The initial implementation uses an unstable sort:

```python

# buggy implementation (unstable)

def sort_scores(scores):
    return sorted(scores, key=lambda x: -x["score"])

```

Running the test first confirms the failure—ordering of equal elements is not guaranteed, potentially causing extra sorting passes or cache misses in production data pipelines.

### The Fix

The corrected implementation adds a tie-breaker to ensure stable ordering:

```python

# stable implementation (passes the test)

def sort_scores(scores):
    """Sort by descending score, then ascending name for ties."""
    return sorted(scores, key=lambda x: (-x["score"], x["name"]))

```

By verifying the fix through the existing test, the developer guarantees deterministic performance. The stable sort prevents subtle regressions where unstable algorithms might trigger additional passes or unpredictable memory access patterns.

## Encoding Performance Constraints in Tests

Beyond functional correctness, the repository demonstrates encoding hard performance requirements directly into test assertions. This practice ensures that optimizations remain intact across refactoring cycles.

Consider the rate limiter example from the codebase:

```python
import time

def test_rate_limiter_performance():
    """Ensure in‑memory rate limiter handles 10 000 requests < 50 ms."""
    limiter = MemoryRateLimiter(limit=100, window_seconds=1)

    start = time.perf_counter()
    for _ in range(10_000):
        limiter.allow()   # should be fast

    elapsed_ms = (time.perf_counter() - start) * 1000

    assert elapsed_ms < 50, f"Too slow: {elapsed_ms:.2f} ms"

```

This test establishes a **performance contract**. If a future developer switches to a slower data structure or adds a heavy locking mechanism, the CI pipeline fails immediately. The test-first approach forces teams to acknowledge and justify any performance degradation before it impacts users.

## Repository Architecture Supporting Test-First Performance

The **multica-ai/andrej-karpathy-skills** repository embeds these principles into its tooling and documentation structure. The following files enforce the test-first pattern that drives performance reliability:

- **[`README.md`](https://github.com/multica-ai/andrej-karpathy-skills/blob/main/README.md)**: Documents the four principles including Goal-Driven Execution, establishing the philosophical foundation for test-first verification.
- **[`CLAUDE.md`](https://github.com/multica-ai/andrej-karpathy-skills/blob/main/CLAUDE.md)**: Provides the instruction set for Claude Code agents, operationalizing the test-first workflow for AI-assisted development.
- **[`skills/karpathy-guidelines/SKILL.md`](https://github.com/multica-ai/andrej-karpathy-skills/blob/main/skills/karpathy-guidelines/SKILL.md)**: Contains the portable skill definition used by both Claude and Cursor, specifying the Goal-Driven Execution principle in a reusable format.
- **[`EXAMPLES.md`](https://github.com/multica-ai/andrej-karpathy-skills/blob/main/EXAMPLES.md)**: Provides the concrete "Sorting breaks when there are duplicate scores" demonstration linking test-first patterns to performance outcomes.
- **`.cursor/rules/karpathy-guidelines.mdc`**: Applies the guidelines automatically within the Cursor IDE, ensuring test-first verification happens at the point of code creation.

This architectural approach ensures that test-first verification is not merely suggested but enforced by the development environment itself, making performance regressions statistically less likely.

## Summary

- **Test-first verification** requires writing a failing test before implementing code changes, as mandated by the Goal-Driven Execution principle in `multica-ai/andrej-karpathy-skills`.
- This pattern **prevents premature optimization** by forcing minimal, justified changes that avoid speculative complexity.
- Tests can encode **performance contracts** with timing assertions, automatically flagging regressions in CI pipelines.
- The **sorting bug example** in [`EXAMPLES.md`](https://github.com/multica-ai/andrej-karpathy-skills/blob/main/EXAMPLES.md) demonstrates how test-first approaches guarantee deterministic performance and prevent non-deterministic ordering issues.
- Repository architecture including [`CLAUDE.md`](https://github.com/multica-ai/andrej-karpathy-skills/blob/main/CLAUDE.md), [`SKILL.md`](https://github.com/multica-ai/andrej-karpathy-skills/blob/main/SKILL.md), and `.cursor/rules/karpathy-guidelines.mdc` **enforces** this workflow, making performance reliability a built-in outcome.

## Frequently Asked Questions

### What is the real-world performance impact of test-first verification?

The real-world performance impact is **measurable reduction in latency spikes and CPU usage**. By preventing speculative optimizations and enforcing minimal changes, test-first verification eliminates hidden performance costs from premature complexity. The `multica-ai/andrej-karpathy-skills` repository documents cases where this pattern preserves deterministic performance characteristics, such as stable sorting algorithms that prevent cache misses and extra passes.

### How does test-first verification prevent performance regressions?

Test-first verification prevents regressions by encoding **performance contracts directly into the test suite**. As shown in the rate limiter example from the repository, tests can assert specific timing requirements (e.g., "10,000 requests must complete in under 50ms"). When a developer introduces a slower implementation, the failing test forces immediate correction before the code reaches production, creating an automated safety net against performance degradation.

### Can test-first verification slow down development?

While writing tests requires initial time investment, the repository guidelines argue that **test-first verification accelerates overall delivery** by eliminating debugging cycles and rework. The Goal-Driven Execution principle documented in [`SKILL.md`](https://github.com/multica-ai/andrej-karpathy-skills/blob/main/SKILL.md) ensures that developers know exactly when a feature is complete (when the test passes), preventing gold-plating and speculative features that consume development time without adding value. The performance benefits—faster feedback loops and reduced regression fixing—further improve velocity.

### What is Goal-Driven Execution in the context of this repository?

**Goal-Driven Execution** is the foundational principle in `multica-ai/andrej-karpathy-skills` that mandates converting every user request into a concrete, verifiable goal before writing implementation code. According to [`README.md`](https://github.com/multica-ai/andrej-karpathy-skills/blob/main/README.md) and [`CLAUDE.md`](https://github.com/multica-ai/andrej-karpathy-skills/blob/main/CLAUDE.md), this means creating a failing test that reproduces the bug or defines desired behavior, then implementing only the minimal code necessary to pass that test. This principle operationalizes test-first verification and ensures that performance characteristics are defined as explicit requirements rather than afterthoughts.