How to Define Success Criteria for Ambiguous Requests: A Goal-Driven Execution Approach
Define success criteria for ambiguous requests by converting vague requirements into explicit, testable metrics using the four-step Goal-Driven Execution loop codified in the forrestchang/andrej-karpathy-skills repository.
When you receive ambiguous instructions like "make the feature faster" or "fix the authentication system," you risk implementing solutions that miss the actual problem or introduce over-engineering. The repository's Goal-Driven Execution principle, documented in README.md (section 7.1) and reinforced in skills/karpathy-guidelines/SKILL.md (section 5.1), provides a systematic workflow to clarify dimensions, capture current behavior in failing tests, and define precise metrics before writing implementation code.
The Four-Step Goal-Driven Execution Workflow
The methodology institutionalized in the forrestchang/andrej-karpathy-skills repository eliminates hidden assumptions by enforcing a strict clarification-to-verification loop. Each step includes an explicit verify: checkpoint that determines when the criteria are satisfied and the loop can terminate.
Step 1: Ask Clarifying Questions to Identify Dimensions
Before touching implementation code, interrogate the request to expose measurable dimensions. For performance tasks, clarify whether the concern is latency, throughput, or resource utilization; for authentication bugs, identify whether the issue involves session state, credential validation, or access control. This step transforms "make it faster" into "reduce average query response time" or "invalidate sessions on password change."
Step 2: Write a Failing Test Capturing Current Behavior
Capture the existing system's behavior in a test that will fail until the success criteria are met. This creates an objective baseline that prevents scope creep. According to the guidelines in SKILL.md, this test must fail under current conditions and pass only when the concrete metric is achieved, serving as the formal definition of "done."
Step 3: Define Numeric Success Metrics
Translate clarified dimensions into verifiable thresholds. Effective criteria include specifics like "response time < 100ms for 95th percentile of requests" or "return HTTP 401 for all sessions created before password change." These metrics appear in the verify: checkpoints of the execution loop described in README.md, allowing the LLM or developer to stop precisely when the ambiguous request is satisfied.
Step 4: Iterate with Surgical Implementation
Implement the smallest possible change to make the failing test pass, then re-run the full test suite to ensure no regressions occur in existing functionality. The repository emphasizes that this loop continues until all verify: conditions return true, at which point the success criteria are definitively met without excess engineering.
Practical Code Examples: From Vague to Verifiable
The following examples from the repository's documentation demonstrate how to apply this framework to common ambiguous requests in production codebases.
Example 1: Converting "Make Search Faster" into Latency < 100ms
When asked to improve search speed without specific targets, define the dimension as query latency and implement a failing performance test that encodes the success criterion:
import time
import pytest
from myapp.search import search
@pytest.mark.parametrize("query", ["cat", "dog", "bird"])
def test_search_latency(query):
"""Verify that search responds within acceptable time bounds."""
start = time.perf_counter()
results = search(query)
elapsed_ms = (time.perf_counter() - start) * 1000
# Success criterion: query must complete under 100ms
assert elapsed_ms < 100, f"Search took {elapsed_ms:.1f}ms, exceeds 100ms threshold"
assert results # Ensure functional correctness alongside performance
After confirming this test fails against the current implementation, apply the smallest optimization—such as adding a database index—until the assertion passes. This satisfies the verify: condition for the latency criterion.
Example 2: Converting "Fix Authentication" into Session Invalidation Logic
For vague authentication repair requests, clarify the specific security dimension and write a test that captures the exact vulnerability:
def test_password_change_invalidates_sessions(client, user):
"""Verify that changing password invalidates all existing sessions."""
# Establish authenticated session
client.login(user)
old_session = client.cookies["session"]
# Execute password change
client.post("/change-password", {"new_password": "new_secure_password123"})
# Attempt to use old session cookie
client.cookies["session"] = old_session
response = client.get("/protected")
# Success criterion: old sessions must return 401 Unauthorized
assert response.status_code == 401, "Session remained valid after password change"
Implementation proceeds by modifying the session store invalidation logic in the authentication module until this test passes, satisfying the security criteria without ambiguous "fix auth" scope creep.
Source Code References
The Goal-Driven Execution framework is implemented across several canonical files in the forrestchang/andrej-karpathy-skills repository:
README.md(section 7.1): Contains the complete four-step workflow description and theverify:checkpoint philosophy for defining concrete success criteria before coding.skills/karpathy-guidelines/SKILL.md(section 5.1): Provides the machine-readable policy that LLMs import to enforce test-first clarification of ambiguous requests.CLAUDE.md: Serves as the single-file policy reference for enforcing these guidelines across distributed projects.EXAMPLES.md: Contains real-world illustrations of converting vague requirements into verifiable test cases, useful as templates for new tasks.
Summary
- Clarify dimensions before coding: Always interrogate ambiguous requests to expose measurable dimensions like latency, throughput, or security boundaries before writing implementation code.
- Encode criteria in failing tests: Capture current behavior in a test that will only pass when explicit numeric criteria (e.g., < 100ms, HTTP 401) are met, establishing an objective definition of success.
- Iterate surgically: Implement the smallest change that satisfies the test, then verify no regressions occur in existing test suites before considering the request complete.
- Reference canonical sources: The forrestchang/andrej-karpathy-skills repository codifies these principles in
README.mdandSKILL.mdto prevent hidden assumptions and over-engineering.
Frequently Asked Questions
What distinguishes a "concrete" success criterion from a vague requirement?
A concrete success criterion contains a measurable threshold and an automated verification method. "Make search faster" is vague; "reduce p95 query latency to under 100 milliseconds as verified by test_search_latency" is concrete. The repository's README.md mandates that every criterion must be expressible as a boolean check in a failing test that turns green upon implementation.
How should you handle requests where stakeholders cannot define specific metrics?
Apply the clarification protocol from SKILL.md to expose implicit expectations. Ask which user actions feel "slow" or what specific security events constitute a "bug," then propose candidate metrics (e.g., "page load under 2 seconds") for stakeholder validation. The Goal-Driven Execution framework explicitly accommodates this negotiation phase as Step 1 before any code is written.
Can Goal-Driven Execution apply to qualitative requirements like UX improvements?
Yes. Convert qualitative feedback into quantitative proxies—such as "reduce time-to-interactive below 3 seconds" or "achieve 90% completion rate in user onboarding flows." The guidelines in SKILL.md encourage defining observability metrics that correlate with user experience, then writing tests against those instrumentation points rather than subjective "look and feel" criteria.
Where is the Goal-Driven Execution framework maintained in the source code?
The framework is primarily documented in README.md at section 7.1, with a machine-readable enforcement policy located in skills/karpathy-guidelines/SKILL.md at section 5.1. The repository forrestchang/andrej-karpathy-skills maintains these files as the canonical reference for transforming ambiguous LLM or human requests into verifiable success criteria loops.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →