4 Common LLM Coding Mistakes the Karpathy Guidelines Prevent
The Karpathy Guidelines prevent LLMs from fabricating APIs, over-engineering solutions, introducing unintended side effects, and shipping unverified code by enforcing explicit assumptions, minimal implementations, surgical precision, and goal-driven validation.
The multica-ai/andrej-karpathy-skills repository codifies a strict behavioral framework for large language models writing production code. These guidelines, documented in skills/karpathy-guidelines/SKILL.md and integrated into the project's CLAUDE.md, specifically target the most frequent LLM coding mistakes that degrade software quality, introduce subtle bugs, and create maintenance nightmares.
What Are the Karpathy Guidelines?
The guidelines consist of four imperative rules that govern how an LLM must behave when generating or modifying code:
- Think Before Coding – Explicitly state assumptions, ask clarifying questions, and surface trade-offs before writing any code.
- Simplicity First – Write the minimum code necessary to solve the problem; avoid speculative abstractions and unnecessary configuration.
- Surgical Changes – Modify only what is required, remove any new dead code created by the edit, and avoid refactoring unrelated working sections.
- Goal-Driven Execution – Define a verifiable success criterion (e.g., a test) and iterate until that criterion is satisfied.
The 4 Critical LLM Coding Mistakes Prevented
1. Fabrication and Hidden Assumptions (Prevented by "Think Before Coding")
LLMs frequently hallucinate APIs or assume implementation details that do not exist in the codebase. Without explicit verification, a model might generate calls to non-existent endpoints or import modules that were never defined.
The "Think Before Coding" guideline forces the model to enumerate all assumptions in comments or explanatory text before generating implementation. For example, instead of immediately writing requests.get("https://api.example.com/fetch_user?id=123"), the model must first state: "Assuming the endpoint is /user not /fetch_user; confirm if different."
This prevents the common LLM coding mistake of silent fabrication, where wrong interfaces are buried in otherwise syntactically correct code.
2. Over-Engineering and Unnecessary Abstraction (Prevented by "Simplicity First")
Large language models often produce speculative architectures—complete class hierarchies, factory patterns, and configuration systems—for trivial one-off tasks. This over-engineering introduces maintenance burden, cognitive complexity, and potential failure points where none are needed.
The "Simplicity First" rule mandates that the model write only the minimum code required to satisfy the specific request. If the task is converting a string to camelCase, the model must provide a single function, not a StringConverter class with plugin architecture.
This guideline directly prevents the LLM coding mistake of excessive abstraction, keeping solutions lean, readable, and directly aligned with the problem scope.
3. Unintended Side Effects and Code Rot (Prevented by "Surgical Changes")
When modifying existing files, LLMs often perform bulk "beautification"—reformatting entire modules, reorganizing imports across the file, or refactoring unrelated functions while making a targeted fix. These collateral changes introduce noise in version control, obscure the actual fix, and risk breaking working code.
The "Surgical Changes" guideline requires the model to touch only what is necessary for the specific task. It must not refactor unrelated sections, must remove any new dead code (such as orphaned imports) created by its edit, and must preserve the existing structure of untouched code.
This prevents the common LLM coding mistake of scope creep and collateral damage, ensuring that diffs are minimal, reviewable, and safe.
4. Unverified Solutions Without Success Criteria (Prevented by "Goal-Driven Execution")
LLMs frequently generate code that appears correct but has no defined validation mechanism. Without a specific test or success criterion, the model cannot verify that its solution actually satisfies the requirements, leading to implementations that compile but fail in production.
The "Goal-Driven Execution" rule forces the model to define a verifiable success criterion before writing code—typically a test case or specific output validation. The model must then iterate on its solution until that criterion passes, ensuring that the final code is proven to work rather than merely syntactically plausible.
This prevents the critical LLM coding mistake of untested deliverables, ensuring that all generated code has validated behavior against concrete requirements.
Practical Examples: Bad vs. Good LLM Output
Example 1: Preventing API Fabrication
Bad (Violates "Think Before Coding"):
# LLM assumes facts without verification
response = requests.get("https://api.example.com/fetch_user?id=123")
data = response.json()
Good (Follows Guidelines):
# ✅ Explicitly state assumption, ask if unsure
# Assume the service provides /user?id=...; confirm if different.
endpoint = "https://api.example.com/user"
params = {"id": 123}
response = requests.get(endpoint, params=params)
response.raise_for_status()
data = response.json()
Example 2: Preventing Over-Engineering
Bad (Violates "Simplicity First"):
// Over-engineered class for a one-off conversion
class Converter {
constructor() { /* ... */ }
toCamelCase(str) { /* ... */ }
toSnakeCase(str) { /* ... */ }
}
export default new Converter();
Good (Follows Guidelines):
// ✅ Minimal solution – one function does the job
export const toCamelCase = (str) =>
str.replace(/_([a-z])/g, (_, c) => c.toUpperCase());
Example 3: Preventing Unintended Side Effects
Bad (Violates "Surgical Changes"):
# Original file has orphaned imports after edit
import json, os, sys # sys is now unused
def load_config(path):
with open(path) as f:
return json.load(f)
# LLM added this but forgot to clean up imports
def normalize_path(p):
return os.path.normpath(p)
Good (Follows Guidelines):
# ✅ Surgical change with cleanup
import json, os # removed unused sys
def load_config(path):
with open(path) as f:
return json.load(f)
def normalize_path(p):
return os.path.normpath(p)
Where to Find the Official Guidelines
The complete behavioral rules are maintained in the multica-ai/andrej-karpathy-skills repository:
skills/karpathy-guidelines/SKILL.md– Contains the full text of the four guidelines (lines 13-67), including detailed explanations of each rule and its intent.CLAUDE.md– Project-wide instruction set that merges the Karpathy guidelines with repository-specific rules, ensuring consistent application across all generated code.README.md– Overview of the repository structure and instructions for invoking the guidelines in your development workflow.
Summary
The Karpathy Guidelines prevent the four most damaging categories of LLM coding mistakes:
- Fabrication and hallucination are blocked by forcing explicit assumptions and clarifying questions before any code is written.
- Over-engineering and bloat are eliminated by mandating minimal solutions that solve only the stated problem.
- Collateral damage and code rot are prevented by requiring surgical precision—touching only necessary lines and cleaning up dead code.
- Untested, unverified deliverables are stopped by requiring concrete success criteria and validation loops before finalizing code.
Frequently Asked Questions
What are the Karpathy Guidelines for LLM coding?
The Karpathy Guidelines are a set of four behavioral rules—Think Before Coding, Simplicity First, Surgical Changes, and Goal-Driven Execution—designed to steer large language models away from common coding pitfalls. They are documented in the multica-ai/andrej-karpathy-skills repository and enforce a defensive, precise approach to code generation.
How do the Karpathy Guidelines prevent LLM hallucinations in code?
The "Think Before Coding" rule prevents hallucinations by forcing the model to enumerate all assumptions explicitly before writing implementation. If the model is unsure about an API endpoint, library function, or data structure, it must state that uncertainty rather than fabricating a plausible-sounding but incorrect solution. This surfaces potential errors before they become committed code.
Can these guidelines be applied to non-LLM software development?
Yes, while designed for LLM behavior, the Karpathy Guidelines serve as excellent general software engineering principles. Simplicity First aligns with YAGNI (You Aren't Gonna Need It), Surgical Changes reflects the boy scout rule in reverse (don't touch what you don't need to fix), and Goal-Driven Execution enforces test-driven development practices. Human developers can adopt these rules to reduce technical debt and improve code review efficiency.
Where are the Karpathy Guidelines located in the repository?
The primary source is skills/karpathy-guidelines/SKILL.md, which contains the complete guideline text and detailed explanations (specifically lines 13-67). The guidelines are also integrated into CLAUDE.md, which serves as the project-wide instruction set for Claude-based workflows. Both files reside in the root and skills/karpathy-guidelines/ directories of the multica-ai/andrej-karpathy-skills repository.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →