How to Debug When LLM Ignores Guidelines: A Step-by-Step Guide to the Karpathy Skills

When an LLM violates coding guidelines, the root cause is almost always that the policy file isn't loaded, the user request lacks critical context, or success criteria are undefined, all of which you can resolve by verifying the CLAUDE.md plugin and explicitly invoking the four Karpathy principles in your prompts.

The forrestchang/andrej-karpathy-skills repository encodes a compact set of behavioral rules that steer LLMs away from hidden assumptions and over-engineering. When you notice an LLM producing output that ignores these constraints—such as refactoring unrelated code or omitting success tests—you need a structured debugging approach to restore alignment with the guidelines.

Understanding the Repository Architecture

Before debugging, confirm you understand which files constitute the knowledge base. The repository contains four critical components:

The Policy Engine (CLAUDE.md)

The CLAUDE.md file at the repository root serves as the single-file policy engine that Claude (or any LLM) loads as a plugin. It enumerates the four principles—Think Before Coding, Simplicity First, Surgical Changes, and Goal-Driven Execution—and provides the behavioral constraints the model must follow.

The Skill Manifest (skills/karpathy-guidelines/SKILL.md)

This file declares the skill name, description, and the same four principles in a machine-readable YAML front-matter block, allowing plugins or toolchains to discover the guidelines automatically.

Examples Library (EXAMPLES.md)

This file contains a curated library of real-world before/after snippets that illustrate each principle and the typical errors LLMs make when violating them.

Installation Reference (README.md)

The README.md provides the installation snippet and a "how to know it's working" checklist to verify the plugin is active.

Why LLMs Appear to Ignore Guidelines

According to the source code analysis, LLMs violate the Karpathy guidelines for three predictable reasons:

  • Prompt Engineering Gap – The model never receives the guidelines because the developer didn't attach the plugin or include CLAUDE.md in the conversation context.
  • Ambiguous User Requests – The prompt omits critical constraints (e.g., privacy requirements), causing the model to default to assumptions and violate Think Before Coding.
  • Missing Verifiable Success Criteria – Without a concrete test or specification, the model cannot apply Goal-Driven Execution, resulting in broad, unfocused changes.

Step-by-Step Debugging Procedure

Follow this sequence to diagnose and resolve guideline violations:

  1. Confirm the Guidelines Are Loaded
    Check the plugin list in Claude Code or verify that CLAUDE.md is present in your working directory. If missing, install it with:

    curl -o CLAUDE.md https://raw.githubusercontent.com/forrestchang/andrej-karpathy-skills/main/CLAUDE.md
  2. Capture the Raw LLM Output
    Save the assistant's response verbatim. Audit it against the four principles:

    • Assumptions listed? → Think Before Coding satisfied.
    • Extraneous abstractions? → Simplicity First violation.
    • Unrelated code sections changed? → Surgical Changes violation.
    • No test or success condition? → Goal-Driven Execution violation.
  3. Map Violations to Source Examples
    Use EXAMPLES.md as a reference library. Match symptoms to known patterns: hidden assumptions map to Example 1 under "Think Before Coding," over-engineered calculators map to "Simplicity First," and drive-by refactoring maps to "Surgical Changes."

  4. Ask Clarifying Questions
    Prompt the LLM with a focused follow-up that explicitly references the violated principle. For hidden assumptions, ask:

    "Before implementing the export feature, could you list the assumptions you're making about scope, format, fields, and volume? (Think Before Coding)"

  5. Define Verifiable Success Criteria
    Convert vague requests into concrete tests. When asked to "make the search faster," specify:

    "Do you want to reduce average latency below 100 ms, increase throughput, or improve perceived speed? (Goal-Driven Execution)"
    Then require the model to write a test before implementing the change.

  6. Validate the Revised Output
    Run any generated tests or perform manual checks. Ensure only the lines required for the fix changed—no extra imports, style changes, or added features.

  7. Iterate with Explicit References
    If the model still deviates, repeat steps 3-6, tightening the prompt with explicit principle references. Persistent failures indicate the plugin may not be attached correctly or the model temperature is too high.

Practical Debugging Code Samples

Automate the detection of guideline violations using these helpers from the repository's debugging patterns.

Detecting Unscoped Assumptions

Use this Python scanner to identify Think Before Coding breaches in the LLM's raw output:

import re
from typing import List

def detect_assumptions(output: str) -> List[str]:
    """
    Scan LLM output for common assumption patterns.
    Returns a list of suspect statements.
    """
    patterns = [
        r"export all users",
        r"write to '?.*\.json'?",
        r"default to .*",
    ]
    return [m for pat in patterns for m in re.findall(pat, output, flags=re.I)]

If this function returns any matches, the model has made hidden assumptions without declaring them.

Enforcing Surgical Changes

Verify that the model adhered to the Surgical Changes principle by checking that only target files were modified:


# After the assistant's diff, run a git diff check:

git diff --diff-filter=AM --unified=0 HEAD | grep -E '^\+{3}|^\-{3}'

Only lines modifying the intended file should appear; any unrelated file modifications indicate a violation.

Goal-Driven Test Wrapper

Force Goal-Driven Execution by wrapping the success criteria in a timeout-based validator:

import time
from typing import Callable

def assert_success(criteria: Callable[[], bool], timeout: int = 30):
    """
    Repeatedly evaluate a success function until it returns True or timeout.
    Encourages LLMs to produce testable outcomes.
    """
    start = time.time()
    while time.time() - start < timeout:
        if criteria():
            return True
        time.sleep(1)
    raise AssertionError("Success criteria not met within timeout")

Ask the LLM to implement the criteria lambda before writing any implementation code.

Summary

  • Verify loading first – Ensure CLAUDE.md is present in the working directory or the plugin is attached before assuming malfeasance.
  • Reference principles explicitly – Cite Think Before Coding, Simplicity First, Surgical Changes, or Goal-Driven Execution by name in follow-up prompts to correct drift.
  • Map to EXAMPLES.md – Match the LLM's error patterns against the curated before/after snippets to diagnose specific violations.
  • Automate detection – Use the detect_assumptions scanner and git diff checks to catch violations programmatically.
  • Demand tests first – Enforce Goal-Driven Execution by requiring verifiable success criteria and test code before implementation.

Frequently Asked Questions

How do I verify that the Karpathy guidelines are actually loaded in my session?

Check the Claude Code plugin panel to confirm the skill is active, or manually verify that the CLAUDE.md file exists in your repository root and is included in the conversation context. If the file is absent, the model has no way to know the guidelines exist.

What is the difference between SKILL.md and CLAUDE.md in the repository?

CLAUDE.md is the human-readable policy text that the LLM consumes directly, while skills/karpathy-guidelines/SKILL.md is a machine-readable YAML manifest that allows automated toolchains to discover and load the skill. For manual debugging, CLAUDE.md is the critical file to verify.

Why does the LLM keep over-engineering solutions even when I mention Simplicity First?

Over-engineering typically occurs when the prompt implies a complex architecture or when the model defaults to pattern completion. Explicitly constrain the solution space by asking the model to list its assumptions first (Think Before Coding) and setting hard limits on dependencies or abstraction layers.

How can I prevent an LLM from changing unrelated code during a fix?

Enforce the Surgical Changes principle by asking the model to identify the exact lines it plans to modify before generating code, then use the git diff command provided above to verify that only the specified files and line ranges appear in the output.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →