Impact of Incorrect Examples on Few-Shot Prompting: Risks and Mitigation Strategies

Incorrect examples in few-shot prompting cause pattern mis-learning, confidence degradation, and cascading errors that can flip a model's output distribution from correct to erroneous behavior.

Few-shot prompting relies on providing language models with example input-output pairs that demonstrate the desired task pattern. When these examples contain errors, the model treats the corrupted signal as ground truth, leading to systematic failures in generation quality. This analysis draws from the Lordog/dive-into-llms repository to examine exactly how incorrect examples impact few-shot prompting performance and how to prevent these failures.

How Few-Shot Prompting Interprets Examples

In few-shot prompting, the model receives a handful of example pairs that function as a micro-training set. The model analyzes these shots to infer the underlying pattern connecting inputs to outputs. According to the Lordog/dive-into-llms source code, this inference process is highly sensitive to the quality of the provided examples—any corruption in the training signal directly pollutes the inferred pattern.

When the model processes these examples, it calculates statistical relationships between the prompt structure and response format. If an example contains an incorrect answer or ambiguous formatting, the model incorporates this noise into its generation parameters, treating the error as a valid data point.

Three Critical Failure Modes from Incorrect Examples

Pattern Mis-learning

The most immediate impact of incorrect examples is pattern mis-learning. When a language model encounters an example where the input maps to a wrong output, it learns the wrong relationship between the query and response. Instead of following the intended behavior, the generated answer mirrors the error present in the examples. This occurs because the model weights the few provided examples heavily when the task is ambiguous or complex.

Confidence Degradation

Incorrect examples introduce noise into the training signal, causing confidence degradation in the model's internal probability distributions. The model's confidence scores become less reliable when the few-shot context contains contradictions or errors. This uncertainty increases the likelihood of hallucinations, contradictory outputs, or responses that vary widely between semantically identical prompts.

Cascading Errors in Chain-of-Thought

In chain-of-thought or multi-step few-shot settings, incorrect examples create cascading errors. An early wrong example can steer subsequent reasoning steps off-track, compounding the initial mistake through the entire generation process. As noted in the repository's documentation for documents/chapter2/assets/understanding-CoT.png, the sequential nature of reasoning chains makes them particularly vulnerable to error propagation from malformed examples.

Repository Evidence on Incorrect Example Impact

The Lordog/dive-into-llms repository explicitly documents this phenomenon in documents/chapter2/README.md at lines 111-113:

"错误范例的影响:把少样本学习中的例子改成错误的答案,结果会发生变化吗?"

This translates to examining whether changing few-shot examples to incorrect answers alters the results—confirming that a single malformed example can flip the model's output distribution. The repository demonstrates that few-shot learning does not automatically filter out incorrect demonstrations; rather, it treats them as authoritative ground truth for the task at hand.

The interactive notebook documents/chapter2/dive-prompting.ipynb provides practical implementations showing how corrupted examples propagate through the generation process, turning correctly behaving prompts into failure cases.

Best Practices to Mitigate Incorrect Example Risks

To prevent the degradation of few-shot prompting performance, implement these verification strategies:

  • Validate every example manually and programmatically. Ensure each example contains the correct answer and consistent formatting before including it in the prompt. This guarantees the model receives a clean, authoritative signal rather than noise.

  • Keep examples minimal and focused. Limit the number of shots to only what is necessary for the task. Fewer examples reduce the probability of hidden errors and make the underlying pattern easier for the model to extract correctly.

  • Implement sanity-check scripts. Create automated tests that verify each example's output matches expected results. For instance, programmatically confirm that translations produce valid target language words or that calculations yield mathematically correct results.

  • Use explicitly labeled negative examples when needed. If demonstrating what not to do, clearly label these as incorrect (e.g., "Incorrect answer: ..."). This signals to the model that the particular output is not the target, preventing accidental adoption of wrong patterns.

Implementing Safe Few-Shot Prompts with Verification

The following Python implementation from the Lordog/dive-into-llms repository demonstrates how to construct a safe few-shot prompt using the dashscope API. This approach first verifies each example before concatenating them into the final prompt.

import dashscope

# --- Step 1: Define and verify examples -------------------------------------------------

examples = [
    {
        "prompt": "Translate English to French: sea otter =>",
        "expected": "loutre de mer"
    },
    {
        "prompt": "Translate English to French: cheese =>",
        "expected": "fromage"
    },
]

def verify_examples(examples):
    """Raise an error if any example output is empty or malformed."""
    for i, ex in enumerate(examples):
        if not ex["expected"].strip():
            raise ValueError(f"Example {i} has an empty expected output")
        # add more domain‑specific checks here

    return True

verify_examples(examples)

# --- Step 2: Build the few‑shot prompt --------------------------------------------------

def build_few_shot_prompt(examples, query):
    prompt = ""
    for ex in examples:
        prompt += f"{ex['prompt']} {ex['expected']}\n"
    prompt += f"{query}"
    return prompt

user_query = "Translate English to French: plush giraffe =>"
prompt_text = build_few_shot_prompt(examples, user_query)

# --- Step 3: Call the model ------------------------------------------------------------

resp = dashscope.Generation.call(
    model="qwen-turbo",
    prompt=prompt_text,
    # optional safety settings

    top_p=0.9,
    temperature=0.0,
)

print("Model response:", resp.output)

This implementation enforces three critical safety layers:

  1. Pre-flight verification through verify_examples() ensures every example contains non-empty expected output before reaching the model.

  2. Deterministic generation using temperature=0.0 isolates the impact of the examples themselves, removing randomness that might mask the effect of incorrect shots.

  3. Consistent formatting preservation in build_few_shot_prompt() maintains the exact structure the model expects, preventing format-based misinterpretation.

If you introduce an incorrect example (e.g., changing "loutre de mer" to "loutre du lac"), the model will begin producing the wrong translation for similar inputs, empirically demonstrating how incorrect examples corrupt few-shot prompting.

Summary

  • Incorrect examples act as corrupted training signals in few-shot prompting, causing models to learn wrong patterns rather than the intended task behavior.
  • Three major failure modes emerge: pattern mis-learning, confidence degradation, and cascading errors in multi-step reasoning chains.
  • The Lordog/dive-into-llms repository documents this risk explicitly in documents/chapter2/README.md, noting that changing examples to wrong answers fundamentally alters model outputs.
  • Verification pipelines are essential for production few-shot prompts, including automated sanity checks and minimal example sets to reduce error surfaces.
  • Explicit labeling of negative examples can mitigate risks when demonstrating incorrect approaches is necessary for the task.

Frequently Asked Questions

How many incorrect examples does it take to degrade few-shot prompting performance?

Even a single incorrect example can significantly alter a model's output distribution, particularly in tasks with clear right-or-wrong answers. The impact scales with the proportion of incorrect examples relative to the total shot count—if you provide three shots and one is wrong, the model has a 33% chance of learning the corrupted pattern. For critical applications, maintain 100% accuracy in your few-shot examples or explicitly label any negative demonstrations.

Can incorrect examples ever be useful in few-shot prompting?

Incorrect examples can be valuable only when explicitly labeled as negative demonstrations. By clearly marking an output as "Incorrect answer:" or "Wrong approach:", you signal to the model that this pattern should be avoided. This technique requires careful formatting to prevent the model from accidentally learning the wrong pattern as correct. Without explicit negative labeling, incorrect examples consistently degrade performance.

How do I detect incorrect examples in my prompt dataset?

Implement programmatic verification scripts that test each example's output against validation rules or reference data. For translation tasks, verify that outputs match dictionary definitions. For mathematical reasoning, confirm calculations evaluate correctly. The verify_examples() function demonstrated in documents/chapter2/dive-prompting.ipynb provides a template for catching empty or malformed outputs before they reach the model.

Does temperature affect how models handle incorrect examples?

Temperature settings do not prevent the impact of incorrect examples—they only control the randomness of the generation process. Setting temperature=0.0 makes the model's reliance on the corrupted examples more visible by removing stochastic variation, but it does not mitigate the underlying pattern mis-learning. High temperatures might occasionally produce correct answers despite incorrect examples, but this is unreliable and does not solve the root problem of pattern corruption.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →