Chain-of-Thought (CoT) Prompting Techniques: Few-Shot, Zero-Shot, and Self-Consistency Explained
Chain-of-Thought (CoT) prompting is a method that encourages large language models (LLMs) to reason step-by-step before producing a final answer, significantly improving accuracy on complex arithmetic, commonsense, and symbolic reasoning tasks.
Chain-of-Thought (CoT) prompting techniques have emerged as a critical strategy for unlocking advanced reasoning capabilities in large language models. According to the documentation in the dair-ai/Prompt-Engineering-Guide repository, these methods work by inserting intermediate reasoning steps into prompts, allowing models to decompose complex problems rather than jumping directly to conclusions. The technique is particularly effective when layered on top of existing prompting strategies like few-shot learning.
Core Chain-of-Thought (CoT) Methods
The guides/prompts-advanced-usage.md file defines three primary variants of CoT prompting, each suited to different use cases based on whether you have example reasoning chains available.
Few-Shot CoT Prompting
Few-shot CoT combines exemplars that explicitly show the reasoning process with your target query. This method requires manually crafting examples where each demonstration includes the intermediate thought process leading to the final answer.
In guides/prompts-advanced-usage.md (lines 146-165), the guide demonstrates this with a mathematical reasoning task. The model is shown how to identify odd numbers, sum them, and determine if the result is even before answering the next query.
Zero-Shot CoT Prompting
Zero-shot CoT requires no example demonstrations. Instead, you append a simple trigger phrase—typically "Let's think step by step."—to the prompt. This single addition forces the model to generate its own reasoning chain without prior examples.
The implementation in guides/prompts-advanced-usage.md (lines 193-215) shows a zero-shot example where adding this trigger phrase enables the model to correctly solve a multi-step apple-counting word problem that it would otherwise fail.
Self-Consistency Sampling
Self-consistency runs the few-shot CoT prompt multiple times, sampling diverse reasoning paths, and selects the answer that appears most frequently across samples. This mitigates "greedy decoding" errors where a single reasoning chain might lead to an incorrect conclusion despite correct intermediate steps.
As documented in guides/prompts-advanced-usage.md (lines 228-277), this ensemble approach generates multiple completions for the same question (such as the tree-planting arithmetic problem), then aggregates results by majority vote.
Implementation Examples
Below are runnable prompt templates derived directly from the Prompt Engineering Guide source files. You can paste these into any LLM playground (OpenAI, Claude, or local models).
Few-Shot CoT Example
The odd numbers in this group add up to an even number: 4, 8, 9, 15, 12, 2, 1.
A: Adding all the odd numbers (9, 15, 1) gives 25. The answer is False.
The odd numbers in this group add up to an even number: 17, 10, 19, 4, 8, 12, 24.
A: Adding all the odd numbers (17, 19) gives 36. The answer is True.
The odd numbers in this group add up to an even number: 15, 32, 5, 13, 82, 7, 1.
A:
Expected model completion:
Adding all the odd numbers (15, 5, 13, 7, 1) gives 41. The answer is False.
Zero-Shot CoT Example
I went to the market and bought 10 apples. I gave 2 apples to the neighbor and 2 to the repairman. I then went and bought 5 more apples and ate 1. How many apples did I remain with?
Let's think step by step.
Expected model completion:
First, you started with 10 apples.
You gave away 2 apples to the neighbor and 2 to the repairman, so you had 6 apples left.
Then you bought 5 more apples, so now you had 11 apples.
Finally you ate 1 apple, so you would remain with 10 apples.
Self-Consistency Setup
Q: There are 15 trees in the grove. Grove workers will plant trees today. After they are done, there will be 21 trees. How many trees did they plant today?
A: We start with 15 trees. Later we have 21 trees. The difference must be the number of trees they planted.
So, they must have planted 21 - 15 = 6 trees. The answer is 6.
To implement self-consistency, repeat this prompt multiple times with temperature > 0 to generate diverse reasoning paths. Collect all final answers (e.g., 6, six, 6 trees) and select the most frequent response as the correct answer.
Key Considerations for CoT Implementation
CoT prompting exhibits emergent abilities that scale with model size. According to the source documentation in dair-ai/Prompt-Engineering-Guide, these techniques work best with sufficiently large models where multi-step reasoning capabilities become apparent. Smaller models may generate fluent but logically inconsistent thought chains.
The repository's main navigation in README.md (line 55) links to these advanced techniques, indicating their importance in the broader prompting hierarchy. For optimal results, combine CoT with clear instruction-following patterns from guides/prompts-basic-usage.md to establish the task context before requesting step-by-step reasoning.
Summary
- Chain-of-Thought (CoT) prompting forces LLMs to generate intermediate reasoning steps, reducing errors on complex tasks.
- Few-shot CoT uses 2-3 exemplars with explicit reasoning chains, documented in
guides/prompts-advanced-usage.mdlines 146-165. - Zero-shot CoT requires only the trigger phrase "Let's think step by step" with no examples, as shown in lines 193-215.
- Self-consistency improves reliability by sampling multiple reasoning paths and taking the majority vote, detailed in lines 228-277.
- These techniques require large-scale models to function effectively and can be combined with other prompting strategies for enhanced performance.
Frequently Asked Questions
What is the difference between few-shot and zero-shot Chain-of-Thought prompting?
Few-shot CoT requires manually crafted examples that include the reasoning process and final answer, allowing the model to pattern-match the thought structure. Zero-shot CoT uses only a trigger phrase like "Let's think step by step" without any examples, forcing the model to generate its own reasoning chain. Zero-shot is faster to implement but may be less consistent than few-shot for highly complex tasks.
When should I use Self-Consistency with CoT?
Use self-consistency when the cost of an incorrect answer is high or when you observe that single-generation CoT occasionally produces logical errors despite correct reasoning patterns. By sampling multiple completions (typically 5-40 paths) and selecting the majority answer, you trade computational cost for accuracy. This is particularly effective for arithmetic and symbolic reasoning benchmarks documented in the guides/prompts-advanced-usage.md self-consistency section.
Does Chain-of-Thought prompting work with all language models?
No. According to the dair-ai/Prompt-Engineering-Guide source, CoT prompting is an emergent capability that appears primarily in large language models. Smaller models (typically under 100B parameters) often fail to generate coherent multi-step reasoning chains or may produce "hallucinated" logic that doesn't actually solve the problem. The technique should be reserved for sufficiently capable models where reasoning patterns naturally develop.
How do I know if my CoT prompt is working correctly?
Verify that the model's generated text contains explicit intermediate steps that logically lead to the final answer, not just the final answer itself. In the few-shot examples from guides/prompts-advanced-usage.md, valid CoT outputs show the calculation process (e.g., "Adding all the odd numbers... gives...") before stating True/False. If the model skips directly to the answer, increase the temperature or refine your exemplars to emphasize the required reasoning format.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →