How to Implement Chain-of-Thought (CoT) Reasoning in LLMs: A Complete Guide

Chain-of-Thought prompting induces large language models to generate step-by-step rationales before final answers by augmenting prompts with reasoning cues like "Let's think step by step" or few-shot exemplars.

Chain-of-Thought (CoT) reasoning is a prompting technique that nudges an LLM to decompose complex problems into intermediate reasoning steps. According to the Lordog/dive-into-llms repository, this approach significantly improves accuracy on arithmetic, logical, and multi-hop tasks by leveraging the model's ability to attend to previously generated reasoning tokens.

Architectural Overview of CoT in LLMs

Implementing CoT reasoning involves understanding how the prompt augmentation flows through the transformer architecture:

  • Input Layer: The user prompt receives a CoT cue such as "Let's think step by step." or a detailed few-shot template, biasing the model toward generating reasoning tokens.

  • Tokenizer & Embedding: The extra tokens from the CoT prompt are converted to embeddings that guide the model away from jumping directly to conclusions.

  • Transformer Decoder: Through autoregressive prediction and self-attention mechanisms, the model attends to previously generated reasoning steps, building upon intermediate conclusions iteratively.

  • Output Layer: After emitting the reasoning trace, the model produces the final answer token(s), which can be post-processed to extract numeric values or final statements.

This architecture enables intermediate supervision—implicitly teaching the model to treat reasoning as a sub-task aligned with its training on chain-like text data.

Implementation Methods

Zero-Shot CoT

The simplest implementation appends a standard reasoning cue to the query. As shown in documents/chapter2/README.md, this requires no task-specific examples.

import openai

def cot_zero_shot(question: str) -> str:
    prompt = f"Question: {question}\nAnswer: Let's think step by step."
    response = openai.ChatCompletion.create(
        model="gpt-3.5-turbo",
        messages=[{"role": "user", "content": prompt}],
        temperature=0.7,  # enable diverse reasoning paths

        max_tokens=500,
    )
    return response.choices[0].message.content.strip()

print(cot_zero_shot("There were 12 apples. I gave 5 to a friend and then bought 3 more. How many apples do I have now?"))

Few-Shot CoT

For complex tasks, provide exemplars that demonstrate the reasoning pattern. The repository's Chapter 2 ("少样本" section) recommends including step-by-step solutions within the prompt context.

EXEMPLARS = """
Q: There are 15 trees in the grove. After planting, there are 21 trees. How many were planted?
A: There were 15 trees originally. After planting there are 21. So 21 - 15 = 6. The answer is 6.

Q: If 3 cars are in a lot and 2 more arrive, how many cars are there?
A: There were 3 cars. 2 more arrived. So 3 + 2 = 5. The answer is 5.
"""

def cot_few_shot(question: str) -> str:
    prompt = f"{EXEMPLARS}\nQ: {question}\nA:"
    response = openai.ChatCompletion.create(
        model="gpt-3.5-turbo",
        messages=[{"role": "user", "content": prompt}],
        temperature=0.0,  # deterministic once the pattern is learned

        max_tokens=300,
    )
    return response.choices[0].message.content.strip()

print(cot_few_shot("A baker uses 20 minutes for batter, 30 minutes for baking, 2 hours to cool, and 10 minutes to frost. When must they start to serve at 5 pm?"))

Self-Consistency Voting

To reduce noise and improve reliability, sample multiple reasoning paths and select the most frequent answer. This technique, documented in the "自洽性提升推理结果" section of the repository, requires temperature > 0 to generate diverse outputs.

import collections

def cot_self_consistency(question: str, n_samples: int = 5) -> str:
    answers = []
    for _ in range(n_samples):
        answer = cot_zero_shot(question)  # reuse zero-shot helper

        # Extract the final value assuming standard format

        ans = answer.split("The answer is")[-1].strip().strip(".")
        answers.append(ans)
    most_common = collections.Counter(answers).most_common(1)[0][0]
    return most_common

print(cot_self_consistency("There are 10 friends each with 8 lives. 7 quit. How many lives remain?"))

Programmatic CoT (PoT)

Programmatic CoT (PoT) generates executable code rather than natural language reasoning. The table at documents/chapter2/README.md#L93-L102 contrasts this approach with standard natural-language CoT, noting that PoT is particularly effective for precise numerical calculations.

def cot_programmatic(question: str) -> str:
    prompt = f"""
Question: {question}
Answer: Let's write a Python program step by step, then return the answer.

```python

# define variables

"""
    response = openai.ChatCompletion.create(
        model="gpt-3.5-turbo",
        messages=[{"role": "user", "content": prompt}],
        temperature=0.3,
        max_tokens=600,
    )
    code = response.choices[0].message.content
    # Execute the generated code in a safe sandbox (implementation omitted)

    return code

print(cot_programmatic("Compute the total minutes needed to bake a cake given the recipe steps."))

Source Code Reference

The Lordog/dive-into-llms repository provides comprehensive resources for implementing these techniques:

  • documents/chapter2/README.md: Contains the complete tutorial on prompting strategies, including zero-shot/few-shot CoT, self-consistency, Auto-CoT, Sum-CoT, and PoT implementations.
  • documents/chapter2/assets/understanding-CoT.png: Visual diagram of the CoT pipeline architecture.
  • documents/chapter2/assets/self-consistency.png: Illustration of the majority voting process used in self-consistency decoding.
  • documents/chapter2/README.md#L122-L124: Links to advanced variants including Auto-CoT, Sum-CoT, and CRITIC implementations for further customization.

Summary

  • Chain-of-Thought reasoning improves LLM accuracy by forcing step-by-step generation before final answers.
  • Zero-shot implementation requires only the phrase "Let's think step by step." added to the prompt.
  • Few-shot templates provide concrete reasoning patterns that help the model generalize to new queries.
  • Self-consistency mitigates sampling errors by aggregating results across multiple generation passes with temperature > 0.
  • Programmatic CoT (PoT) generates executable Python code for tasks requiring precise numerical computation rather than natural language reasoning.

Frequently Asked Questions

What is the difference between zero-shot and few-shot CoT?

Zero-shot CoT uses a generic reasoning cue like "Let's think step by step." without task examples, while few-shot CoT includes specific question-answer pairs with detailed reasoning steps in the prompt context. Few-shot approaches typically yield higher accuracy on complex tasks but require careful construction of exemplars.

How does self-consistency improve CoT performance?

Self-consistency works by sampling multiple reasoning paths (using temperature > 0) and selecting the most frequent final answer through majority voting. This reduces the impact of random noise in the decoding process and selects the most robust reasoning chain, as implemented in the cot_self_consistency function.

When should I use Programmatic CoT instead of natural language CoT?

Use Programmatic CoT (PoT) when the task requires precise numerical calculations, symbolic manipulation, or algorithmic execution that natural language might approximate incorrectly. PoT generates Python code that can be executed in a sandbox to derive exact answers, whereas natural language CoT is better for explanatory reasoning and commonsense tasks.

What temperature setting should I use for CoT prompting?

Use temperature = 0 for deterministic few-shot scenarios where you want consistent templated outputs, and temperature between 0.3 and 0.7 for zero-shot or self-consistency approaches that benefit from diverse reasoning paths. Higher temperatures enable exploration of alternative solution strategies when using self-consistency voting.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →