# Program-of-Thought (PoT) vs Chain-of-Thought (CoT): Key Differences Explained

> Discover the key differences between Program-of-Thought PoT and Chain-of-Thought CoT prompting techniques. Learn how PoT uses code generation and CoT uses natural language for LLM problem solving.

- Repository: [Tongxin Yuan/dive-into-llms](https://github.com/Lordog/dive-into-llms)
- Tags: deep-dive
- Published: 2026-04-16

---

**Program-of-Thought (PoT) is a prompting technique where large language models generate executable code to solve problems, whereas Chain-of-Thought (CoT) uses step-by-step natural language reasoning to reach conclusions.**

Both techniques improve multi-step reasoning in large language models, but they differ fundamentally in execution methodology. According to the `Lordog/dive-into-llms` repository—specifically in [`documents/chapter2/README.md`](https://github.com/Lordog/dive-into-llms/blob/main/documents/chapter2/README.md)—PoT moves computation out of the model and into an external interpreter, while CoT relies entirely on the model's internal text generation capabilities. Understanding these distinctions helps practitioners select the right approach for mathematical, logical, or algorithmic tasks.

## What Is Chain-of-Thought (CoT)?

Chain-of-Thought prompting instructs models to produce a **natural-language reasoning trace** before delivering the final answer. The model essentially "thinks out loud," breaking complex problems into textual intermediate steps. As documented in the repository's Chapter 2 README, CoT improves performance on multi-step reasoning by encouraging logical inference through prose explanations.

In [`documents/chapter2/README.md`](https://github.com/Lordog/dive-into-llms/blob/main/documents/chapter2/README.md), the CoT approach appears in a comparison table illustrating how models handle word problems through English explanations. This method excels when logical consistency and explanatory transparency matter more than computational precision, though it may produce approximate or inconsistent numeric results on complex calculations.

## What Is Program-of-Thought (PoT)?

Program-of-Thought extends the CoT concept by replacing natural-language reasoning with **programmatic reasoning traces**. Instead of writing explanations in prose, the model generates executable code—typically Python—that solves the problem algorithmically. The final answer is obtained by running the generated code in an external interpreter.

As shown in [`documents/chapter2/README.md`](https://github.com/Lordog/dive-into-llms/blob/main/documents/chapter2/README.md), PoT prompts often ask the model to implement a `solver()` function. For example, the repository demonstrates a cake-baking timing problem where the model generates Python code to calculate start times based on preparation steps. This approach leverages the model's ability to write syntactically correct programs while offloading actual computation to a Python runtime.

## Key Differences Between PoT and CoT

The architectural distinction between these methods centers on where computation occurs and how reasoning is represented.

### Output Format

**CoT** produces plain English reasoning steps that read like a human explanation. **PoT** generates valid source code, such as Python scripts, that can be parsed and executed by an interpreter.

### Computation Method

In **CoT**, computation remains implicit and occurs within the language model's internal reasoning process. In **PoT**, computation is explicit—performed by an external interpreter when the generated code executes. This separation ensures precise arithmetic and algorithmic logic.

### Execution Requirements

**CoT** requires only a language model capable of text generation. **PoT** requires a secure runtime environment capable of executing the generated code, along with error handling for syntax mistakes.

## Practical Code Examples

The `dive-into-llms` repository provides concrete implementations demonstrating both approaches.

### Chain-of-Thought Example

Consider a word problem about players quitting a game. A CoT prompt elicits the following reasoning:

```text
Q: There were 10 friends playing a video game online when 7 players quit.  
If each player left had 8 lives, how many lives did they have total?

A: Let's think step by step.  
- First, compute the total lives at the start: 10 friends × 8 lives = 80 lives.  
- Then, 7 players quit, removing 7 × 8 = 56 lives.  
- Subtract the lost lives: 80 − 56 = 24 lives remaining.  
Therefore, the answer is 24.

```

This example demonstrates how CoT breaks arithmetic into verbal steps, though the model performs the calculation internally.

### Program-of-Thought Example

For temporal calculation problems, PoT generates executable Python. The repository's [`documents/chapter2/README.md`](https://github.com/Lordog/dive-into-llms/blob/main/documents/chapter2/README.md) illustrates this with a cake-baking scenario:

```text
Q: Jordan wants to bake a cake that will be ready by 5:00 pm.  
The recipe requires 20 min for batter, 30 min for baking, 2 h to cool, and 10 min to frost.  
When is the latest time Jordan can start the cake?

A: Write a Python function `solver()` that computes the start time.

```

The model generates:

```python
def solver():
    # minutes for each step

    minutes_to_make_batter = 20
    minutes_to_bake_cake = 30
    minutes_to_cool_cake = 2 * 60   # 2 hours → minutes

    minutes_to_frost_cake = 10

    # total minutes needed

    total_minutes = (minutes_to_make_batter + minutes_to_bake_cake +
                     minutes_to_cool_cake + minutes_to_frost_cake)

    # target serving time in hours (5 pm)

    target_hour = 17  # 24-hour clock

    # compute start hour

    start_hour = target_hour - total_minutes / 60
    return start_hour

print(solver())   # → 13.5 (i.e., 1:30 pm)

```

Running this code yields `13.5`, indicating a 1:30 pm start time. The repository also includes a minimal implementation in [`documents/chapter2/README.md`](https://github.com/Lordog/dive-into-llms/blob/main/documents/chapter2/README.md):

```python
minutes_to_make_batter = 20
minutes_to_bake_cake = 30
minutes_to_cool_cake = 2 * 60
minutes_to_frost_cake = 10
total_minutes = minutes_to_make_batter + minutes_to_bake_cake + minutes_to_cool_cake + minutes_to_frost_cake
total_hours = total_minutes / 60
ans = 5 - total_hours

```

Here, the variable `ans` holds the computed start time, demonstrating how PoT offloads arithmetic to the Python interpreter.

## Implementation Files

The `Lordog/dive-into-llms` repository contains primary source materials for both techniques:

- **[`documents/chapter2/README.md`](https://github.com/Lordog/dive-into-llms/blob/main/documents/chapter2/README.md)**: Contains the CoT vs PoT comparison table and concrete prompt examples, including the cake-baking Python implementation.
- **`documents/chapter2/dive-prompting.ipynb`**: A Jupyter notebook demonstrating PoT prompting with runnable code cells that execute generated programs, allowing direct experimentation with the technique.

## When to Use CoT vs PoT

Select **Chain-of-Thought** when tasks require explanatory reasoning, commonsense inference, or when no code execution environment is available. Use **Program-of-Thought** when problems involve precise calculations, date-time arithmetic, or algorithmic steps that benefit from external verification and exact computation.

## Summary

- **Program-of-Thought (PoT)** generates executable code rather than natural language, moving computation to an external interpreter for precise results.
- **Chain-of-Thought (CoT)** relies on step-by-step English reasoning processed entirely within the language model, suitable for explanatory tasks.
- PoT achieves higher accuracy on mathematical and algorithmic tasks by eliminating arithmetic errors through external code execution.
- CoT provides greater flexibility for tasks where textual explanations are valuable and no runtime environment exists.
- The `dive-into-llms` repository implements both in [`documents/chapter2/README.md`](https://github.com/Lordog/dive-into-llms/blob/main/documents/chapter2/README.md) and `documents/chapter2/dive-prompting.ipynb`, providing working Python examples like the `solver()` function.

## Frequently Asked Questions

### Can Program-of-Thought and Chain-of-Thought be combined?

Yes, hybrid approaches exist where models generate both natural language explanations and executable code. However, according to the `dive-into-llms` documentation, these are typically treated as distinct strategies because PoT specifically offloads computation to an interpreter while CoT keeps reasoning internal to the model.

### What programming languages work best for PoT?

Python is the dominant language for Program-of-Thought prompting because large language models are trained extensively on Python syntax. The examples in [`documents/chapter2/README.md`](https://github.com/Lordog/dive-into-llms/blob/main/documents/chapter2/README.md) use Python exclusively, leveraging its readability and mathematical precision.

### Why does PoT improve accuracy on math problems?

PoT eliminates arithmetic errors by delegating calculations to a deterministic interpreter rather than relying on the language model's internal weights. When the model generates code like `total_minutes / 60` in the repository's cake example, the Python interpreter computes the exact floating-point result, avoiding approximation errors common in pure text generation.

### Is PoT more expensive than CoT?

PoT incurs additional overhead because it requires a secure sandbox environment to execute generated code safely. While CoT only requires text generation inference, PoT needs both code generation and execution infrastructure, potentially increasing latency and computational costs depending on the runtime environment setup.