# How to Implement a ReAct Agent Loop from Scratch in Python

> Build a ReAct agent loop from scratch in Python. Understand the five core components and cycle through observation, reasoning, and action for task completion.

- Repository: [Rohit Ghumare/ai-engineering-from-scratch](https://github.com/rohitg00/ai-engineering-from-scratch)
- Tags: how-to-guide
- Published: 2026-07-25

---

**A ReAct agent loop repeatedly cycles through observation, reasoning, and action until a task is complete, using five core components: a message buffer, tool registry, stop condition, turn budget, and observation formatter.**

The ReAct (Reason + Act) pattern is the canonical architecture for autonomous language-model agents. This guide walks through the reference implementation in the `rohitg00/ai-engineering-from-scratch` repository, showing you how to build a framework-agnostic agent loop that works with any LLM provider.

## The Five Essential Components of a ReAct Loop

Every production-grade ReAct implementation requires five structural elements. In [`phases/14-agent-engineering/01-the-agent-loop/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/14-agent-engineering/01-the-agent-loop/code/main.py), these are encapsulated in the `AgentLoop` class and its supporting utilities.

### Message Buffer

The **message buffer** stores the complete turn-by-turn transcript, including user inputs, model thoughts, tool actions, and observations. In the reference code, this is implemented as `AgentLoop.history`, which maintains a list of `Turn` objects representing the entire conversation state.

### Tool Registry

The **tool registry** maps string tool names to executable callables. The `ToolRegistry` class in [`code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/code/main.py) provides a `dispatch` method that routes function calls from the LLM to the appropriate Python function, handling argument parsing and error catching.

### Stop Condition and Turn Budget

The loop terminates when the model emits a `finish` signal or when the **turn budget** is exhausted. The `AgentLoop.run` method checks `reply["kind"] == "finish"` after each LLM response and enforces `max_turns` (defaulting to 12) to prevent infinite loops.

### Observation Formatter

After a tool executes, its return value must be converted into a string observation that the LLM can consume on the next iteration. The reference implementation stores this in `Turn.observation`, populated by the string returned from `ToolRegistry.dispatch`.

## Step-by-Step Execution Flow

The ReAct loop follows a strict deterministic sequence implemented in the `AgentLoop.run` method (lines 99-121 of [`main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/main.py)):

1. **User turn** – Append the user message to the history buffer.
2. **LLM turn** – Call `ToyLLM.respond(history)` to receive either a thought+action pair or a finish signal.
3. **Thought recording** – Log the thought string as a `Turn` of kind `thought`.
4. **Tool dispatch** – Wrap the action name and arguments in a `ToolCall` and execute via `ToolRegistry.dispatch`.
5. **Observation recording** – Store the tool's return value (or error string) as the observation for this action turn.
6. **Loop back** – Feed the updated history to the LLM for the next iteration.
7. **Termination** – Return the final answer when the LLM emits `finish` or the turn budget is exceeded.

## Complete Python Implementation

The repository provides a minimal, deterministic implementation using `ToyLLM`, which emits scripted sequences of dictionaries containing `kind`, `thought`, `action`, and `args` keys.

```python
from phases.14_agent_engineering.01_the_agent_loop.code.main import (
    build_demo_agent, pretty_trace,
)

# Build the demo agent (toy LLM + three simple tools)

agent = build_demo_agent()

# Run the loop with a user query

final_answer = agent.run("What is 120 plus 15% tax, stored in kv?")

# Print the full trace

pretty_trace(agent.history)
print(f"\nfinal answer: {final_answer}")

```

Running this script produces a structured trace showing each reasoning step and tool invocation:

```

[00  user] What is 120 plus 15% tax, stored in kv?
[01 thought] store the base price
[02  action] kv_set({'key': 'base', 'value': '120'}) -> stored base
[03 thought] compute 15% tax
[04  action] calculator({'expr': '120 * 0.15'}) -> 18.0
[05 thought] store the tax
[06  action] kv_set({'key': 'tax', 'value': '18.0'}) -> stored tax
[07 thought] compute total
[08  action] calculator({'expr': '120 + 18.0'}) -> 138.0
[09 thought] confirm stored values
[10  action] kv_get({'key': 'base'}) -> 120
[11 final] the total including 15% tax is 138.0

```

The `build_demo_agent()` function wires together three default tools (`calculator`, `kv_get`, `kv_set`) with the toy LLM and the `AgentLoop` controller.

## Extending the Agent with Custom Tools

You can register additional tools without modifying the core loop logic. The `ToolRegistry` accepts any callable and exposes it to the LLM via the dispatch mechanism:

```python
def reverse(text: str) -> str:
    return text[::-1]

# Register the new tool

agent.tools.register("reverse", reverse)

# Extend the toy LLM's script with a new step

agent.llm.script.append(
    {"kind": "action", "thought": "reverse the word 'hello'",
     "action": "reverse", "args": {"text": "hello"}}
)

# Run again

final = agent.run("Run the reverse tool.")
pretty_trace(agent.history)

```

The loop automatically records the new `reverse` call and its observation in the history buffer, demonstrating how the architecture remains unchanged regardless of tool complexity.

## Integrating Production LLM Providers

To use a real model, replace `ToyLLM` with a client that implements the `respond(history)` interface. The control flow in `AgentLoop` remains identical:

```python
from openai import OpenAI

class RealLLM:
    def __init__(self, model="gpt-4o-mini"):
        self.client = OpenAI()
        self.model = model

    def respond(self, history):
        # Serialize history into a prompt (implementation omitted for brevity)

        # Call the Responses API and parse structured output

        # Must return a dict with keys: kind, thought, action, args, or finish

        ...

# Replace the toy LLM

real_llm = RealLLM()
agent = AgentLoop(llm=real_llm, tools=agent.tools, max_turns=20)

```

Only the `respond` method changes; the tool dispatch, observation handling, and budgeting logic stay exactly as implemented in the reference file.

## Summary

- The **ReAct loop** alternates between LLM reasoning and tool execution until a stop condition is met.
- **Five components** are required: message buffer (`AgentLoop.history`), tool registry (`ToolRegistry`), stop condition (`reply["kind"] == "finish"`), turn budget (`max_turns`), and observation formatter.
- The reference implementation in `rohitg00/ai-engineering-from-scratch` provides a **deterministic toy LLM** for testing and a pluggable architecture for production models.
- **Tool registration** is dynamic—add new capabilities by calling `tools.register()` without changing the loop logic.
- The code is **framework-agnostic**; the same `AgentLoop` structure powers Claude Agent SDK, LangGraph, and CrewAI under the hood.

## Frequently Asked Questions

### What is the ReAct pattern in AI agents?

The ReAct pattern combines reasoning traces (thoughts) with task-specific actions in an interleaved loop. The agent thinks about what to do, performs an action using a tool, observes the result, and repeats until the task is complete. This architecture appears in virtually every modern agent framework because it mirrors human cognitive workflows.

### How does the stop condition work in a ReAct loop?

The loop checks two termination criteria after each LLM response. First, it inspects `reply["kind"]` for the string `"finish"`, which signals the model has generated a final answer. Second, it compares the current turn count against `max_turns` (default 12) to prevent infinite execution if the model fails to conclude.

### Can I use this implementation with OpenAI or Anthropic models?

Yes. The `AgentLoop` class is designed to be provider-agnostic. You only need to implement a class with a `respond(history)` method that returns a dictionary containing `kind`, `thought`, `action`, `args`, or `finish` keys. Replace the `ToyLLM` instance with your custom client that calls the OpenAI Responses API or Anthropic's Messages API.

### What prevents the agent from running forever?

The **turn budget** (`max_turns`) acts as a circuit breaker. In `AgentLoop.run`, the loop counter increments with each iteration, and if it exceeds the budget, the loop terminates regardless of the LLM's state. This safety mechanism ensures that malfunctioning tool calls or repetitive reasoning patterns cannot cause infinite loops.