How to Implement a ReAct Agent Loop from Scratch in Python
A ReAct agent loop repeatedly cycles through observation, reasoning, and action until a task is complete, using five core components: a message buffer, tool registry, stop condition, turn budget, and observation formatter.
The ReAct (Reason + Act) pattern is the canonical architecture for autonomous language-model agents. This guide walks through the reference implementation in the rohitg00/ai-engineering-from-scratch repository, showing you how to build a framework-agnostic agent loop that works with any LLM provider.
The Five Essential Components of a ReAct Loop
Every production-grade ReAct implementation requires five structural elements. In phases/14-agent-engineering/01-the-agent-loop/code/main.py, these are encapsulated in the AgentLoop class and its supporting utilities.
Message Buffer
The message buffer stores the complete turn-by-turn transcript, including user inputs, model thoughts, tool actions, and observations. In the reference code, this is implemented as AgentLoop.history, which maintains a list of Turn objects representing the entire conversation state.
Tool Registry
The tool registry maps string tool names to executable callables. The ToolRegistry class in code/main.py provides a dispatch method that routes function calls from the LLM to the appropriate Python function, handling argument parsing and error catching.
Stop Condition and Turn Budget
The loop terminates when the model emits a finish signal or when the turn budget is exhausted. The AgentLoop.run method checks reply["kind"] == "finish" after each LLM response and enforces max_turns (defaulting to 12) to prevent infinite loops.
Observation Formatter
After a tool executes, its return value must be converted into a string observation that the LLM can consume on the next iteration. The reference implementation stores this in Turn.observation, populated by the string returned from ToolRegistry.dispatch.
Step-by-Step Execution Flow
The ReAct loop follows a strict deterministic sequence implemented in the AgentLoop.run method (lines 99-121 of main.py):
- User turn – Append the user message to the history buffer.
- LLM turn – Call
ToyLLM.respond(history)to receive either a thought+action pair or a finish signal. - Thought recording – Log the thought string as a
Turnof kindthought. - Tool dispatch – Wrap the action name and arguments in a
ToolCalland execute viaToolRegistry.dispatch. - Observation recording – Store the tool's return value (or error string) as the observation for this action turn.
- Loop back – Feed the updated history to the LLM for the next iteration.
- Termination – Return the final answer when the LLM emits
finishor the turn budget is exceeded.
Complete Python Implementation
The repository provides a minimal, deterministic implementation using ToyLLM, which emits scripted sequences of dictionaries containing kind, thought, action, and args keys.
from phases.14_agent_engineering.01_the_agent_loop.code.main import (
build_demo_agent, pretty_trace,
)
# Build the demo agent (toy LLM + three simple tools)
agent = build_demo_agent()
# Run the loop with a user query
final_answer = agent.run("What is 120 plus 15% tax, stored in kv?")
# Print the full trace
pretty_trace(agent.history)
print(f"\nfinal answer: {final_answer}")
Running this script produces a structured trace showing each reasoning step and tool invocation:
[00 user] What is 120 plus 15% tax, stored in kv?
[01 thought] store the base price
[02 action] kv_set({'key': 'base', 'value': '120'}) -> stored base
[03 thought] compute 15% tax
[04 action] calculator({'expr': '120 * 0.15'}) -> 18.0
[05 thought] store the tax
[06 action] kv_set({'key': 'tax', 'value': '18.0'}) -> stored tax
[07 thought] compute total
[08 action] calculator({'expr': '120 + 18.0'}) -> 138.0
[09 thought] confirm stored values
[10 action] kv_get({'key': 'base'}) -> 120
[11 final] the total including 15% tax is 138.0
The build_demo_agent() function wires together three default tools (calculator, kv_get, kv_set) with the toy LLM and the AgentLoop controller.
Extending the Agent with Custom Tools
You can register additional tools without modifying the core loop logic. The ToolRegistry accepts any callable and exposes it to the LLM via the dispatch mechanism:
def reverse(text: str) -> str:
return text[::-1]
# Register the new tool
agent.tools.register("reverse", reverse)
# Extend the toy LLM's script with a new step
agent.llm.script.append(
{"kind": "action", "thought": "reverse the word 'hello'",
"action": "reverse", "args": {"text": "hello"}}
)
# Run again
final = agent.run("Run the reverse tool.")
pretty_trace(agent.history)
The loop automatically records the new reverse call and its observation in the history buffer, demonstrating how the architecture remains unchanged regardless of tool complexity.
Integrating Production LLM Providers
To use a real model, replace ToyLLM with a client that implements the respond(history) interface. The control flow in AgentLoop remains identical:
from openai import OpenAI
class RealLLM:
def __init__(self, model="gpt-4o-mini"):
self.client = OpenAI()
self.model = model
def respond(self, history):
# Serialize history into a prompt (implementation omitted for brevity)
# Call the Responses API and parse structured output
# Must return a dict with keys: kind, thought, action, args, or finish
...
# Replace the toy LLM
real_llm = RealLLM()
agent = AgentLoop(llm=real_llm, tools=agent.tools, max_turns=20)
Only the respond method changes; the tool dispatch, observation handling, and budgeting logic stay exactly as implemented in the reference file.
Summary
- The ReAct loop alternates between LLM reasoning and tool execution until a stop condition is met.
- Five components are required: message buffer (
AgentLoop.history), tool registry (ToolRegistry), stop condition (reply["kind"] == "finish"), turn budget (max_turns), and observation formatter. - The reference implementation in
rohitg00/ai-engineering-from-scratchprovides a deterministic toy LLM for testing and a pluggable architecture for production models. - Tool registration is dynamic—add new capabilities by calling
tools.register()without changing the loop logic. - The code is framework-agnostic; the same
AgentLoopstructure powers Claude Agent SDK, LangGraph, and CrewAI under the hood.
Frequently Asked Questions
What is the ReAct pattern in AI agents?
The ReAct pattern combines reasoning traces (thoughts) with task-specific actions in an interleaved loop. The agent thinks about what to do, performs an action using a tool, observes the result, and repeats until the task is complete. This architecture appears in virtually every modern agent framework because it mirrors human cognitive workflows.
How does the stop condition work in a ReAct loop?
The loop checks two termination criteria after each LLM response. First, it inspects reply["kind"] for the string "finish", which signals the model has generated a final answer. Second, it compares the current turn count against max_turns (default 12) to prevent infinite execution if the model fails to conclude.
Can I use this implementation with OpenAI or Anthropic models?
Yes. The AgentLoop class is designed to be provider-agnostic. You only need to implement a class with a respond(history) method that returns a dictionary containing kind, thought, action, args, or finish keys. Replace the ToyLLM instance with your custom client that calls the OpenAI Responses API or Anthropic's Messages API.
What prevents the agent from running forever?
The turn budget (max_turns) acts as a circuit breaker. In AgentLoop.run, the loop counter increments with each iteration, and if it exceeds the budget, the loop terminates regardless of the LLM's state. This safety mechanism ensures that malfunctioning tool calls or repetitive reasoning patterns cannot cause infinite loops.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →