Implementing the Agent Loop in AI Systems: A Complete Guide to ReAct Architecture
The agent loop—also known as the ReAct loop—interleaves reasoning (thought), action (tool execution), and observation (tool results) in a cyclic control flow until a stop condition terminates execution, powering autonomous agents from Claude Code to LangGraph.
The agent loop forms the computational backbone of every modern autonomous AI system. In the rohitg00/ai-engineering-from-scratch repository, this pattern is implemented as a deterministic control structure that manages the conversation state between a language model and its available tools. Mastering this implementation provides a portable foundation for understanding any major agent framework, from the OpenAI Agents SDK to AutoGen v0.4.
The ReAct Pattern: Thought, Action, Observation
The ReAct (Reasoning + Acting) pattern, introduced by Yao et al. in 2022, structures agent behavior as an iterative sequence of three distinct phases. First, the Thought phase where the LLM reasons about the current state and plans the next step. Second, the Action phase where the model emits a tool call with specific arguments. Third, the Observation phase where the system executes the tool and returns results to the model. This triad repeats until the agent determines it has sufficient information to provide a final answer.
In phases/14-agent-engineering/01-the-agent-loop/code/main.py, the AgentLoop class orchestrates this cycle through its run method (lines 104-121). The loop maintains a history attribute (lines 99-103)—a chronologically ordered list of Turn objects—that serves as the message buffer, allowing the LLM to see its own reasoning trajectory across multiple iterations.
Core Architectural Components
Message Buffer and Turn Management
The message buffer stores the entire transcript of the agent's execution as a list of Turn dataclasses. Each turn captures the kind (user, thought, action, observation, or final), content, and optional tool call metadata. In the reference implementation, AgentLoop.history appends each phase of the ReAct cycle, ensuring the LLM receives complete context for every subsequent decision.
Tool Registry and Deterministic Dispatch
The Tool Registry guarantees deterministic dispatch from tool name to callable. Implemented as the ToolRegistry class (lines 34-44), it maintains an internal dictionary mapping string names to functions. The dispatch method (lines 45-53) handles execution and error normalization: if a tool is missing, it returns an error string; if execution raises an exception, it catches and formats the error message. This prevents tool failures from crashing the agent loop and allows the LLM to recover or adjust its strategy.
Stop Conditions and Safety Guardrails
The loop terminates when any stop condition evaluates to true. Valid termination triggers include: an explicit finish token emitted by the LLM, absence of tool calls indicating a direct answer, a turn budget exhaustion, or a guardrail violation. The AgentLoop.max_turns attribute (line 101) defaults to 12 iterations, providing a hard ceiling on computational cost. This budget mechanism prevents runaway loops when agents enter unproductive reasoning cycles.
Observation Formatting
Raw tool outputs often contain complex data structures, null values, or exceptions. The Observation Formatter normalizes these into string representations that the LLM can process. In the reference code, ToolRegistry.dispatch returns stringified results or error messages, ensuring the AgentLoop can always append a valid observation to the history buffer and continue execution.
The Control Flow Implementation
The agent loop follows this precise execution pattern, implemented in the AgentLoop.run method:
def run(self, user_message: str) -> str:
self.history.append(Turn(kind="user", content=user_message))
for _ in range(self.max_turns):
reply = self.llm.respond(self.history)
if reply["kind"] == "finish":
self.history.append(Turn(kind="final", content=reply["content"]))
return reply["content"]
# Record reasoning
self.history.append(Turn(kind="thought", content=reply.get("thought", "")))
# Execute action
call = ToolCall(name=reply["action"], args=reply.get("args", {}))
observation = self.tools.dispatch(call)
# Record result
self.history.append(
Turn(kind="action", content=call.name,
tool_call=call, observation=observation)
)
return "budget exhausted"
This structure remains consistent across frameworks. The LLM receives the accumulated history, emits either a final answer or a tool request, and the loop handles the execution plumbing. The ToyLLM class in the repository demonstrates this contract using a scripted policy for deterministic testing, but replacing it with a production LLM client (Claude, OpenAI, etc.) requires no changes to the loop logic.
Universal Pattern Across Frameworks
The 2025–2026 evolution of agent APIs introduced native reasoning channels—replacing explicit "Thought:" tokens with structured reasoning fields—yet the underlying control flow remains identical. Whether using Claude's Agent SDK, OpenAI's Agents SDK, LangGraph's state graphs, or AutoGen v0.4's actor model, all implementations wrap this same Thought-Action-Observation cycle with additional scaffolding for state checkpointing, distributed messaging, or role-based templates.
The repository's minimal implementation in main.py strips away framework-specific abstractions to reveal the portable core. Understanding this pure-Python reference enables engineers to debug complex agent behaviors in production systems by tracing them back to these fundamental components: message buffering, tool dispatch, and iteration control.
Complete Working Example
Below is the minimal ReAct implementation from the repository, runnable without external dependencies:
from dataclasses import dataclass, field
from typing import Any, Callable
@dataclass
class Turn:
kind: str
content: str
tool_call: "ToolCall" | None = None
observation: str | None = None
@dataclass
class ToolCall:
name: str
args: dict[str, Any]
class ToolRegistry:
def __init__(self) -> None:
self._tools: dict[str, Callable[..., str]] = {}
def register(self, name: str, fn: Callable[..., str]) -> None:
self._tools[name] = fn
def dispatch(self, call: ToolCall) -> str:
fn = self._tools.get(call.name)
if fn is None:
return f"error: unknown tool {call.name!r}"
try:
return fn(**call.args)
except Exception as e:
return f"error: {type(e).__name__}: {e}"
@dataclass
class AgentLoop:
llm: Any # ToyLLM or production client
tools: ToolRegistry
max_turns: int = 12
history: list[Turn] = field(default_factory=list)
def run(self, user_message: str) -> str:
self.history.append(Turn(kind="user", content=user_message))
for _ in range(self.max_turns):
reply = self.llm.respond(self.history)
if reply["kind"] == "finish":
self.history.append(Turn(kind="final", content=reply["content"]))
return reply["content"]
self.history.append(Turn(kind="thought", content=reply.get("thought", "")))
call = ToolCall(name=reply["action"], args=reply.get("args", {}))
observation = self.tools.dispatch(call)
self.history.append(
Turn(kind="action", content=call.name,
tool_call=call, observation=observation)
)
return "budget exhausted"
This code mirrors the reference implementation in phases/14-agent-engineering/01-the-agent-loop/code/main.py, providing a drop-in template for custom agent development.
Summary
- The ReAct loop structures agent execution as alternating phases of reasoning, tool action, and observation observation, implemented in the
AgentLoopclass. - Message history (stored in
AgentLoop.history) maintains chronological context across turns, essential for multi-step problem solving. - ToolRegistry provides deterministic dispatch and error normalization through its
registeranddispatchmethods, preventing tool failures from terminating execution. - Stop conditions include explicit finish tokens, missing tool calls, and the
max_turnsbudget (defaulting to 12), which prevents infinite loops. - The pattern is universal across modern frameworks—understanding the raw implementation in
rohitg00/ai-engineering-from-scratchprovides debugging capabilities when working with higher-level abstractions like LangGraph or AutoGen.
Frequently Asked Questions
What is the difference between the agent loop and the ReAct pattern?
The ReAct pattern (Reasoning + Acting) refers to the cognitive structure of interleaving thought and action, while the agent loop is the programmatic implementation that executes this pattern cyclically. The ReAct pattern defines what the agent does (think, act, observe), whereas the agent loop defines how the computer manages the iteration, state persistence, and termination conditions.
How does the turn budget prevent runaway agents?
The max_turns parameter in AgentLoop (defaulting to 12 iterations) acts as a circuit breaker. If an agent enters a recursive reasoning pattern or fails to converge on an answer, the hard limit forces termination after the specified number of cycles. This prevents excessive API costs and computational resource consumption when agents encounter edge cases or ambiguous instructions.
Can this implementation work with Claude, GPT-4, or other production LLMs?
Yes. The AgentLoop class is agnostic to the LLM implementation. The ToyLLM class in the repository demonstrates the required interface—accepting a history of turns and returning a dictionary with kind, thought, action, and args keys. Replacing ToyLLM with a client that calls Anthropic's Claude, OpenAI's GPT-4, or any other model requires only implementing the respond method with the appropriate API client, while the loop logic remains unchanged.
What happens when a tool throws an exception?
The ToolRegistry.dispatch method catches all exceptions during tool execution and returns a formatted error string rather than propagating the exception. This error string becomes the observation appended to the history buffer. The LLM then receives this error message in the next iteration, allowing it to adjust its approach, request different parameters, or acknowledge the failure—maintaining the flow of the agent loop without crashes.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →