What Is the ReAct Loop? A Deep Dive into AI Agent Task Execution

The ReAct loop (Reason-Act-Observe) is an iterative orchestration pattern that enables large language models to solve complex tasks by alternating between internal reasoning, external tool calls, and observation processing until generating a final answer.

Modern AI agents require more than single-shot inference to handle real-world tasks effectively. The ReAct loop, extensively documented in the bojieli/ai-agent-book repository, provides a lightweight text-based protocol that allows LLMs to plan actions, execute tools, and adapt based on environmental feedback through a structured cycle of reasoning and acting.

How the ReAct Loop Works

The ReAct pattern follows a strict three-phase cycle that repeats until task completion. Each iteration appends new context to a running transcript, enabling the model to maintain state across multiple tool invocations.

The Reasoning Phase

During the Reasoning phase, the LLM generates a natural-language thought that describes the current plan or analysis. This internal monologue might state: "I need to search Wikipedia for the capital of France before answering."

According to the repository's implementation in chapter2/local_llm_serving/agent.py, this reasoning step is appended to the conversation history as an assistant message, creating a persistent chain of thought visible to subsequent iterations.

The Action Phase

The Act phase involves parsing the LLM's output for executable commands. The repository uses a specific text protocol where tool calls follow the syntax [[tool_name]] {"arg": "value"}.

In chapter2/prompt-engineering/tau_bench/agents/chat_react_agent.py, the ChatReActAgent class implements this parsing using a regular expression pattern—TOOL_PATTERN = r'\[\[([\w_]+)\]\]\s*(\{.*?\})'—to identify tool names and their JSON arguments. When a match is found, the agent invokes the corresponding function from the available toolkit.

The Observation Phase

After tool execution, the Observe phase feeds the raw result back into the prompt context as a structured observation. Whether the result is a successful API response, a calculation error, or empty data, this observation becomes part of the next iteration's input.

The implementation in chapter2/kv-cache/agent.py demonstrates how these observations are cached alongside previous thoughts to optimize subsequent LLM calls, while chapter2/attention_visualization/main.py provides ReActStep and ReActAttentionAgent classes for visualizing how attention weights shift across different phases of the loop.

Implementation Architecture

The ReAct loop's power lies in its simplicity—a text-driven protocol that works across any LLM backend supporting conversational interfaces.

Streaming Iteration Control

The core loop logic resides in chapter2/local_llm_serving/agent.py, where each iteration performs:

  1. Logging: Records the current step via logger.info(f"ReAct iteration {iteration}")
  2. Prompt Construction: Assembles the accumulated conversation including all prior thoughts, tool calls, and observations
  3. LLM Inference: Sends the compiled prompt to the model backend
  4. Pattern Matching: Searches for [[tool_name]] syntax to determine if another action is required
  5. Termination Check: Returns the final answer when no tool pattern is detected or when max_iterations is reached

This streaming approach allows the agent to handle arbitrarily long trajectories while maintaining a constant memory footprint through careful KV-cache management.

Text-Based Protocol Advantages

Because the ReAct loop relies purely on text parsing rather than structured API schemas, it remains backend-agnostic. The pattern works identically across OpenAI-compatible APIs, local Ollama deployments, or custom vLLM servers. The chapter6/phone-agent/agent.py file demonstrates this flexibility in a production voice-based agent that processes speech-to-text outputs through the same [[tool]] protocol.

Practical Code Examples

Minimal ReAct Implementation

The following Python snippet illustrates the core loop mechanics using the regex pattern found in the repository:

import re
import json
from typing import Any, Dict, List, Tuple

TOOL_PATTERN = r'\[\[([\w_]+)\]\]\s*(\{.*?\})'

def react_loop(
    llm_call: Any,
    tools: Dict[str, Any],
    user_query: str,
    max_iters: int = 10
) -> Tuple[str, List[dict]]:
    """Execute ReAct loop with text-based tool calling."""
    history = [{"role": "user", "content": user_query}]
    trajectory = []
    
    for i in range(max_iters):
        prompt = "\n".join(f"{m['role']}: {m['content']}" for m in history)
        response = llm_call(prompt)
        
        match = re.search(TOOL_PATTERN, response)
        if not match:
            return response.strip(), trajectory
            
        tool_name, args_json = match.groups()
        args = json.loads(args_json)
        
        try:
            observation = tools[tool_name](**args) if tool_name in tools else f"Error: unknown tool {tool_name}"
        except Exception as e:
            observation = f"Error: {e}"
            
        trajectory.append({
            "iteration": i + 1,
            "thought": response,
            "tool": tool_name,
            "observation": observation
        })
        
        history.extend([
            {"role": "assistant", "content": response},
            {"role": "observation", "content": str(observation)}
        ])
    
    raise RuntimeError(f"Max iterations ({max_iters}) exceeded")

Production Agent Classes

For production use, the repository provides specialized agent classes that handle streaming and error recovery:

from chapter2.local_llm_serving.agent import ReActLLMAgent
from chapter2.local_llm_serving.tools import search, calculate

agent = ReActLLMAgent(
    llm_provider="ollama",
    tools={"search": search, "calc": calculate},
    max_iterations=12
)

answer, steps = agent.run("What is the population density of Tokyo?")

The ReActAttentionAgent class in chapter2/attention_visualization/main.py extends this pattern with attention weight logging, enabling researchers to trace which tokens the model focuses on during reasoning versus tool selection phases.

Advanced Applications Across the Repository

The ReAct pattern serves as the foundation for several specialized agent architectures in the ai-agent-book codebase.

Agentic RAG Integration

In chapter3/agentic-rag/agent.py, the ReAct loop augments retrieval-augmented generation systems. Rather than performing a single retrieval, the agent can iteratively refine search queries based on intermediate findings, using the Observe phase to evaluate document relevance before generating final answers.

Dynamic Tool Discovery

The implementation in chapter4/active-tool-discovery/agent.py leverages the flexible [[tool_name]] syntax to discover and invoke tools that weren't explicitly programmed at startup. The agent reasons about available capabilities and constructs appropriate tool calls during execution, demonstrating the protocol's extensibility.

Multimodal Phone Agents

chapter6/phone-agent/agent.py applies the ReAct loop to voice-based interactions, where observations include transcribed speech and system state changes. This production-grade example shows how the same text protocol handles real-time constraints and error recovery in telephony environments.

Summary

  • The ReAct loop cycles through Reasoning (thought generation), Acting (tool execution), and Observing (result processing) to solve complex, multi-step tasks.
  • Text-based protocol using [[tool_name]] {"args"} syntax ensures backend-agnostic compatibility across OpenAI, Ollama, and vLLM providers.
  • Termination conditions occur when the LLM returns text without tool patterns or when max_iterations limits are reached.
  • Repository implementations span from basic loops in chapter2/local_llm_serving/agent.py to specialized agents for RAG, tool discovery, and phone systems.
  • KV-cache optimization in chapter2/kv-cache/agent.py demonstrates efficient context management during long-running agent trajectories.

Frequently Asked Questions

How does ReAct differ from standard chain-of-thought prompting?

Chain-of-thought prompting generates intermediate reasoning steps but produces a final answer in a single pass. ReAct interleaves reasoning with active tool execution, allowing the model to gather external data mid-stream and adjust its plan based on observations. This feedback loop enables ReAct agents to handle tasks requiring real-time information or calculations that exceed the LLM's parametric knowledge.

What determines when the ReAct loop stops executing?

The loop terminates under two conditions according to the source code: when the LLM response contains no tool-call pattern (indicating the model has provided a final answer), or when the iteration count reaches the max_iterations threshold defined in configuration files like chapter2/kv-cache/config.py. This dual safety mechanism prevents infinite loops while allowing natural completion.

Can ReAct agents work with local open-source models?

Yes. The repository demonstrates ReAct compatibility with local inference engines including Ollama and vLLM. Since the protocol relies solely on text generation and regex parsing—without requiring structured function-calling schemas—it works with any model capable of producing the [[tool_name]] syntax. The chapter2/local_llm_serving/agent.py file specifically implements streaming interfaces for local deployment scenarios.

How does the observation phase handle tool failures?

When a tool raises an exception or returns an error, the observation phase captures this as text (e.g., "Error while running search: timeout"). This error observation is appended to the conversation history, allowing the LLM to reason about the failure in the next iteration and potentially retry with different parameters or select an alternative tool. This error-handling capability, visible in the loop implementations, makes ReAct robust to unreliable external APIs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →