# What Is the ReAct Loop? A Deep Dive into AI Agent Task Execution

> Explore the ReAct loop and master AI agent task execution. Understand how Reason-Act-Observe empowers LLMs to solve complex problems by integrating reasoning, tool use, and observation.

- Repository: [Bojie Li/ai-agent-book](https://github.com/bojieli/ai-agent-book)
- Tags: deep-dive
- Published: 2026-08-26

---

**The ReAct loop** (Reason-Act-Observe) is an iterative orchestration pattern that enables large language models to solve complex tasks by alternating between internal reasoning, external tool calls, and observation processing until generating a final answer.

Modern AI agents require more than single-shot inference to handle real-world tasks effectively. The ReAct loop, extensively documented in the `bojieli/ai-agent-book` repository, provides a lightweight text-based protocol that allows LLMs to plan actions, execute tools, and adapt based on environmental feedback through a structured cycle of reasoning and acting.

## How the ReAct Loop Works

The ReAct pattern follows a strict three-phase cycle that repeats until task completion. Each iteration appends new context to a running transcript, enabling the model to maintain state across multiple tool invocations.

### The Reasoning Phase

During the **Reasoning** phase, the LLM generates a natural-language thought that describes the current plan or analysis. This internal monologue might state: *"I need to search Wikipedia for the capital of France before answering."*

According to the repository's implementation in [`chapter2/local_llm_serving/agent.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter2/local_llm_serving/agent.py), this reasoning step is appended to the conversation history as an assistant message, creating a persistent chain of thought visible to subsequent iterations.

### The Action Phase

The **Act** phase involves parsing the LLM's output for executable commands. The repository uses a specific text protocol where tool calls follow the syntax `[[tool_name]] {"arg": "value"}`.

In [`chapter2/prompt-engineering/tau_bench/agents/chat_react_agent.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter2/prompt-engineering/tau_bench/agents/chat_react_agent.py), the `ChatReActAgent` class implements this parsing using a regular expression pattern—`TOOL_PATTERN = r'\[\[([\w_]+)\]\]\s*(\{.*?\})'`—to identify tool names and their JSON arguments. When a match is found, the agent invokes the corresponding function from the available toolkit.

### The Observation Phase

After tool execution, the **Observe** phase feeds the raw result back into the prompt context as a structured observation. Whether the result is a successful API response, a calculation error, or empty data, this observation becomes part of the next iteration's input.

The implementation in [`chapter2/kv-cache/agent.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter2/kv-cache/agent.py) demonstrates how these observations are cached alongside previous thoughts to optimize subsequent LLM calls, while [`chapter2/attention_visualization/main.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter2/attention_visualization/main.py) provides `ReActStep` and `ReActAttentionAgent` classes for visualizing how attention weights shift across different phases of the loop.

## Implementation Architecture

The ReAct loop's power lies in its simplicity—a text-driven protocol that works across any LLM backend supporting conversational interfaces.

### Streaming Iteration Control

The core loop logic resides in [`chapter2/local_llm_serving/agent.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter2/local_llm_serving/agent.py), where each iteration performs:

1. **Logging**: Records the current step via `logger.info(f"ReAct iteration {iteration}")`
2. **Prompt Construction**: Assembles the accumulated conversation including all prior thoughts, tool calls, and observations
3. **LLM Inference**: Sends the compiled prompt to the model backend
4. **Pattern Matching**: Searches for `[[tool_name]]` syntax to determine if another action is required
5. **Termination Check**: Returns the final answer when no tool pattern is detected or when `max_iterations` is reached

This streaming approach allows the agent to handle arbitrarily long trajectories while maintaining a constant memory footprint through careful KV-cache management.

### Text-Based Protocol Advantages

Because the ReAct loop relies purely on text parsing rather than structured API schemas, it remains backend-agnostic. The pattern works identically across OpenAI-compatible APIs, local Ollama deployments, or custom vLLM servers. The [`chapter6/phone-agent/agent.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter6/phone-agent/agent.py) file demonstrates this flexibility in a production voice-based agent that processes speech-to-text outputs through the same `[[tool]]` protocol.

## Practical Code Examples

### Minimal ReAct Implementation

The following Python snippet illustrates the core loop mechanics using the regex pattern found in the repository:

```python
import re
import json
from typing import Any, Dict, List, Tuple

TOOL_PATTERN = r'\[\[([\w_]+)\]\]\s*(\{.*?\})'

def react_loop(
    llm_call: Any,
    tools: Dict[str, Any],
    user_query: str,
    max_iters: int = 10
) -> Tuple[str, List[dict]]:
    """Execute ReAct loop with text-based tool calling."""
    history = [{"role": "user", "content": user_query}]
    trajectory = []
    
    for i in range(max_iters):
        prompt = "\n".join(f"{m['role']}: {m['content']}" for m in history)
        response = llm_call(prompt)
        
        match = re.search(TOOL_PATTERN, response)
        if not match:
            return response.strip(), trajectory
            
        tool_name, args_json = match.groups()
        args = json.loads(args_json)
        
        try:
            observation = tools[tool_name](**args) if tool_name in tools else f"Error: unknown tool {tool_name}"
        except Exception as e:
            observation = f"Error: {e}"
            
        trajectory.append({
            "iteration": i + 1,
            "thought": response,
            "tool": tool_name,
            "observation": observation
        })
        
        history.extend([
            {"role": "assistant", "content": response},
            {"role": "observation", "content": str(observation)}
        ])
    
    raise RuntimeError(f"Max iterations ({max_iters}) exceeded")

```

### Production Agent Classes

For production use, the repository provides specialized agent classes that handle streaming and error recovery:

```python
from chapter2.local_llm_serving.agent import ReActLLMAgent
from chapter2.local_llm_serving.tools import search, calculate

agent = ReActLLMAgent(
    llm_provider="ollama",
    tools={"search": search, "calc": calculate},
    max_iterations=12
)

answer, steps = agent.run("What is the population density of Tokyo?")

```

The `ReActAttentionAgent` class in [`chapter2/attention_visualization/main.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter2/attention_visualization/main.py) extends this pattern with attention weight logging, enabling researchers to trace which tokens the model focuses on during reasoning versus tool selection phases.

## Advanced Applications Across the Repository

The ReAct pattern serves as the foundation for several specialized agent architectures in the `ai-agent-book` codebase.

### Agentic RAG Integration

In [`chapter3/agentic-rag/agent.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/agentic-rag/agent.py), the ReAct loop augments retrieval-augmented generation systems. Rather than performing a single retrieval, the agent can iteratively refine search queries based on intermediate findings, using the **Observe** phase to evaluate document relevance before generating final answers.

### Dynamic Tool Discovery

The implementation in [`chapter4/active-tool-discovery/agent.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter4/active-tool-discovery/agent.py) leverages the flexible `[[tool_name]]` syntax to discover and invoke tools that weren't explicitly programmed at startup. The agent reasons about available capabilities and constructs appropriate tool calls during execution, demonstrating the protocol's extensibility.

### Multimodal Phone Agents

[`chapter6/phone-agent/agent.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter6/phone-agent/agent.py) applies the ReAct loop to voice-based interactions, where observations include transcribed speech and system state changes. This production-grade example shows how the same text protocol handles real-time constraints and error recovery in telephony environments.

## Summary

- **The ReAct loop** cycles through Reasoning (thought generation), Acting (tool execution), and Observing (result processing) to solve complex, multi-step tasks.
- **Text-based protocol** using `[[tool_name]] {"args"}` syntax ensures backend-agnostic compatibility across OpenAI, Ollama, and vLLM providers.
- **Termination conditions** occur when the LLM returns text without tool patterns or when `max_iterations` limits are reached.
- **Repository implementations** span from basic loops in [`chapter2/local_llm_serving/agent.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter2/local_llm_serving/agent.py) to specialized agents for RAG, tool discovery, and phone systems.
- **KV-cache optimization** in [`chapter2/kv-cache/agent.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter2/kv-cache/agent.py) demonstrates efficient context management during long-running agent trajectories.

## Frequently Asked Questions

### How does ReAct differ from standard chain-of-thought prompting?

**Chain-of-thought** prompting generates intermediate reasoning steps but produces a final answer in a single pass. **ReAct** interleaves reasoning with active tool execution, allowing the model to gather external data mid-stream and adjust its plan based on observations. This feedback loop enables ReAct agents to handle tasks requiring real-time information or calculations that exceed the LLM's parametric knowledge.

### What determines when the ReAct loop stops executing?

The loop terminates under two conditions according to the source code: when the LLM response contains **no tool-call pattern** (indicating the model has provided a final answer), or when the iteration count reaches the `max_iterations` threshold defined in configuration files like [`chapter2/kv-cache/config.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter2/kv-cache/config.py). This dual safety mechanism prevents infinite loops while allowing natural completion.

### Can ReAct agents work with local open-source models?

Yes. The repository demonstrates ReAct compatibility with local inference engines including Ollama and vLLM. Since the protocol relies solely on text generation and regex parsing—without requiring structured function-calling schemas—it works with any model capable of producing the `[[tool_name]]` syntax. The [`chapter2/local_llm_serving/agent.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter2/local_llm_serving/agent.py) file specifically implements streaming interfaces for local deployment scenarios.

### How does the observation phase handle tool failures?

When a tool raises an exception or returns an error, the observation phase captures this as text (e.g., `"Error while running search: timeout"`). This error observation is appended to the conversation history, allowing the LLM to reason about the failure in the next iteration and potentially retry with different parameters or select an alternative tool. This error-handling capability, visible in the loop implementations, makes ReAct robust to unreliable external APIs.