How Trajectory is Managed in the ReAct Loop: A Complete Technical Breakdown

In the ReAct (Reason → Act → Observe) pattern, trajectory is managed by incrementally appending step dictionaries containing thought, action, action_input, and observation to a running Python list inside the loop, then persisting the complete sequence via save_react_trajectory and validating it with TrajectoryConsistencyChecker.

The bojieli/ai-agent-book repository implements a concrete ReAct agent that records every reasoning step into a structured trajectory. This article examines exactly how the trajectory is built, stored, and verified using the production code found in the attention visualization and trajectory verifier modules.

Building the Trajectory Inside the ReAct Loop

The core trajectory construction logic resides in chapter2/attention_visualization/main.py. Here, the agent initializes an empty list before entering the reasoning loop and appends a structured record after every tool invocation.

Step Initialization

Before the first iteration, the system prepares an empty container to hold the reasoning history:

trajectory: List[Dict[str, Any]] = []
iteration = 0

This trajectory variable will eventually contain the complete ordered list of steps taken by the agent.

Iterative Step Recording

Inside the while loop (approximately lines 380–420 in main.py), each iteration captures the full ReAct cycle. The model generates a thought, optionally selects a tool, executes that tool, and records the observation:

while not done and iteration < max_iterations:
    # 1️⃣ Reasoning – the model returns a "thought" string

    thought = model.think(...)
    
    # 2️⃣ Action – the model may emit a tool call

    action, action_input = extract_action(response)
    
    # 3️⃣ Observation – the tool is executed and its result captured

    observation = run_tool(action, action_input)
    
    # 4️⃣ Record the step

    trajectory.append({
        "thought": thought,
        "action": action,
        "action_input": action_input,
        "observation": observation,
    })
    iteration += 1

Each dictionary appended to the list represents a single ReActStep containing the four essential fields that reconstruct the agent's reasoning path.

Loop Termination and Completion

The loop terminates when the model returns a final answer without requesting another tool, or when the iteration count hits max_iterations. At this point, the trajectory variable holds the complete, ordered record of the agent's reasoning process, ready for persistence or analysis.

Persisting and Verifying the Trajectory

Once the ReAct loop finishes, the accumulated trajectory moves through two additional stages: JSON serialization and semantic validation.

Saving to JSON

Around line 480 in chapter2/attention_visualization/main.py, the helper function save_react_trajectory handles persistence. It accepts the original query, the list of steps, and the final answer, then writes a structured JSON file:

def save_react_trajectory(query: str, steps: List[ReActStep], final_answer: str) -> None:
    trajectory = {
        "query": query,
        "steps": [s.asdict() for s in steps],
        "final_answer": final_answer,
    }
    with open(Path("trajectories") / f"{timestamp()}.json", "w") as f:
        json.dump(trajectory, f, indent=2)

This produces a timestamped JSON file in the trajectories/ directory, preserving the complete execution context for later audit or replay.

Trajectory Validation

The trajectory-verifier package in chapter9/trajectory-verifier/verifier.py provides the TrajectoryConsistencyChecker class to validate saved trajectories. The verifier loads the JSON and runs a battery of structural and semantic checks:

checker = TrajectoryConsistencyChecker()
report = checker.check_trajectory(loaded_trajectory)

The validator performs four critical checks:

  • Schema Verification (verify_step_schema): Ensures every step contains exactly the keys thought, action, action_input, and observation.
  • Action-Observation Pairing (verify_action_observation_order): Confirms that observations follow their corresponding actions in the correct sequence.
  • Termination Validation (verify_termination): Verifies the trajectory ends with a plain final answer rather than an unfinished tool call.
  • Score Aggregation (aggregate_scores): Computes a compliance score from 0 to 100 based on violation severity.

The demonstration script chapter9/trajectory-verifier/demo.py illustrates loading a saved trajectory and invoking these checks to generate a human-readable report.

Complete Execution Flow

Putting the components together, a full ReAct execution with trajectory management follows this pattern:


# 1️⃣ Run the ReAct agent

steps, answer = run_react_agent(user_query)

# 2️⃣ Persist the trajectory

save_react_trajectory(user_query, steps, answer)

# 3️⃣ Verify the saved trajectory (optional, used in the book's tests)

from trajectory_verifier import TrajectoryConsistencyChecker
checker = TrajectoryConsistencyChecker()
report = checker.check_trajectory(load_latest_trajectory())
print(report.summary())

The run_react_agent function is implemented in chapter2/attention_visualization/main.py and serves as the main entry point that orchestrates the loop, step collection, and handoff to persistence.

Summary

  • Trajectory construction occurs incrementally inside the ReAct loop via trajectory.append() in chapter2/attention_visualization/main.py, creating a list of dictionaries with thought, action, action_input, and observation.
  • Persistence is handled by save_react_trajectory, which serializes the complete trajectory plus metadata to timestamped JSON files.
  • Validation is performed by TrajectoryConsistencyChecker in chapter9/trajectory-verifier/verifier.py, ensuring structural integrity, correct ordering, and proper termination.
  • Verification checks include schema validation, action-observation pairing, termination verification, and compliance scoring.

Frequently Asked Questions

What data structure holds the ReAct trajectory?

The trajectory is stored as a Python List[Dict[str, Any]] where each dictionary represents one ReAct step. According to the source code in chapter2/attention_visualization/main.py, every dictionary contains four required string keys: thought, action, action_input, and observation.

How does the trajectory verifier check for structural errors?

The TrajectoryConsistencyChecker class in chapter9/trajectory-verifier/verifier.py implements verify_step_schema to ensure each step contains exactly the required keys without extras. It also runs verify_action_observation_order to confirm that every observation follows its generating action in the sequence.

Can the trajectory be saved without running the verifier?

Yes. The save_react_trajectory function in chapter2/attention_visualization/main.py operates independently of the verification system. Verification via TrajectoryConsistencyChecker is optional and typically used for testing or audit purposes, as demonstrated in chapter9/trajectory-verifier/demo.py.

What determines when the ReAct loop stops recording steps?

The loop terminates when the model returns a final answer string instead of a tool call, or when the iteration counter reaches max_iterations. At that moment, the complete trajectory list is finalized and passed to the persistence layer, ensuring no partial or dangling tool calls remain in the recorded history.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →