# How the AI Agent Book Defines the Agent-Environment Interaction Loop

> The AI Agent Book explains the agent-environment interaction loop: observations from the environment inform agent actions, which in turn modify the environment for the next cycle.

- Repository: [Bojie Li/ai-agent-book](https://github.com/bojieli/ai-agent-book)
- Tags: deep-dive
- Published: 2026-08-25

---

**The AI Agent Book defines the Agent-Environment interaction loop as a closed-loop system where the Environment supplies observations, the Agent (comprising an LLM, context, and tools) processes them to choose actions, and those actions change the Environment to produce the next observation.**

The Agent-Environment interaction loop forms the foundational architecture of autonomous AI systems according to bojieli/ai-agent-book. This conceptual framework, introduced in Chapter 1, bridges classical reinforcement learning with modern LLM-based agents to create a practical design pattern for building autonomous systems.

## Core Definition: A Closed-Loop System

In [`book-en/chapter1.md`](https://github.com/bojieli/ai-agent-book/blob/main/book-en/chapter1.md) (lines 21-26), the book establishes the fundamental relationship:

> "From the classical reinforcement-learning and control perspective, the **Agent and the Environment are two sides of a closed-loop interaction**, not components of one another. The Environment returns an observation, the Agent uses its context to choose the next action, and that action changes the Environment's state, producing the next observation."

This definition explicitly separates the **Agent** from the **Environment**, treating them as distinct entities that communicate through a well-defined interface rather than nested components.

## Two Abstraction Levels

The book structures the Agent-Environment interaction loop across two nested levels:

### Outer Level: Agent ↔ Environment

At this level, the interaction occurs between the complete Agent system and the external world:

- **Environment scope**: Filesystems, databases, web pages, users, other agents, and physical or simulated worlds
- **Agent receives**: Observations from the external Environment
- **Agent produces**: Actions via tool invocations that modify the Environment

### Inner Level: Model ↔ Harness

Inside the Agent boundary, another loop operates between components:

- **Model**: The LLM that makes policy decisions based on context
- **Harness**: The runtime that builds context, exposes tool interfaces, maintains state, and enforces safety constraints

The Harness acts as the intermediary between the raw LLM and the external Environment, handling validation, execution, and feedback collection.

## The Four-Step Loop Cycle

The complete Agent-Environment interaction loop executes continuously through these phases:

1. **Observe** — Gather the latest state from the Environment
2. **Reason** — The LLM processes current context (observation + conversation history)
3. **Act** — The Agent invokes one or more tools, producing actions affecting the Environment
4. **Feedback** — The Environment's response feeds back into context, closing the loop

## ReAct Pattern Implementation

The book formalizes this as the **ReAct** (Reason → Act → Observe) pattern, detailed in [`book-en/chapter1.md`](https://github.com/bojieli/ai-agent-book/blob/main/book-en/chapter1.md) (lines 61-66). This pattern iterates until task completion, with each cycle potentially generating tool calls or final answers.

### Pseudocode Implementation

The concrete implementation from Figure 1-4 ([`book-en/chapter1.md`](https://github.com/bojieli/ai-agent-book/blob/main/book-en/chapter1.md), lines 73-85) demonstrates this structure:

```python

# Initialise the trajectory with the user's request

trajectory = [user_request]

while True:
    # 1️⃣ Build the full context: static prefix (system prompt + tool defs) + trajectory

    context = stable_prefix + trajectory

    # 2️⃣ LLM decides what to do next (reasoning + optional tool calls)

    decision = Model(context)

    # Append the LLM's response to the trajectory

    trajectory.append(decision)

    # 3️⃣ If the decision contains no tool calls, we have reached a final answer

    if not decision.tool_calls:
        return decision.content   # final answer to the user

    # 4️⃣ Execute each tool call, observe results, and feed them back

    for call in decision.tool_calls:
        validated = Harness.validate(call)          # safety checks

        observation = Environment.execute(validated) # perform action

        trajectory.append(observation)               # add result to history

```

The **trajectory** serves as the persistent memory structure, accumulating the full history of thoughts, actions, and observations across loop iterations. The **stable_prefix** contains system prompts and tool definitions that remain constant.

## Key Source Files

| File | Significance |
|------|------------|
| [`book-en/chapter1.md`](https://github.com/bojieli/ai-agent-book/blob/main/book-en/chapter1.md) | Core definition of the Agent-Environment loop; outer/inner abstraction layers; ReAct control flow |
| [`book-en/chapter1.md`](https://github.com/bojieli/ai-agent-book/blob/main/book-en/chapter1.md) (Figures 1-1 & 1-2) | Visual diagrams of the loop architecture and three-level update hierarchy |
| [`book-en/chapter1.md`](https://github.com/bojieli/ai-agent-book/blob/main/book-en/chapter1.md) ("The ReAct Loop" section) | Executable pseudocode implementing the interaction loop |
| [`extras/agent-lab/SCHEMA.md`](https://github.com/bojieli/ai-agent-book/blob/main/extras/agent-lab/SCHEMA.md) | Message schema definitions for tool calls within the loop |
| [`slides/lesson-29.md`](https://github.com/bojieli/ai-agent-book/blob/main/slides/lesson-29.md) | Reinforcement learning visualization of the agent-environment loop (Figure 7-1) |

## Summary

- The **Agent-Environment interaction loop** is a closed-loop system with bidirectional information flow, not a hierarchical containment structure
- Two abstraction levels separate external interaction (Agent-Environment) from internal processing (Model-Harness)
- The **ReAct pattern** provides a concrete implementation: Reason → Act → Observe, iterating until completion
- The **trajectory** data structure maintains state across loop iterations, enabling multi-step reasoning
- **Safety enforcement** occurs at the Harness level through validation before Environment execution

## Frequently Asked Questions

### What separates the Agent from the Environment in this framework?

The book explicitly treats them as "two sides of a closed-loop interaction, not components of one another." This means the Agent does not contain the Environment, nor vice versa—they communicate through a defined observation-action interface. The Environment includes all external systems the Agent can perceive and modify: files, databases, APIs, other agents, and physical systems.

### How does the Harness differ from the Model in the inner loop?

The **Model** is purely the LLM that generates text and decisions. The **Harness** is the operational runtime—it manages the trajectory (context history), validates tool calls for safety, executes actions against the Environment, and packages observations back into the context. The Harness implements the "plumbing" that makes the loop functional and safe.

### Why does the loop use a trajectory structure rather than simple state?

The **trajectory** preserves the full historical record of observations, reasoning steps, and action results. This enables:
- In-context learning from previous loop iterations
- Debugging and transparency of the Agent's reasoning chain
- Recovery from errors by maintaining context across failures
- Multi-turn interactions where earlier observations inform later decisions

According to the source code analysis, currency conversion examples demonstrate how tool results append to the trajectory after each execution, building progressively toward task completion.