How the AI Agent Book Defines the Agent-Environment Interaction Loop

The AI Agent Book defines the Agent-Environment interaction loop as a closed-loop system where the Environment supplies observations, the Agent (comprising an LLM, context, and tools) processes them to choose actions, and those actions change the Environment to produce the next observation.

The Agent-Environment interaction loop forms the foundational architecture of autonomous AI systems according to bojieli/ai-agent-book. This conceptual framework, introduced in Chapter 1, bridges classical reinforcement learning with modern LLM-based agents to create a practical design pattern for building autonomous systems.

Core Definition: A Closed-Loop System

In book-en/chapter1.md (lines 21-26), the book establishes the fundamental relationship:

"From the classical reinforcement-learning and control perspective, the Agent and the Environment are two sides of a closed-loop interaction, not components of one another. The Environment returns an observation, the Agent uses its context to choose the next action, and that action changes the Environment's state, producing the next observation."

This definition explicitly separates the Agent from the Environment, treating them as distinct entities that communicate through a well-defined interface rather than nested components.

Two Abstraction Levels

The book structures the Agent-Environment interaction loop across two nested levels:

Outer Level: Agent ↔ Environment

At this level, the interaction occurs between the complete Agent system and the external world:

  • Environment scope: Filesystems, databases, web pages, users, other agents, and physical or simulated worlds
  • Agent receives: Observations from the external Environment
  • Agent produces: Actions via tool invocations that modify the Environment

Inner Level: Model ↔ Harness

Inside the Agent boundary, another loop operates between components:

  • Model: The LLM that makes policy decisions based on context
  • Harness: The runtime that builds context, exposes tool interfaces, maintains state, and enforces safety constraints

The Harness acts as the intermediary between the raw LLM and the external Environment, handling validation, execution, and feedback collection.

The Four-Step Loop Cycle

The complete Agent-Environment interaction loop executes continuously through these phases:

  1. Observe — Gather the latest state from the Environment
  2. Reason — The LLM processes current context (observation + conversation history)
  3. Act — The Agent invokes one or more tools, producing actions affecting the Environment
  4. Feedback — The Environment's response feeds back into context, closing the loop

ReAct Pattern Implementation

The book formalizes this as the ReAct (Reason → Act → Observe) pattern, detailed in book-en/chapter1.md (lines 61-66). This pattern iterates until task completion, with each cycle potentially generating tool calls or final answers.

Pseudocode Implementation

The concrete implementation from Figure 1-4 (book-en/chapter1.md, lines 73-85) demonstrates this structure:


# Initialise the trajectory with the user's request

trajectory = [user_request]

while True:
    # 1️⃣ Build the full context: static prefix (system prompt + tool defs) + trajectory

    context = stable_prefix + trajectory

    # 2️⃣ LLM decides what to do next (reasoning + optional tool calls)

    decision = Model(context)

    # Append the LLM's response to the trajectory

    trajectory.append(decision)

    # 3️⃣ If the decision contains no tool calls, we have reached a final answer

    if not decision.tool_calls:
        return decision.content   # final answer to the user

    # 4️⃣ Execute each tool call, observe results, and feed them back

    for call in decision.tool_calls:
        validated = Harness.validate(call)          # safety checks

        observation = Environment.execute(validated) # perform action

        trajectory.append(observation)               # add result to history

The trajectory serves as the persistent memory structure, accumulating the full history of thoughts, actions, and observations across loop iterations. The stable_prefix contains system prompts and tool definitions that remain constant.

Key Source Files

File Significance
book-en/chapter1.md Core definition of the Agent-Environment loop; outer/inner abstraction layers; ReAct control flow
book-en/chapter1.md (Figures 1-1 & 1-2) Visual diagrams of the loop architecture and three-level update hierarchy
book-en/chapter1.md ("The ReAct Loop" section) Executable pseudocode implementing the interaction loop
extras/agent-lab/SCHEMA.md Message schema definitions for tool calls within the loop
slides/lesson-29.md Reinforcement learning visualization of the agent-environment loop (Figure 7-1)

Summary

  • The Agent-Environment interaction loop is a closed-loop system with bidirectional information flow, not a hierarchical containment structure
  • Two abstraction levels separate external interaction (Agent-Environment) from internal processing (Model-Harness)
  • The ReAct pattern provides a concrete implementation: Reason → Act → Observe, iterating until completion
  • The trajectory data structure maintains state across loop iterations, enabling multi-step reasoning
  • Safety enforcement occurs at the Harness level through validation before Environment execution

Frequently Asked Questions

What separates the Agent from the Environment in this framework?

The book explicitly treats them as "two sides of a closed-loop interaction, not components of one another." This means the Agent does not contain the Environment, nor vice versa—they communicate through a defined observation-action interface. The Environment includes all external systems the Agent can perceive and modify: files, databases, APIs, other agents, and physical systems.

How does the Harness differ from the Model in the inner loop?

The Model is purely the LLM that generates text and decisions. The Harness is the operational runtime—it manages the trajectory (context history), validates tool calls for safety, executes actions against the Environment, and packages observations back into the context. The Harness implements the "plumbing" that makes the loop functional and safe.

Why does the loop use a trajectory structure rather than simple state?

The trajectory preserves the full historical record of observations, reasoning steps, and action results. This enables:

  • In-context learning from previous loop iterations
  • Debugging and transparency of the Agent's reasoning chain
  • Recovery from errors by maintaining context across failures
  • Multi-turn interactions where earlier observations inform later decisions

According to the source code analysis, currency conversion examples demonstrate how tool results append to the trajectory after each execution, building progressively toward task completion.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →