How the Learning from Experience Experiment Evaluates an Agent's Ability to Learn from Its Running Trajectory

The I-S experiment uses a two-phase protocol where ExperienceAgent captures execution trajectories into a JSON database during an initial learning run, then retrieves and injects those experiences into subsequent prompts to measure whether the agent improves its performance by learning from its own running trajectory.

The bojieli/ai-agent-book repository implements a concrete methodology for testing whether AI agents can learn from their own execution history. The learning from experience experiment—designated as I-S in the codebase—evaluates an agent's capacity to capture, summarize, and reuse knowledge from its running trajectory to improve task completion. This evaluation framework is implemented in the ExperienceAgent class and provides a reproducible benchmark for trajectory-based learning capabilities.

The Two-Phase Evaluation Methodology

The learning from experience experiment operates through distinct capture and reuse phases that together measure whether an agent can effectively learn from its own execution trace.

Phase 1: Capture and Learn

In the first phase, the agent executes a task with learning_mode=True. During execution, each action is recorded via the capture_action method and stored in a current_trajectory list. When the run completes successfully (determined by _is_successful), the trajectory is passed to a TrajectorySummarizer, which generates a concise summary capturing key insights, approaches, and tools used. This summary is persisted to an on-disk JSON database (experience_db.json) keyed by a hash of the original question.

According to chapter9/gaia-experience/experience_agent.py (lines 54-66), the learning_mode flag activates trajectory capture, while lines 93-110 handle the _learn_from_success method that triggers summarization and storage. Success in this phase is confirmed by the presence of a new entry in the experience DB, the logged message "Learned from successful execution", and the entry's success flag set to True.

Phase 2: Apply and Evaluate

In the second phase, the same or a similar task is executed with apply_experience=True. Before invoking the LLM, the agent queries the experience database (and optionally a KnowledgeBase) via _get_relevant_experiences to retrieve past trajectories. These experiences are formatted using _format_experiences and prepended to the system prompt (lines 102-108 in experience_agent.py). The LLM then generates its answer under the guidance of the retrieved experiences.

Improvement is judged by comparing answer quality against a baseline run without injected experiences. The test harness verifies that the answer is non-empty, the success flag remains True, and the num_steps matches expectations, confirming that the agent successfully leveraged its stored trajectory knowledge.

Core Implementation Details

The experiment relies on specific components within the ExperienceAgent class:

  • Trajectory Capture (lines 54-66): The current_trajectory list stores actions when learning_mode is active
  • Summarization & Storage (lines 93-110): The _learn_from_success method persists summarized experiences to experience_db.json
  • Experience Retrieval (lines 102-108): Relevant past experiences are injected into prompts when apply_experience is enabled

Practical Code Examples

The following examples demonstrate the two-step evaluation process used in the I-S experiment.

First, run a task and capture the trajectory:


# Run a task and let the agent learn from its own trajectory

agent = ExperienceAgent(
    conf=agent_cfg,
    learning_mode=True,          # Capture & learn

    apply_experience=False,      # No reuse on this run

    experience_db_path="./experience_db.json",
    summarizer=MyTrajectorySummarizer(),
)

task = Task(input="What is the capital of France?")
response = await agent.execute_task(task)   # → stores a new experience

Then, run the same task and apply the learned experience:


# Run the same task and let the agent apply the learned experience

agent = ExperienceAgent(
    conf=agent_cfg,
    learning_mode=False,
    apply_experience=True,       # Reuse past experiences

    experience_db_path="./experience_db.json",
    summarizer=MyTrajectorySummarizer(),
)

task = Task(input="What is the capital of France?")
response = await agent.execute_task(task)   # → answer benefits from stored experience

The automated test suite mirrors this flow in tests/test_ch9_trajectory_consistency_checker.py:


# test_ch9_trajectory_consistency_checker.py (excerpt)

agent = ExperienceAgent(
    conf=cfg,
    learning_mode=True,
    apply_experience=False,
    experience_db_path=temp_db,
    summarizer=FakeSummarizer(),
)
await agent.execute_task(simple_task)   # learns

agent2 = ExperienceAgent(
    conf=cfg,
    learning_mode=False,
    apply_experience=True,
    experience_db_path=temp_db,
    summarizer=FakeSummarizer(),
)
resp = await agent2.execute_task(simple_task)   # should use stored experience

assert resp.success is True

Summary

  • The I-S experiment evaluates whether an agent can learn from its own execution trajectory through a capture-and-reuse protocol.
  • Phase 1 uses learning_mode=True to record actions via capture_action and store summarized trajectories in experience_db.json via _learn_from_success.
  • Phase 2 uses apply_experience=True to retrieve relevant experiences with _get_relevant_experiences and inject them into the system prompt.
  • Success is measured by comparing performance metrics (answer quality, success flags, and num_steps) between baseline runs and experience-enhanced runs.
  • The implementation resides primarily in chapter9/gaia-experience/experience_agent.py with automated tests in tests/test_ch9_trajectory_consistency_checker.py.

Frequently Asked Questions

What does the I-S designation stand for in the learning from experience experiment?

The I-S designation refers to the "Learning from Experience" experiment identifier used in the bojieli/ai-agent-book codebase. It represents a specific experimental condition where the agent is tested on its ability to learn from its own execution trajectory, as opposed to learning from external data or static examples.

How does the agent determine which experiences to retrieve from the database?

The agent uses the _get_relevant_experiences method to query the experience_db.json file using a hash of the current task's question as the lookup key. It may also query an optional KnowledgeBase for additional context. Retrieved experiences are then formatted via _format_experiences before being injected into the system prompt.

Can the experience database be shared across different agent instances?

Yes, the experience database is stored as a JSON file on disk (experience_db.json), making it persistent across different ExperienceAgent instances. Multiple agents can read from the same database path, enabling shared learning across distributed or sequential agent executions, provided they use compatible TrajectorySummarizer implementations.

What prevents the agent from being negatively influenced by failed trajectories?

The implementation specifically checks for successful execution via _is_successful before calling _learn_from_success. Only trajectories where the success flag is True are summarized and stored in the experience database. This ensures that the agent learns exclusively from successful execution paths and does not propagate errors or failed strategies to future tasks.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →