How the Learning from Experience Experiment Evaluates an Agent's Ability to Learn from Its Running Trajectory
The I-S experiment uses a two-phase protocol where ExperienceAgent captures execution trajectories into a JSON database during an initial learning run, then retrieves and injects those experiences into subsequent prompts to measure whether the agent improves its performance by learning from its own running trajectory.
The bojieli/ai-agent-book repository implements a concrete methodology for testing whether AI agents can learn from their own execution history. The learning from experience experiment—designated as I-S in the codebase—evaluates an agent's capacity to capture, summarize, and reuse knowledge from its running trajectory to improve task completion. This evaluation framework is implemented in the ExperienceAgent class and provides a reproducible benchmark for trajectory-based learning capabilities.
The Two-Phase Evaluation Methodology
The learning from experience experiment operates through distinct capture and reuse phases that together measure whether an agent can effectively learn from its own execution trace.
Phase 1: Capture and Learn
In the first phase, the agent executes a task with learning_mode=True. During execution, each action is recorded via the capture_action method and stored in a current_trajectory list. When the run completes successfully (determined by _is_successful), the trajectory is passed to a TrajectorySummarizer, which generates a concise summary capturing key insights, approaches, and tools used. This summary is persisted to an on-disk JSON database (experience_db.json) keyed by a hash of the original question.
According to chapter9/gaia-experience/experience_agent.py (lines 54-66), the learning_mode flag activates trajectory capture, while lines 93-110 handle the _learn_from_success method that triggers summarization and storage. Success in this phase is confirmed by the presence of a new entry in the experience DB, the logged message "Learned from successful execution", and the entry's success flag set to True.
Phase 2: Apply and Evaluate
In the second phase, the same or a similar task is executed with apply_experience=True. Before invoking the LLM, the agent queries the experience database (and optionally a KnowledgeBase) via _get_relevant_experiences to retrieve past trajectories. These experiences are formatted using _format_experiences and prepended to the system prompt (lines 102-108 in experience_agent.py). The LLM then generates its answer under the guidance of the retrieved experiences.
Improvement is judged by comparing answer quality against a baseline run without injected experiences. The test harness verifies that the answer is non-empty, the success flag remains True, and the num_steps matches expectations, confirming that the agent successfully leveraged its stored trajectory knowledge.
Core Implementation Details
The experiment relies on specific components within the ExperienceAgent class:
- Trajectory Capture (lines 54-66): The
current_trajectorylist stores actions whenlearning_modeis active - Summarization & Storage (lines 93-110): The
_learn_from_successmethod persists summarized experiences toexperience_db.json - Experience Retrieval (lines 102-108): Relevant past experiences are injected into prompts when
apply_experienceis enabled
Practical Code Examples
The following examples demonstrate the two-step evaluation process used in the I-S experiment.
First, run a task and capture the trajectory:
# Run a task and let the agent learn from its own trajectory
agent = ExperienceAgent(
conf=agent_cfg,
learning_mode=True, # Capture & learn
apply_experience=False, # No reuse on this run
experience_db_path="./experience_db.json",
summarizer=MyTrajectorySummarizer(),
)
task = Task(input="What is the capital of France?")
response = await agent.execute_task(task) # → stores a new experience
Then, run the same task and apply the learned experience:
# Run the same task and let the agent apply the learned experience
agent = ExperienceAgent(
conf=agent_cfg,
learning_mode=False,
apply_experience=True, # Reuse past experiences
experience_db_path="./experience_db.json",
summarizer=MyTrajectorySummarizer(),
)
task = Task(input="What is the capital of France?")
response = await agent.execute_task(task) # → answer benefits from stored experience
The automated test suite mirrors this flow in tests/test_ch9_trajectory_consistency_checker.py:
# test_ch9_trajectory_consistency_checker.py (excerpt)
agent = ExperienceAgent(
conf=cfg,
learning_mode=True,
apply_experience=False,
experience_db_path=temp_db,
summarizer=FakeSummarizer(),
)
await agent.execute_task(simple_task) # learns
agent2 = ExperienceAgent(
conf=cfg,
learning_mode=False,
apply_experience=True,
experience_db_path=temp_db,
summarizer=FakeSummarizer(),
)
resp = await agent2.execute_task(simple_task) # should use stored experience
assert resp.success is True
Summary
- The I-S experiment evaluates whether an agent can learn from its own execution trajectory through a capture-and-reuse protocol.
- Phase 1 uses
learning_mode=Trueto record actions viacapture_actionand store summarized trajectories inexperience_db.jsonvia_learn_from_success. - Phase 2 uses
apply_experience=Trueto retrieve relevant experiences with_get_relevant_experiencesand inject them into the system prompt. - Success is measured by comparing performance metrics (answer quality,
successflags, andnum_steps) between baseline runs and experience-enhanced runs. - The implementation resides primarily in
chapter9/gaia-experience/experience_agent.pywith automated tests intests/test_ch9_trajectory_consistency_checker.py.
Frequently Asked Questions
What does the I-S designation stand for in the learning from experience experiment?
The I-S designation refers to the "Learning from Experience" experiment identifier used in the bojieli/ai-agent-book codebase. It represents a specific experimental condition where the agent is tested on its ability to learn from its own execution trajectory, as opposed to learning from external data or static examples.
How does the agent determine which experiences to retrieve from the database?
The agent uses the _get_relevant_experiences method to query the experience_db.json file using a hash of the current task's question as the lookup key. It may also query an optional KnowledgeBase for additional context. Retrieved experiences are then formatted via _format_experiences before being injected into the system prompt.
Can the experience database be shared across different agent instances?
Yes, the experience database is stored as a JSON file on disk (experience_db.json), making it persistent across different ExperienceAgent instances. Multiple agents can read from the same database path, enabling shared learning across distributed or sequential agent executions, provided they use compatible TrajectorySummarizer implementations.
What prevents the agent from being negatively influenced by failed trajectories?
The implementation specifically checks for successful execution via _is_successful before calling _learn_from_success. Only trajectories where the success flag is True are summarized and stored in the experience database. This ensures that the agent learns exclusively from successful execution paths and does not propagate errors or failed strategies to future tasks.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →