How the Context Ablation Experiment Validates Essential Context Components for Agent Performance
The context ablation experiment systematically removes specific components from an AI agent's interaction history to prove that tool definitions, tool results, reasoning traces, and conversation history are each essential for reliable task completion.
The context ablation experiment in Chapter 1 of the bojieli/ai-agent-book repository isolates the impact of individual context components on agent reasoning. By running identical tasks under five distinct ContextMode configurations, the experiment demonstrates that stripping any single element—from tool schemas to conversation history—degrades performance or causes complete task failure. This systematic approach, implemented in chapter1/context/run_experiment_1_1.py, provides quantitative evidence for which parts of the communication history actually matter.
The Five Context Modes Tested
The experiment defines five distinct ablation arms in the ContextMode enum (located in chapter1/context/agent.py). Each arm removes a specific layer of context while keeping others intact, allowing precise measurement of individual component contributions.
FULL Mode: The Performance Baseline
The FULL mode retains all context components: tool definitions, tool results, provider-generated reasoning, and complete conversation history.
According to the source code in chapter1/context/run_experiment_1_1.py, this mode verifies three critical contracts:
tools_present_every_turn: Tool definitions are included in every API requesthistory_present_after_first_turn: Previous conversation turns are retained after the initial exchangereasoning_retained_after_first_turn: Provider reasoning traces persist in the context window
With full context available, the agent achieves canonical_answer_correct = True, establishing the baseline for successful task completion.
NO_TOOL_CALLS: Removing Tool Definitions
When running in NO_TOOL_CALLS mode, the experiment strips all "tools" and "tool_choice" fields from API requests before sending them to the provider.
The contract verification in evaluate_context_contract() asserts tools_absent_every_turn to confirm complete removal. This ablation demonstrates catastrophic failure: the agent produces zero tool actions (tool_action_count == 0) because it literally cannot see the available functions. Without tool definitions, the agent fails to generate the correct answer, proving that tool schemas are prerequisite for functional tool use.
NO_TOOL_RESULTS: Hiding Execution Feedback
The NO_TOOL_RESULTS mode retains tool definitions and allows the agent to invoke tools, but replaces the "content" of tool result messages with "[Tool result hidden due to context mode]".
The contract checks confirm:
tool_calls_retained: The agent can still see that it made a tool calltool_results_hidden: The actual execution output is masked
This creates a specific failure pattern: the agent enters a repetition loop (has_repeated_tool_action == True). Because it cannot see the results of its previous invocations, it calls the same tool repeatedly with identical parameters, unable to progress toward the solution. This demonstrates that tool execution feedback is necessary for adaptive reasoning.
NO_REASONING: Stripping Provider Reasoning Traces
In NO_REASONING mode, all reasoning content generated by the provider is filtered from the conversation history before the next turn. The contract verifies that provider_generated_reasoning exists in the raw output but reasoning_removed_from_history confirms it never reaches subsequent context.
The impact manifests as degraded canonical correctness (without_reasoning_degraded == True). While the agent may still execute tool actions, the quality of decision-making deteriorates because it lacks access to its own chain-of-thought from previous steps. This proves that reasoning traces are not debug output but functional dependencies for multi-step tasks.
NO_HISTORY: Eliminating Prior Conversation Turns
The NO_HISTORY mode implements a sliding window that retains only the static system prompt and the current user message, dropping all intermediate turns. The contract only_static_prefix_and_user_every_turn validates this isolation while tools_still_present confirms that tool definitions remain available.
Despite having access to tools, the agent exhibits repeated tool actions (has_repeated_tool_action == True) because it cannot reference previous steps or results. This mirrors the NO_TOOL_RESULTS failure mode but stems from different root causes—loss of temporal continuity rather than masked execution output—demonstrating that conversation history provides essential episodic memory for coherent task execution.
Running the Context Ablation Experiment
You can reproduce these findings using the orchestration script provided in the repository.
Executing the Full Five-Arm Experiment
Run all context modes against a specific provider and model:
python chapter1/context/run_experiment_1_1.py \
--provider kimi \
--model kimi-k3 \
--max-iterations 5 \
--output-dir ./validation/example_run
This generates an evidence.json file containing detailed results for each arm, including contract verification results and behavioral metrics.
Analyzing Specific Ablation Results
Inspect individual experiment arms programmatically to verify specific hypotheses:
import json
from pathlib import Path
evidence = json.loads(Path("./validation/example_run/evidence.json").read_text())
no_tool_results_arm = next(
a for a in evidence["arms"] if a["mode"] == "no_tool_results"
)
print("Tool calls retained:", no_tool_results_arm["context_contract"]["tool_calls_retained"])
print("Tool results hidden:", no_tool_results_arm["context_contract"]["tool_results_hidden"])
print("Repeated tool actions:", no_tool_results_arm["behavior"]["has_repeated_tool_action"])
Using Contract Evaluation for Custom Tests
The evaluate_context_contract() function is exposed for validating custom agent implementations:
from chapter1.context.run_experiment_1_1 import evaluate_context_contract, ContextMode
# Assuming `turns` is a list of recorded API interactions from your custom run
contract = evaluate_context_contract(ContextMode.NO_REASONING.value, turns)
print(contract["reasoning_removed_from_history"])
print(contract["provider_generated_reasoning"])
Key Source Files
Understanding the experiment architecture requires examining these specific files:
chapter1/context/run_experiment_1_1.py: Contains theevaluate_context_contract(),summarize_arm(), andanalyze()functions that orchestrate the five-arm experiment and verify context integrity.chapter1/context/agent.py: ImplementsContextAwareAgentand theContextModeenum that controls context filtering logic.chapter1/context/test_experiment_1_1.py: Unit tests verifying that each ablation arm behaves according to specification.chapter1/context/tests/test_agent.py: Tests for the underlying agent logic ensuring context mode implementations are correct.
Summary
The context ablation experiment provides empirical validation that agent performance depends on the complete interaction history, not just the current prompt:
- Tool definitions (
NO_TOOL_CALLS): Removal prevents all tool use, resulting in zero actions and task failure. - Tool results (
NO_TOOL_RESULTS): Hiding execution output causes repetition loops as the agent cannot learn from previous invocations. - Reasoning traces (
NO_REASONING): Stripping chain-of-thought degrades decision quality even when tools remain accessible. - Conversation history (
NO_HISTORY): Eliminating prior turns destroys episodic memory, causing the agent to repeat steps unnecessarily. - Full context (
FULL): The only configuration achieving consistent canonical correctness, proving that every component serves a functional purpose.
Frequently Asked Questions
What is a context ablation experiment?
A context ablation experiment is a controlled study that systematically removes specific components from an AI agent's input context to measure the impact on task performance. In the bojieli/ai-agent-book implementation, the experiment uses five distinct ContextMode settings to isolate the contributions of tool definitions, tool results, reasoning traces, and conversation history. This approach reveals which context elements are essential versus merely present.
Which context component causes the most severe performance degradation when removed?
Removing tool definitions (NO_TOOL_CALLS mode) causes total task failure with zero tool actions, representing the most severe degradation. While other ablations cause repetition or quality degradation, the absence of tool schemas prevents the agent from invoking capabilities entirely. This demonstrates that tool definitions are the foundational prerequisite for functional agent behavior.
How does the experiment verify that specific context components are actually removed?
The experiment uses contract verification functions defined in evaluate_context_contract() within run_experiment_1_1.py. These contracts assert specific conditions such as tools_absent_every_turn, tool_results_hidden, or only_static_prefix_and_user_every_turn. By programmatically inspecting the actual API request payloads sent to the provider, the experiment confirms that ablations are implemented correctly and not merely simulated.
Can the context ablation framework test custom agent implementations?
Yes. The evaluate_context_contract() function accepts a ContextMode value and a list of recorded API turns, allowing you to validate custom agents against the same contracts. Import the function from chapter1.context.run_experiment_1_1 and pass your agent's interaction history to verify compliance with specific context preservation requirements or to test how your agent behaves under specific ablation scenarios.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →