How the Context Ablation Experiment Validates Essential Context Components for Agent Performance

The context ablation experiment systematically removes specific components from an AI agent's interaction history to prove that tool definitions, tool results, reasoning traces, and conversation history are each essential for reliable task completion.

The context ablation experiment in Chapter 1 of the bojieli/ai-agent-book repository isolates the impact of individual context components on agent reasoning. By running identical tasks under five distinct ContextMode configurations, the experiment demonstrates that stripping any single element—from tool schemas to conversation history—degrades performance or causes complete task failure. This systematic approach, implemented in chapter1/context/run_experiment_1_1.py, provides quantitative evidence for which parts of the communication history actually matter.

The Five Context Modes Tested

The experiment defines five distinct ablation arms in the ContextMode enum (located in chapter1/context/agent.py). Each arm removes a specific layer of context while keeping others intact, allowing precise measurement of individual component contributions.

FULL Mode: The Performance Baseline

The FULL mode retains all context components: tool definitions, tool results, provider-generated reasoning, and complete conversation history.

According to the source code in chapter1/context/run_experiment_1_1.py, this mode verifies three critical contracts:

  • tools_present_every_turn: Tool definitions are included in every API request
  • history_present_after_first_turn: Previous conversation turns are retained after the initial exchange
  • reasoning_retained_after_first_turn: Provider reasoning traces persist in the context window

With full context available, the agent achieves canonical_answer_correct = True, establishing the baseline for successful task completion.

NO_TOOL_CALLS: Removing Tool Definitions

When running in NO_TOOL_CALLS mode, the experiment strips all "tools" and "tool_choice" fields from API requests before sending them to the provider.

The contract verification in evaluate_context_contract() asserts tools_absent_every_turn to confirm complete removal. This ablation demonstrates catastrophic failure: the agent produces zero tool actions (tool_action_count == 0) because it literally cannot see the available functions. Without tool definitions, the agent fails to generate the correct answer, proving that tool schemas are prerequisite for functional tool use.

NO_TOOL_RESULTS: Hiding Execution Feedback

The NO_TOOL_RESULTS mode retains tool definitions and allows the agent to invoke tools, but replaces the "content" of tool result messages with "[Tool result hidden due to context mode]".

The contract checks confirm:

  • tool_calls_retained: The agent can still see that it made a tool call
  • tool_results_hidden: The actual execution output is masked

This creates a specific failure pattern: the agent enters a repetition loop (has_repeated_tool_action == True). Because it cannot see the results of its previous invocations, it calls the same tool repeatedly with identical parameters, unable to progress toward the solution. This demonstrates that tool execution feedback is necessary for adaptive reasoning.

NO_REASONING: Stripping Provider Reasoning Traces

In NO_REASONING mode, all reasoning content generated by the provider is filtered from the conversation history before the next turn. The contract verifies that provider_generated_reasoning exists in the raw output but reasoning_removed_from_history confirms it never reaches subsequent context.

The impact manifests as degraded canonical correctness (without_reasoning_degraded == True). While the agent may still execute tool actions, the quality of decision-making deteriorates because it lacks access to its own chain-of-thought from previous steps. This proves that reasoning traces are not debug output but functional dependencies for multi-step tasks.

NO_HISTORY: Eliminating Prior Conversation Turns

The NO_HISTORY mode implements a sliding window that retains only the static system prompt and the current user message, dropping all intermediate turns. The contract only_static_prefix_and_user_every_turn validates this isolation while tools_still_present confirms that tool definitions remain available.

Despite having access to tools, the agent exhibits repeated tool actions (has_repeated_tool_action == True) because it cannot reference previous steps or results. This mirrors the NO_TOOL_RESULTS failure mode but stems from different root causes—loss of temporal continuity rather than masked execution output—demonstrating that conversation history provides essential episodic memory for coherent task execution.

Running the Context Ablation Experiment

You can reproduce these findings using the orchestration script provided in the repository.

Executing the Full Five-Arm Experiment

Run all context modes against a specific provider and model:

python chapter1/context/run_experiment_1_1.py \
  --provider kimi \
  --model kimi-k3 \
  --max-iterations 5 \
  --output-dir ./validation/example_run

This generates an evidence.json file containing detailed results for each arm, including contract verification results and behavioral metrics.

Analyzing Specific Ablation Results

Inspect individual experiment arms programmatically to verify specific hypotheses:

import json
from pathlib import Path

evidence = json.loads(Path("./validation/example_run/evidence.json").read_text())
no_tool_results_arm = next(
    a for a in evidence["arms"] if a["mode"] == "no_tool_results"
)

print("Tool calls retained:", no_tool_results_arm["context_contract"]["tool_calls_retained"])
print("Tool results hidden:", no_tool_results_arm["context_contract"]["tool_results_hidden"])
print("Repeated tool actions:", no_tool_results_arm["behavior"]["has_repeated_tool_action"])

Using Contract Evaluation for Custom Tests

The evaluate_context_contract() function is exposed for validating custom agent implementations:

from chapter1.context.run_experiment_1_1 import evaluate_context_contract, ContextMode

# Assuming `turns` is a list of recorded API interactions from your custom run

contract = evaluate_context_contract(ContextMode.NO_REASONING.value, turns)
print(contract["reasoning_removed_from_history"])
print(contract["provider_generated_reasoning"])

Key Source Files

Understanding the experiment architecture requires examining these specific files:

Summary

The context ablation experiment provides empirical validation that agent performance depends on the complete interaction history, not just the current prompt:

  • Tool definitions (NO_TOOL_CALLS): Removal prevents all tool use, resulting in zero actions and task failure.
  • Tool results (NO_TOOL_RESULTS): Hiding execution output causes repetition loops as the agent cannot learn from previous invocations.
  • Reasoning traces (NO_REASONING): Stripping chain-of-thought degrades decision quality even when tools remain accessible.
  • Conversation history (NO_HISTORY): Eliminating prior turns destroys episodic memory, causing the agent to repeat steps unnecessarily.
  • Full context (FULL): The only configuration achieving consistent canonical correctness, proving that every component serves a functional purpose.

Frequently Asked Questions

What is a context ablation experiment?

A context ablation experiment is a controlled study that systematically removes specific components from an AI agent's input context to measure the impact on task performance. In the bojieli/ai-agent-book implementation, the experiment uses five distinct ContextMode settings to isolate the contributions of tool definitions, tool results, reasoning traces, and conversation history. This approach reveals which context elements are essential versus merely present.

Which context component causes the most severe performance degradation when removed?

Removing tool definitions (NO_TOOL_CALLS mode) causes total task failure with zero tool actions, representing the most severe degradation. While other ablations cause repetition or quality degradation, the absence of tool schemas prevents the agent from invoking capabilities entirely. This demonstrates that tool definitions are the foundational prerequisite for functional agent behavior.

How does the experiment verify that specific context components are actually removed?

The experiment uses contract verification functions defined in evaluate_context_contract() within run_experiment_1_1.py. These contracts assert specific conditions such as tools_absent_every_turn, tool_results_hidden, or only_static_prefix_and_user_every_turn. By programmatically inspecting the actual API request payloads sent to the provider, the experiment confirms that ablations are implemented correctly and not merely simulated.

Can the context ablation framework test custom agent implementations?

Yes. The evaluate_context_contract() function accepts a ContextMode value and a list of recorded API turns, allowing you to validate custom agents against the same contracts. Import the function from chapter1.context.run_experiment_1_1 and pass your agent's interaction history to verify compliance with specific context preservation requirements or to test how your agent behaves under specific ablation scenarios.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →