How Harvey AI Handles Context Window Overflow and Token Limit Exceedances: 3-Layer Protection System
Harvey AI prevents context window overflow and token limit breaches through a coordinated three-layer system: adapter-level token capping, runtime token accounting, and overflow detection with graceful failure handling.
Built for reliable enterprise AI workloads, the harveyai/harvey-labs repository implements robust safeguards to protect against the most common failure modes in LLM-powered applications: exceeding model context windows and blowing through token quotas. These mechanisms operate at the adapter, agent loop, and evaluation layers to ensure predictable behavior even when prompts grow unexpectedly large.
Layer 1: Model-Adapter Token Capping
The first line of defense lives in the adapter layer, where each provider-specific adapter enforces hard limits on both input context and output generation.
OpenAI Adapter: max_tokens Enforcement
In harness/adapters/openai.py, the OpenAIAdapter class accepts a max_tokens parameter that defaults to 128,000 tokens. This value is passed directly to the OpenAI API as max_output_tokens, capping how much the model can generate in a single turn:
# harness/adapters/openai.py, lines 20-21
def __init__(..., max_tokens: int = 128_000, ...):
self.max_tokens = max_tokens
# Lines 50-51: passed to API call
response = client.chat.completions.create(
...,
max_output_tokens=self.max_tokens,
)
Anthropic Adapter: Model-Specific Ceilings
The AnthropicAdapter in harness/adapters/anthropic.py takes a more nuanced approach, deriving context window limits based on the specific Claude model variant. For Claude-Fable-5, it sets a 128,000-token ceiling:
# harness/adapters/anthropic.py, lines 53-55
max_tokens = self._get_model_context_window(self.model) # e.g., 128_000 for Claude-Fable-5
response = client.messages.create(
...,
max_tokens=max_tokens,
)
This model-aware mapping prevents the "one size fits all" problem that can waste tokens on smaller-context models or cause hard failures on larger ones.
Layer 2: Runtime Token Accounting
While adapters prevent per-turn excesses, the agent loop tracks cumulative usage across multi-turn conversations. Every ModelResponse carries precise token metadata, which the loop aggregates into running totals.
In harness/agent_loop.py, lines 76-78 update counters after each model invocation:
# harness/agent_loop.py
total_input_tokens += response.input_tokens
total_output_tokens += response.output_tokens
These counters surface in the final result dictionary, enabling cost tracking, debugging, and prompt optimization. Developers can inspect exactly where token budgets are spent and trim accordingly.
Layer 3: Context-Window Overflow Detection
When accumulated conversation history exceeds a provider's hard context limit, the adapter raises an exception. The agent loop catches these provider-specific error messages and converts them into structured failure states.
The detection logic in harness/agent_loop.py, lines 70-73, identifies two key error signatures:
"prompt is too long"— generic length rejection"context_length_exceeded"— explicit window overflow
# harness/adapters/base.py defines ModelResponse with usage metadata
@dataclass
class ModelResponse:
content: str
input_tokens: int
output_tokens: int
stop_reason: Optional[str] = None
When caught, the loop sets context_overflow = True, aborts further turns, and surfaces the condition in the final result. This prevents silent truncation that could corrupt outputs or trigger cascading errors in downstream processing.
Judge-Side Truncation Guard
A fourth safeguard operates in the evaluation layer. When "judge" models (used for quality scoring or validation) hit their own token limits, the system checks stop_reason == "max_tokens" in evaluation/judge.py. This prevents hidden test failures where a truncated judgment would be misinterpreted as a valid conclusion.
Practical Implementation Examples
Setting Custom Token Limits
Override defaults when instantiating adapters for cost-sensitive or latency-critical workloads:
from harness.adapters.openai import OpenAIAdapter
# Cap output at 32,000 tokens for faster, cheaper responses
adapter = OpenAIAdapter(
model="gpt-5.0",
max_tokens=32_000 # Down from 128,000 default
)
result = run_agent(
adapter,
system_prompt="You are a contract analysis assistant.",
user_prompt="Review this 500-page merger agreement...",
tool_executor=my_tool_executor,
)
print(f"Tokens: {result['input_tokens']} in, {result['output_tokens']} out")
Checking for Overflow in Results
Always inspect the overflow flag before using agent outputs:
result = run_agent(adapter, ...)
if result.get("context_overflow"):
print("⚠️ Context window overflow — prompt history exceeded model limits")
# Trigger prompt compression or conversation summarization
else:
process_output(result["content"])
Monitoring Cumulative Usage
For long-running sessions, track trends across multiple runs:
session_tokens = {"input": 0, "output": 0}
for task in task_queue:
result = run_agent(adapter, user_prompt=task.description)
session_tokens["input"] += result["input_tokens"]
session_tokens["output"] += result["output_tokens"]
if session_tokens["input"] > 800_000:
print("Approaching context limit — forcing conversation reset")
adapter.reset_conversation()
Key Source Files
| File | Responsibility |
|---|---|
harness/adapters/openai.py |
OpenAI-specific token capping via max_output_tokens |
harness/adapters/anthropic.py |
Model-aware context window derivation |
harness/adapters/base.py |
ModelResponse dataclass with token metadata |
harness/agent_loop.py |
Overflow detection, aggregation of total_input_tokens/total_output_tokens |
evaluation/judge.py |
Truncation detection via stop_reason analysis |
Summary
Harvey AI's approach to context window overflow and token limit exceedances rests on three coordinated mechanisms:
- Adapter-level caps enforce provider-specific hard limits before any API call
- Runtime accounting exposes cumulative usage for cost control and debugging
- Graceful overflow handling converts provider errors into structured state rather than crashes
These layers work together to prevent silent failures, enable transparent monitoring, and maintain system reliability under unpredictable load conditions.
Frequently Asked Questions
What happens when a prompt exceeds the context window limit?
The agent loop catches the provider's error (either "prompt is too long" or "context_length_exceeded"), sets context_overflow = True in the result object, and halts further interaction. The caller receives a structured failure state rather than an unhandled exception or silently truncated output.
How does Harvey AI track token usage across multiple turns?
After each model call, harness/agent_loop.py accumulates input_tokens and output_tokens from the ModelResponse object into running totals (total_input_tokens and total_output_tokens). These values appear in the final result dictionary for post-run analysis.
Can I customize token limits for different models?
Yes. Both OpenAIAdapter and AnthropicAdapter accept configuration parameters — max_tokens for OpenAI, model-specific context windows for Anthropic — allowing per-deployment tuning based on cost, latency, or output length requirements.
What's the difference between max_tokens in adapters and context overflow detection?
max_tokens (or max_output_tokens) limits how many tokens the model generates in a single response. Context overflow detection catches when the input prompt exceeds the model's total context window. These are complementary: one prevents runaway generation, the other handles accumulated conversation history that grows too large.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →