How Harvey AI Handles Context Window Overflow and Token Limit Exceedances: 3-Layer Protection System

Harvey AI prevents context window overflow and token limit breaches through a coordinated three-layer system: adapter-level token capping, runtime token accounting, and overflow detection with graceful failure handling.

Built for reliable enterprise AI workloads, the harveyai/harvey-labs repository implements robust safeguards to protect against the most common failure modes in LLM-powered applications: exceeding model context windows and blowing through token quotas. These mechanisms operate at the adapter, agent loop, and evaluation layers to ensure predictable behavior even when prompts grow unexpectedly large.

Layer 1: Model-Adapter Token Capping

The first line of defense lives in the adapter layer, where each provider-specific adapter enforces hard limits on both input context and output generation.

OpenAI Adapter: max_tokens Enforcement

In harness/adapters/openai.py, the OpenAIAdapter class accepts a max_tokens parameter that defaults to 128,000 tokens. This value is passed directly to the OpenAI API as max_output_tokens, capping how much the model can generate in a single turn:


# harness/adapters/openai.py, lines 20-21

def __init__(..., max_tokens: int = 128_000, ...):
    self.max_tokens = max_tokens

# Lines 50-51: passed to API call

response = client.chat.completions.create(
    ...,
    max_output_tokens=self.max_tokens,
)

Anthropic Adapter: Model-Specific Ceilings

The AnthropicAdapter in harness/adapters/anthropic.py takes a more nuanced approach, deriving context window limits based on the specific Claude model variant. For Claude-Fable-5, it sets a 128,000-token ceiling:


# harness/adapters/anthropic.py, lines 53-55

max_tokens = self._get_model_context_window(self.model)  # e.g., 128_000 for Claude-Fable-5

response = client.messages.create(
    ...,
    max_tokens=max_tokens,
)

This model-aware mapping prevents the "one size fits all" problem that can waste tokens on smaller-context models or cause hard failures on larger ones.

Layer 2: Runtime Token Accounting

While adapters prevent per-turn excesses, the agent loop tracks cumulative usage across multi-turn conversations. Every ModelResponse carries precise token metadata, which the loop aggregates into running totals.

In harness/agent_loop.py, lines 76-78 update counters after each model invocation:


# harness/agent_loop.py

total_input_tokens += response.input_tokens
total_output_tokens += response.output_tokens

These counters surface in the final result dictionary, enabling cost tracking, debugging, and prompt optimization. Developers can inspect exactly where token budgets are spent and trim accordingly.

Layer 3: Context-Window Overflow Detection

When accumulated conversation history exceeds a provider's hard context limit, the adapter raises an exception. The agent loop catches these provider-specific error messages and converts them into structured failure states.

The detection logic in harness/agent_loop.py, lines 70-73, identifies two key error signatures:

  • "prompt is too long" — generic length rejection
  • "context_length_exceeded" — explicit window overflow

# harness/adapters/base.py defines ModelResponse with usage metadata

@dataclass
class ModelResponse:
    content: str
    input_tokens: int
    output_tokens: int
    stop_reason: Optional[str] = None

When caught, the loop sets context_overflow = True, aborts further turns, and surfaces the condition in the final result. This prevents silent truncation that could corrupt outputs or trigger cascading errors in downstream processing.

Judge-Side Truncation Guard

A fourth safeguard operates in the evaluation layer. When "judge" models (used for quality scoring or validation) hit their own token limits, the system checks stop_reason == "max_tokens" in evaluation/judge.py. This prevents hidden test failures where a truncated judgment would be misinterpreted as a valid conclusion.

Practical Implementation Examples

Setting Custom Token Limits

Override defaults when instantiating adapters for cost-sensitive or latency-critical workloads:

from harness.adapters.openai import OpenAIAdapter

# Cap output at 32,000 tokens for faster, cheaper responses

adapter = OpenAIAdapter(
    model="gpt-5.0",
    max_tokens=32_000  # Down from 128,000 default

)

result = run_agent(
    adapter,
    system_prompt="You are a contract analysis assistant.",
    user_prompt="Review this 500-page merger agreement...",
    tool_executor=my_tool_executor,
)

print(f"Tokens: {result['input_tokens']} in, {result['output_tokens']} out")

Checking for Overflow in Results

Always inspect the overflow flag before using agent outputs:

result = run_agent(adapter, ...)

if result.get("context_overflow"):
    print("⚠️ Context window overflow — prompt history exceeded model limits")
    # Trigger prompt compression or conversation summarization

else:
    process_output(result["content"])

Monitoring Cumulative Usage

For long-running sessions, track trends across multiple runs:

session_tokens = {"input": 0, "output": 0}

for task in task_queue:
    result = run_agent(adapter, user_prompt=task.description)
    session_tokens["input"] += result["input_tokens"]
    session_tokens["output"] += result["output_tokens"]
    
    if session_tokens["input"] > 800_000:
        print("Approaching context limit — forcing conversation reset")
        adapter.reset_conversation()

Key Source Files

File Responsibility
harness/adapters/openai.py OpenAI-specific token capping via max_output_tokens
harness/adapters/anthropic.py Model-aware context window derivation
harness/adapters/base.py ModelResponse dataclass with token metadata
harness/agent_loop.py Overflow detection, aggregation of total_input_tokens/total_output_tokens
evaluation/judge.py Truncation detection via stop_reason analysis

Summary

Harvey AI's approach to context window overflow and token limit exceedances rests on three coordinated mechanisms:

  • Adapter-level caps enforce provider-specific hard limits before any API call
  • Runtime accounting exposes cumulative usage for cost control and debugging
  • Graceful overflow handling converts provider errors into structured state rather than crashes

These layers work together to prevent silent failures, enable transparent monitoring, and maintain system reliability under unpredictable load conditions.

Frequently Asked Questions

What happens when a prompt exceeds the context window limit?

The agent loop catches the provider's error (either "prompt is too long" or "context_length_exceeded"), sets context_overflow = True in the result object, and halts further interaction. The caller receives a structured failure state rather than an unhandled exception or silently truncated output.

How does Harvey AI track token usage across multiple turns?

After each model call, harness/agent_loop.py accumulates input_tokens and output_tokens from the ModelResponse object into running totals (total_input_tokens and total_output_tokens). These values appear in the final result dictionary for post-run analysis.

Can I customize token limits for different models?

Yes. Both OpenAIAdapter and AnthropicAdapter accept configuration parameters — max_tokens for OpenAI, model-specific context windows for Anthropic — allowing per-deployment tuning based on cost, latency, or output length requirements.

What's the difference between max_tokens in adapters and context overflow detection?

max_tokens (or max_output_tokens) limits how many tokens the model generates in a single response. Context overflow detection catches when the input prompt exceeds the model's total context window. These are complementary: one prevents runaway generation, the other handles accumulated conversation history that grows too large.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →