# How Harvey AI Handles Context Window Overflow and Token Limit Exceedances: 3-Layer Protection System

> Discover how Harvey AI tackles context window overflow with its 3-layer protection system, including token capping, runtime accounting, and graceful failure.

- Repository: [Harvey/harvey-labs](https://github.com/harveyai/harvey-labs)
- Tags: deep-dive
- Published: 2026-08-11

---

**Harvey AI prevents context window overflow and token limit breaches through a coordinated three-layer system: adapter-level token capping, runtime token accounting, and overflow detection with graceful failure handling.**

Built for reliable enterprise AI workloads, the `harveyai/harvey-labs` repository implements robust safeguards to protect against the most common failure modes in LLM-powered applications: exceeding model context windows and blowing through token quotas. These mechanisms operate at the adapter, agent loop, and evaluation layers to ensure predictable behavior even when prompts grow unexpectedly large.

## Layer 1: Model-Adapter Token Capping

The first line of defense lives in the adapter layer, where each provider-specific adapter enforces hard limits on both input context and output generation.

### OpenAI Adapter: max_tokens Enforcement

In [`harness/adapters/openai.py`](https://github.com/harveyai/harvey-labs/blob/main/harness/adapters/openai.py), the `OpenAIAdapter` class accepts a `max_tokens` parameter that defaults to **128,000 tokens**. This value is passed directly to the OpenAI API as `max_output_tokens`, capping how much the model can generate in a single turn:

```python

# harness/adapters/openai.py, lines 20-21

def __init__(..., max_tokens: int = 128_000, ...):
    self.max_tokens = max_tokens

# Lines 50-51: passed to API call

response = client.chat.completions.create(
    ...,
    max_output_tokens=self.max_tokens,
)

```

### Anthropic Adapter: Model-Specific Ceilings

The `AnthropicAdapter` in [`harness/adapters/anthropic.py`](https://github.com/harveyai/harvey-labs/blob/main/harness/adapters/anthropic.py) takes a more nuanced approach, deriving context window limits based on the specific Claude model variant. For Claude-Fable-5, it sets a 128,000-token ceiling:

```python

# harness/adapters/anthropic.py, lines 53-55

max_tokens = self._get_model_context_window(self.model)  # e.g., 128_000 for Claude-Fable-5

response = client.messages.create(
    ...,
    max_tokens=max_tokens,
)

```

This model-aware mapping prevents the "one size fits all" problem that can waste tokens on smaller-context models or cause hard failures on larger ones.

## Layer 2: Runtime Token Accounting

While adapters prevent per-turn excesses, the **agent loop** tracks cumulative usage across multi-turn conversations. Every `ModelResponse` carries precise token metadata, which the loop aggregates into running totals.

In [`harness/agent_loop.py`](https://github.com/harveyai/harvey-labs/blob/main/harness/agent_loop.py), lines 76-78 update counters after each model invocation:

```python

# harness/agent_loop.py

total_input_tokens += response.input_tokens
total_output_tokens += response.output_tokens

```

These counters surface in the final result dictionary, enabling **cost tracking, debugging, and prompt optimization**. Developers can inspect exactly where token budgets are spent and trim accordingly.

## Layer 3: Context-Window Overflow Detection

When accumulated conversation history exceeds a provider's hard context limit, the adapter raises an exception. The agent loop catches these provider-specific error messages and converts them into structured failure states.

The detection logic in [`harness/agent_loop.py`](https://github.com/harveyai/harvey-labs/blob/main/harness/agent_loop.py), lines 70-73, identifies two key error signatures:

- `"prompt is too long"` — generic length rejection
- `"context_length_exceeded"` — explicit window overflow

```python

# harness/adapters/base.py defines ModelResponse with usage metadata

@dataclass
class ModelResponse:
    content: str
    input_tokens: int
    output_tokens: int
    stop_reason: Optional[str] = None

```

When caught, the loop sets `context_overflow = True`, aborts further turns, and surfaces the condition in the final result. This prevents **silent truncation** that could corrupt outputs or trigger cascading errors in downstream processing.

## Judge-Side Truncation Guard

A fourth safeguard operates in the evaluation layer. When "judge" models (used for quality scoring or validation) hit their own token limits, the system checks `stop_reason == "max_tokens"` in [`evaluation/judge.py`](https://github.com/harveyai/harvey-labs/blob/main/evaluation/judge.py). This prevents hidden test failures where a truncated judgment would be misinterpreted as a valid conclusion.

## Practical Implementation Examples

### Setting Custom Token Limits

Override defaults when instantiating adapters for cost-sensitive or latency-critical workloads:

```python
from harness.adapters.openai import OpenAIAdapter

# Cap output at 32,000 tokens for faster, cheaper responses

adapter = OpenAIAdapter(
    model="gpt-5.0",
    max_tokens=32_000  # Down from 128,000 default

)

result = run_agent(
    adapter,
    system_prompt="You are a contract analysis assistant.",
    user_prompt="Review this 500-page merger agreement...",
    tool_executor=my_tool_executor,
)

print(f"Tokens: {result['input_tokens']} in, {result['output_tokens']} out")

```

### Checking for Overflow in Results

Always inspect the overflow flag before using agent outputs:

```python
result = run_agent(adapter, ...)

if result.get("context_overflow"):
    print("⚠️ Context window overflow — prompt history exceeded model limits")
    # Trigger prompt compression or conversation summarization

else:
    process_output(result["content"])

```

### Monitoring Cumulative Usage

For long-running sessions, track trends across multiple runs:

```python
session_tokens = {"input": 0, "output": 0}

for task in task_queue:
    result = run_agent(adapter, user_prompt=task.description)
    session_tokens["input"] += result["input_tokens"]
    session_tokens["output"] += result["output_tokens"]
    
    if session_tokens["input"] > 800_000:
        print("Approaching context limit — forcing conversation reset")
        adapter.reset_conversation()

```

## Key Source Files

| File | Responsibility |
|------|----------------|
| [`harness/adapters/openai.py`](https://github.com/harveyai/harvey-labs/blob/main/harness/adapters/openai.py) | OpenAI-specific token capping via `max_output_tokens` |
| [`harness/adapters/anthropic.py`](https://github.com/harveyai/harvey-labs/blob/main/harness/adapters/anthropic.py) | Model-aware context window derivation |
| [`harness/adapters/base.py`](https://github.com/harveyai/harvey-labs/blob/main/harness/adapters/base.py) | `ModelResponse` dataclass with token metadata |
| [`harness/agent_loop.py`](https://github.com/harveyai/harvey-labs/blob/main/harness/agent_loop.py) | Overflow detection, aggregation of `total_input_tokens`/`total_output_tokens` |
| [`evaluation/judge.py`](https://github.com/harveyai/harvey-labs/blob/main/evaluation/judge.py) | Truncation detection via `stop_reason` analysis |

## Summary

Harvey AI's approach to **context window overflow and token limit exceedances** rests on three coordinated mechanisms:

- **Adapter-level caps** enforce provider-specific hard limits before any API call
- **Runtime accounting** exposes cumulative usage for cost control and debugging
- **Graceful overflow handling** converts provider errors into structured state rather than crashes

These layers work together to prevent silent failures, enable transparent monitoring, and maintain system reliability under unpredictable load conditions.

## Frequently Asked Questions

### What happens when a prompt exceeds the context window limit?

The agent loop catches the provider's error (either `"prompt is too long"` or `"context_length_exceeded"`), sets `context_overflow = True` in the result object, and halts further interaction. The caller receives a structured failure state rather than an unhandled exception or silently truncated output.

### How does Harvey AI track token usage across multiple turns?

After each model call, [`harness/agent_loop.py`](https://github.com/harveyai/harvey-labs/blob/main/harness/agent_loop.py) accumulates `input_tokens` and `output_tokens` from the `ModelResponse` object into running totals (`total_input_tokens` and `total_output_tokens`). These values appear in the final result dictionary for post-run analysis.

### Can I customize token limits for different models?

Yes. Both `OpenAIAdapter` and `AnthropicAdapter` accept configuration parameters — `max_tokens` for OpenAI, model-specific context windows for Anthropic — allowing per-deployment tuning based on cost, latency, or output length requirements.

### What's the difference between max_tokens in adapters and context overflow detection?

`max_tokens` (or `max_output_tokens`) limits how many tokens the **model generates** in a single response. Context overflow detection catches when the **input prompt** exceeds the model's total context window. These are complementary: one prevents runaway generation, the other handles accumulated conversation history that grows too large.