# How Token Limit Exceeded Handling Works in Final Report Generation

> Learn how the open deep research framework handles token limit exceeded errors by automatically reducing max_tokens for seamless final report generation across top AI providers.

- Repository: [LangChain/open_deep_research](https://github.com/langchain-ai/open_deep_research)
- Tags: deep-dive
- Published: 2026-07-23

---

**The `open_deep_research` framework implements an adaptive retry mechanism that detects context-window violations across OpenAI, Anthropic, and Gemini providers, then automatically reduces the `max_tokens` parameter by 20% until the final report fits within the model's limits or exhausts retries.**

When generating comprehensive final reports in the `langchain-ai/open_deep_research` repository, research contexts often exceed LLM token boundaries. The system’s **token limit exceeded handling** prevents workflow crashes by intercepting provider-specific errors, recalibrating request sizes, and gracefully degrading when physical limits cannot be satisfied.

## Detecting Token Limit Errors Across Providers

The detection logic resides in [`src/open_deep_research/utils.py`](https://github.com/langchain-ai/open_deep_research/blob/main/src/open_deep_research/utils.py) within the `is_token_limit_exceeded` function (lines 665-703). This utility inspects exception messages for provider-specific patterns to determine whether a failure stems from context-window constraints rather than authentication or content-policy violations.

### Provider-Specific Pattern Matching

The function aggregates checks for OpenAI, Anthropic, and Gemini error signatures:

```python
def is_token_limit_exceeded(exception: Exception, model_name: str = None) -> bool:
    error_str = str(exception).lower()
    # Provider‑specific checks

    return (
        _check_openai_token_limit(exception, error_str) or
        _check_anthropic_token_limit(exception, error_str) or
        _check_gemini_token_limit(exception, error_str)
    )

```

Each helper (`_check_openai_token_limit`, `_check_anthropic_token_limit`, `_check_gemini_token_limit`) parses the exception string for keywords like "maximum context length", "token limit", or "too many tokens" depending on the provider’s error format. This abstraction allows the final report node to handle token errors generically without provider-specific branching logic.

## Adaptive Retry Logic in Final Report Generation

The `final_report_generation` node in [`src/open_deep_research/deep_researcher.py`](https://github.com/langchain-ai/open_deep_research/blob/main/src/open_deep_research/deep_researcher.py) (lines 607-718) orchestrates the actual report creation and implements the recovery loop. It combines the detection utility with dynamic configuration adjustment.

### Model Invocation and Exception Handling

The node wraps the asynchronous model call in a `try/except` block, using prompts defined in [`prompts.py`](https://github.com/langchain-ai/open_deep_research/blob/main/prompts.py) and configuration from [`configuration.py`](https://github.com/langchain-ai/open_deep_research/blob/main/configuration.py):

```python
try:
    final_report = await configurable_model.with_config(writer_model_config).ainvoke([
        HumanMessage(content=final_report_prompt)
    ])
    state["final_report"] = final_report.content
except Exception as e:
    if is_token_limit_exceeded(e, configurable.final_report_model):
        # Adaptive retry logic triggers here

        model_token_limit = get_model_token_limit(configurable.final_report_model)
        if model_token_limit:
            # Reduce by 20% and retry

            writer_model_config["max_tokens"] = int(model_token_limit * 0.8)
            # Retry mechanism continues...

        else:
            state["final_report"] = (
                "Error generating final report: Token limit exceeded, however, "
                "we could not determine the model's maximum context length. Please "
                "update the model map in deep_researcher/utils.py with this information."
            )
    else:
        state["final_report"] = f"Error generating final report: {e}"

```

### Dynamic Token Reduction Strategy

When `is_token_limit_exceeded` returns **True**, the system queries `get_model_token_limit` to retrieve the model’s maximum capacity from an internal mapping. It then recalculates `writer_model_config["max_tokens"]` to 80% of the original limit, effectively shrinking the generation window to fit within the context boundary. The node retries the invocation with this reduced allocation, repeating the cycle if subsequent attempts still exceed limits until either success or the maximum retry count is reached.

## Configuration and Error Messaging

The [`configuration.py`](https://github.com/langchain-ai/open_deep_research/blob/main/configuration.py) file defines tunable parameters such as `final_report_model_max_tokens`, which influences the initial request size. If the system cannot resolve the model’s token limit from its internal map, or if all retry attempts fail, it stores explicit error messages in the state graph:

- **Undetermined limit**: `"Error generating final report: Token limit exceeded, however, we could not determine the model's maximum context length..."`
- **Exhausted retries**: `"Error generating final report: Maximum retries exceeded"`

These messages propagate to the user interface, providing actionable feedback rather than raw stack traces.

## Summary

- **Provider-agnostic detection**: The `is_token_limit_exceeded` utility in [`utils.py`](https://github.com/langchain-ai/open_deep_research/blob/main/utils.py) normalizes token errors from OpenAI, Anthropic, and Gemini into a single boolean check.
- **Automatic throttling**: The `final_report_generation` node reduces `max_tokens` by 20% of the model’s total capacity upon detecting a limit violation, then retries.
- **Graceful degradation**: When limits are unknown or retries exhaust, the system stores descriptive error messages instead of crashing the research workflow.
- **Configurable boundaries**: Token limits and model parameters are defined in [`configuration.py`](https://github.com/langchain-ai/open_deep_research/blob/main/configuration.py) and referenced dynamically during the retry loop.

## Frequently Asked Questions

### How does the system distinguish token limit errors from other LLM failures?

The `is_token_limit_exceeded` function inspects the exception string for provider-specific keywords. It delegates to internal helpers like `_check_openai_token_limit` and `_check_anthropic_token_limit` that match against known error substrings, returning **True** only when the error explicitly indicates context-length or token-count violations.

### What happens if the model's token limit is not mapped in the utilities?

If `get_model_token_limit` returns `None` because the model name is absent from the internal mapping, the system stores an error message instructing the user to update the model map in [`utils.py`](https://github.com/langchain-ai/open_deep_research/blob/main/utils.py). It does not attempt blind retries without knowing the boundary constraints.

### Can the retry reduction percentage be configured?

The source code implements a fixed 20% reduction (80% of the original limit) calculated as `int(model_token_limit * 0.8)`. This value is hardcoded in the retry logic within [`deep_researcher.py`](https://github.com/langchain-ai/open_deep_research/blob/main/deep_researcher.py) and is not currently exposed as a user-configurable parameter in [`configuration.py`](https://github.com/langchain-ai/open_deep_research/blob/main/configuration.py).

### Where is the final report prompt defined when retries occur?

The prompt template used for all attempts—including retries—is defined in [`src/open_deep_research/prompts.py`](https://github.com/langchain-ai/open_deep_research/blob/main/src/open_deep_research/prompts.py) as `final_report_generation_prompt`. The node passes this prompt unchanged to the model; only the `max_tokens` parameter in `writer_model_config` varies between attempts.