How Token Limit Exceeded Handling Works in Final Report Generation
The open_deep_research framework implements an adaptive retry mechanism that detects context-window violations across OpenAI, Anthropic, and Gemini providers, then automatically reduces the max_tokens parameter by 20% until the final report fits within the model's limits or exhausts retries.
When generating comprehensive final reports in the langchain-ai/open_deep_research repository, research contexts often exceed LLM token boundaries. The system’s token limit exceeded handling prevents workflow crashes by intercepting provider-specific errors, recalibrating request sizes, and gracefully degrading when physical limits cannot be satisfied.
Detecting Token Limit Errors Across Providers
The detection logic resides in src/open_deep_research/utils.py within the is_token_limit_exceeded function (lines 665-703). This utility inspects exception messages for provider-specific patterns to determine whether a failure stems from context-window constraints rather than authentication or content-policy violations.
Provider-Specific Pattern Matching
The function aggregates checks for OpenAI, Anthropic, and Gemini error signatures:
def is_token_limit_exceeded(exception: Exception, model_name: str = None) -> bool:
error_str = str(exception).lower()
# Provider‑specific checks
return (
_check_openai_token_limit(exception, error_str) or
_check_anthropic_token_limit(exception, error_str) or
_check_gemini_token_limit(exception, error_str)
)
Each helper (_check_openai_token_limit, _check_anthropic_token_limit, _check_gemini_token_limit) parses the exception string for keywords like "maximum context length", "token limit", or "too many tokens" depending on the provider’s error format. This abstraction allows the final report node to handle token errors generically without provider-specific branching logic.
Adaptive Retry Logic in Final Report Generation
The final_report_generation node in src/open_deep_research/deep_researcher.py (lines 607-718) orchestrates the actual report creation and implements the recovery loop. It combines the detection utility with dynamic configuration adjustment.
Model Invocation and Exception Handling
The node wraps the asynchronous model call in a try/except block, using prompts defined in prompts.py and configuration from configuration.py:
try:
final_report = await configurable_model.with_config(writer_model_config).ainvoke([
HumanMessage(content=final_report_prompt)
])
state["final_report"] = final_report.content
except Exception as e:
if is_token_limit_exceeded(e, configurable.final_report_model):
# Adaptive retry logic triggers here
model_token_limit = get_model_token_limit(configurable.final_report_model)
if model_token_limit:
# Reduce by 20% and retry
writer_model_config["max_tokens"] = int(model_token_limit * 0.8)
# Retry mechanism continues...
else:
state["final_report"] = (
"Error generating final report: Token limit exceeded, however, "
"we could not determine the model's maximum context length. Please "
"update the model map in deep_researcher/utils.py with this information."
)
else:
state["final_report"] = f"Error generating final report: {e}"
Dynamic Token Reduction Strategy
When is_token_limit_exceeded returns True, the system queries get_model_token_limit to retrieve the model’s maximum capacity from an internal mapping. It then recalculates writer_model_config["max_tokens"] to 80% of the original limit, effectively shrinking the generation window to fit within the context boundary. The node retries the invocation with this reduced allocation, repeating the cycle if subsequent attempts still exceed limits until either success or the maximum retry count is reached.
Configuration and Error Messaging
The configuration.py file defines tunable parameters such as final_report_model_max_tokens, which influences the initial request size. If the system cannot resolve the model’s token limit from its internal map, or if all retry attempts fail, it stores explicit error messages in the state graph:
- Undetermined limit:
"Error generating final report: Token limit exceeded, however, we could not determine the model's maximum context length..." - Exhausted retries:
"Error generating final report: Maximum retries exceeded"
These messages propagate to the user interface, providing actionable feedback rather than raw stack traces.
Summary
- Provider-agnostic detection: The
is_token_limit_exceededutility inutils.pynormalizes token errors from OpenAI, Anthropic, and Gemini into a single boolean check. - Automatic throttling: The
final_report_generationnode reducesmax_tokensby 20% of the model’s total capacity upon detecting a limit violation, then retries. - Graceful degradation: When limits are unknown or retries exhaust, the system stores descriptive error messages instead of crashing the research workflow.
- Configurable boundaries: Token limits and model parameters are defined in
configuration.pyand referenced dynamically during the retry loop.
Frequently Asked Questions
How does the system distinguish token limit errors from other LLM failures?
The is_token_limit_exceeded function inspects the exception string for provider-specific keywords. It delegates to internal helpers like _check_openai_token_limit and _check_anthropic_token_limit that match against known error substrings, returning True only when the error explicitly indicates context-length or token-count violations.
What happens if the model's token limit is not mapped in the utilities?
If get_model_token_limit returns None because the model name is absent from the internal mapping, the system stores an error message instructing the user to update the model map in utils.py. It does not attempt blind retries without knowing the boundary constraints.
Can the retry reduction percentage be configured?
The source code implements a fixed 20% reduction (80% of the original limit) calculated as int(model_token_limit * 0.8). This value is hardcoded in the retry logic within deep_researcher.py and is not currently exposed as a user-configurable parameter in configuration.py.
Where is the final report prompt defined when retries occur?
The prompt template used for all attempts—including retries—is defined in src/open_deep_research/prompts.py as final_report_generation_prompt. The node passes this prompt unchanged to the model; only the max_tokens parameter in writer_model_config varies between attempts.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →