How the clarify_with_user Phase Works in Open Deep Research

The clarify_with_user phase serves as the mandatory entry gate in the Open Deep Research LangGraph workflow, triggering on every user request to determine—via structured LLM inference—whether sufficient context exists to begin research or if disambiguating questions must be posed before proceeding.

The clarify_with_user phase operates as the first node in the langchain-ai/open_deep_research pipeline, acting as an intelligent router that prevents ambiguous queries from entering expensive research loops. When enabled, this phase analyzes the conversation history using a typed Pydantic schema to decide between immediate research initiation and targeted user clarification. Developers configuring self-hosted instances or customizing the research agent must understand this gatekeeper logic to control latency and user experience.

When the clarify_with_user Phase Is Triggered

The phase triggers automatically at the start of every workflow invocation unless explicitly disabled via configuration. In src/open_deep_research/deep_researcher.py, the graph construction explicitly wires the START edge to the clarification node:

deep_researcher_builder.add_edge(START, "clarify_with_user")

This edge definition (located around line 714 in deep_researcher.py) ensures that clarify_with_user is the first executable node when the compiled graph receives a new AgentState. The phase runs regardless of query complexity—whether the user submits a vague request like "research AI" or a specific prompt—because the LLM evaluation determines subsequent routing.

Configuration-Based Short-Circuiting

While the phase triggers by default, you can bypass it entirely using the allow_clarification configuration flag. Inside the clarify_with_user function (lines 60–78), the code inspects RunnableConfig to check this boolean:

if not config.allow_clarification:
    return Command(goto="write_research_brief")

When allow_clarification is set to False—via environment variables or the Configuration schema—the node immediately returns a Command object directing the graph to skip ahead to the "write_research_brief" node. This short-circuit eliminates the LLM latency associated with clarification checks, making it ideal for automated pipelines where queries are pre-validated.

Prompt Engineering and Context Injection

When clarification is enabled, the phase constructs a specialized prompt using the clarify_with_user_instructions template defined in src/open_deep_research/prompts.py (lines 3–41). The function injects two critical context variables:

  • messages: The full conversation history formatted via get_buffer_string(state["messages"])
  • date: The current date string from get_today_str() to ensure temporal relevance

This prompt is then passed to the configurable_model—initialized at the top of deep_researcher.py—which is wrapped with three production-ready modifiers:

  1. with_structured_output(ClarifyWithUser): Enforces JSON schema compliance using the Pydantic model defined in src/open_deep_research/state.py (lines 30–41)
  2. with_retry(stop_after_attempt=...): Implements automatic retries based on configuration settings
  3. with_config(model_config): Applies the selected research model, token limits, and API credentials

The model executes asynchronously via await clarification_model.ainvoke([HumanMessage(content=prompt_content)]).

Structured Decision Logic and State Transitions

The ClarifyWithUser schema returned by the LLM contains three fields that drive deterministic branching logic (lines 104–115 in deep_researcher.py):

  • need_clarification (bool): The primary decision flag
  • question (str): The clarifying question presented to the user when needed
  • verification (str): The confirmation message displayed when proceeding directly to research

The node implements conditional routing based on the need_clarification value:

If clarification is required:

return Command(
    goto=END,
    update={"messages": [AIMessage(content=response.question)]},
)

This halts the graph at the END node, surfacing the clarifying question as an AIMessage in the conversation state. The workflow pauses until the user responds, creating a human-in-the-loop checkpoint.

If clarification is unnecessary:

return Command(
    goto="write_research_brief",
    update={"messages": [AIMessage(content=response.verification)]},
)

This transitions immediately to the research brief generation phase, including a verification message in the message history to indicate that the research process has commenced.

Summary

  • The clarify_with_user phase is the mandatory entry node in the Open Deep Research LangGraph workflow, triggered by the START → "clarify_with_user" edge in deep_researcher.py.
  • Set allow_clarification to False in the configuration to skip this phase and route directly to "write_research_brief", eliminating the associated LLM latency.
  • The phase uses the clarify_with_user_instructions template from prompts.py, injecting conversation history and current date context before calling the model.
  • A structured output schema (ClarifyWithUser from state.py) enforces type-safe responses containing need_clarification, question, and verification fields.
  • The node either halts at END with a clarifying question or proceeds to "write_research_brief" based on the boolean decision flag, updating the message state accordingly in both paths.

Frequently Asked Questions

Can I completely disable the clarify_with_user phase in Open Deep Research?

Yes. Set the allow_clarification configuration parameter to False in your RunnableConfig or environment settings. When disabled, the clarify_with_user node immediately returns Command(goto="write_research_brief") without invoking the LLM, effectively removing the phase from the execution path while maintaining graph structure integrity.

What Pydantic model structures the output of the clarify_with_user phase?

The phase uses the ClarifyWithUser model defined in src/open_deep_research/state.py (lines 30–41). This schema requires three fields: a boolean need_clarification, a string question (populated when clarification is needed), and a string verification (populated when proceeding directly to research). The model is passed to with_structured_output() to enforce JSON schema compliance from the LLM.

How does the clarify_with_user phase handle conversation history?

The phase retrieves the full message buffer using get_buffer_string(state["messages"]) and injects it into the clarify_with_user_instructions template alongside the current date from get_today_str(). This historical context allows the LLM to detect ambiguities, references to previous messages, or sufficient detail in multi-turn conversations before deciding whether to request clarification.

What happens if the structured LLM call fails during the clarify_with_user phase?

The code wraps the model invocation with with_retry(stop_after_attempt=...), which automatically retries the call according to the retry configuration specified in model_config. If all retry attempts exhaust, the underlying LangChain exception propagates upward, halting the workflow and requiring error handling at the application layer or configuration adjustment of retry parameters.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →