DeepSeek-Reasonix Multi-Model Executor and Planner Collaboration Architecture Explained
DeepSeek-Reasonix coordinates parallel planner and executor language model instances through a Coordinator component that maintains separate, cache-stable sessions while enabling structured plan hand-offs between models.
DeepSeek-Reasonix implements a sophisticated multi-model executor and planner collaboration architecture that runs two distinct language model instances in parallel to optimize reasoning and execution. The system separates high-level planning from tool invocation by isolating each model in its own session while allowing structured information exchange through a central Coordinator. This design, as implemented in the esengine/DeepSeek-Reasonix repository, maximizes cache efficiency and safety by ensuring that neither model's prompt prefix pollutes the other's context.
Core Architecture Components
The collaboration architecture centers on three distinct entities working in concert: the planner for strategy, the executor for action, and the Coordinator for orchestration.
The Planner Model
The planner operates as a low-frequency model instance responsible for generating structured execution plans. According to internal/agent/planner_registry.go, the system creates a dedicated plannerAgent with its own session, system prompt, and prepend-only prefix that preserves cache stability. The planner accesses only read-only research tools—such as search and memory lookup—to construct plans without triggering side effects.
The planner outputs a structured plan containing ordered steps, verification criteria, and optional synthetic turns. This output adheres to a defined schema that the Coordinator passes to the executor for validation.
The Executor Model
The executor functions as a high-frequency model instance that consumes plans and performs actual tool invocations. Defined in the routing logic of internal/agent/planner_route.go, the executor maintains its own executorAgent session completely distinct from the planner's context. Unlike the planner, the executor can invoke mutable tools including file operations, memory writes, and external API calls.
Before execution, the executor validates incoming plans by checking tool availability and schema compliance. If validation fails, the executor rejects the plan, which may trigger a user approval flow rather than autonomous execution.
The Coordinator Component
The Coordinator serves as the orchestration layer that instantiates and manages both models simultaneously. As implemented in internal/control/turn_orchestrator.go, the Coordinator maintains two separate Agent instances and handles the deterministic hand-off between planning and execution phases. It ensures that conversation turns from the planner never append to the executor's message history, preventing context contamination and preserving each model's prefix cache.
Coordination Flow and Session Management
The collaboration follows a strict pipeline that activates when configuration demands multi-model operation.
Two-Model Mode Activation
The system inspects the agent.planner_model configuration parameter to determine whether to activate two-model mode. When the planner model differs from the executor model, the Controller instantiates the Coordinator before processing user input. This initialization creates isolated contexts for both models while establishing the communication protocol for plan exchange.
Plan Generation and Validation
During the planner turn, the model receives the original user prompt plus trusted turn metadata. It performs read-only tool calls to gather context, then produces a structured plan. The host applies deterministic policy rules—documented in docs/SPEC.md#35-two-model-collaboration-coordinator—to classify the plan type as executor-only, light planning, full planning, or plan-for-approval.
The executor then receives this plan as structured text and performs validation checks to ensure required tools are available and the schema matches expected formats. Upon successful validation, the executor proceeds with tool invocation and result aggregation.
Execution and Fallback Handling
If the planner exhausts its token budget without completing a plan, the Coordinator falls back to a pure-executor run, bypassing the planning phase entirely. Conversely, if the executor rejects a plan due to validation failures, execution halts pending user intervention or plan regeneration. This bidirectional fallback mechanism ensures robust task completion while maintaining strict safety boundaries.
Why Separate Sessions Matter
The architectural decision to isolate planner and executor sessions provides three critical operational advantages.
Cache Efficiency: By keeping each model's system prompt and prefix immutable and prepend-only, the backend can reuse previous KV-cache completions effectively. Neither session's cache invalidates when the other model generates tokens, significantly reducing inference latency.
Safety Isolation: The planner's read-only tool restriction prevents unintended side effects during the planning phase. Only the executor can perform mutating operations, and only after explicit plan validation, creating a deterministic security boundary.
Policy Flexibility: The host can dynamically switch between planner-first, executor-only, or mixed collaboration modes without restarting the conversation context. This allows runtime adaptation based on task complexity, cost constraints, or user preferences.
Implementation Details
The multi-model collaboration logic resides in specific source files within the repository:
internal/agent/planner_route.go— Routes turns to the planner when configuration requests planninginternal/agent/planner_registry.go— Registers available planner models and instantiates the plannerAgentinternal/control/turn_orchestrator.go— Orchestrates hand-offs between planner and executor, managing readiness states and progress reportingdocs/SPEC.md#35-two-model-collaboration-coordinator— Formal specification documenting policy decisions, hand-off formats, and session isolation requirements
Summary
- DeepSeek-Reasonix employs a Coordinator to manage parallel planner and executor language model instances in isolated, cache-stable sessions.
- The planner generates structured execution plans using read-only tools, while the executor validates and executes plans using mutable tools.
- Separate sessions maintain cache efficiency by preserving immutable prompt prefixes for each model, preventing context contamination.
- Fallback mechanisms handle planner budget exhaustion and executor plan rejection gracefully, ensuring task completion.
- Implementation spans
internal/agentrouting logic andinternal/control/turn_orchestrator.go, with specifications documented indocs/SPEC.md.
Frequently Asked Questions
How does the Coordinator prevent context contamination between models?
The Coordinator maintains physically separate sessions for the planner and executor, each with distinct Agent instances. By ensuring that the planner's turns never append to the executor's prompt prefix (and vice versa), the system prevents context pollution. This isolation allows each model's prefix cache to remain stable, as implemented in internal/control/turn_orchestrator.go.
What happens when the planner cannot complete a plan within its token budget?
When the planner exhausts its allocated budget without producing a valid plan, the Coordinator triggers a fallback to pure-executor mode. In this scenario, the system bypasses the planning phase entirely and routes the user request directly to the executor model, ensuring the task still completes without planner intervention.
Can the executor reject plans, and what triggers rejection?
Yes, the executor validates all incoming plans before execution. It rejects plans that specify unavailable tools, violate schema constraints, or exceed safety boundaries. Upon rejection, the executor does not proceed with tool invocation; instead, the system may prompt the user for approval or request plan regeneration, depending on the host policy configuration defined in docs/SPEC.md.
Why use a low-frequency model for planning and high-frequency for execution?
The architecture assigns low-frequency, cost-efficient models to planning tasks because plan generation requires reasoning but minimal tool invocation. High-frequency, capable models handle execution because they perform the actual work of calling external APIs, writing files, and processing complex outputs. This separation optimizes cost while ensuring execution quality.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →