How OpenMAIC's Two-Stage Generation Pipeline Works: Draft-to-Refine Architecture
OpenMAIC employs a two-stage generation pipeline that separates pure language draft creation from tool-augmented refinement, allowing the system to minimize expensive LLM calls while ensuring factual, safely formatted final outputs.
The THU-MAIC/OpenMAIC repository implements this architecture to balance computational efficiency with high-quality response generation. By decoupling initial reasoning from external tool execution, the pipeline keeps inference costs low while enabling complex multi-step workflows involving search, code execution, and content safety checks.
Stage 1: Draft Generation
The first stage generates a pure language-only draft using the primary large language model (e.g., Claude, GPT-4). This initial pass captures the high-level intent, reasoning steps, and potential actions without invoking any external services.
In lib/orchestration/generation.ts, the generateDraft() method handles this phase:
if (stage === 1) {
return await llm.generateDraft(prompt);
}
This draft serves as a computational inexpensive prototype that defines the response structure before committing to costly tool operations.
Stage 2: Tool-Augmented Refinement
The second stage examines the draft for tool-call markers (e.g., <tool>search</tool>) and orchestrates backend service integration. According to the source implementation in lib/orchestration/generation.ts, this phase extracts tool calls, executes them, and performs a second LLM pass to integrate results:
const toolCalls = extractToolCalls(prompt);
const toolResults = await runTools(toolCalls);
return await llm.refineWithTools(prompt, toolResults);
Tool Call Detection and Execution
During refinement, the system parses the draft for specific XML-style markers indicating required external actions. When detected, OpenMAIC invokes appropriate backend services—ranging from web search and code interpreters to image generation APIs. The runTools() function executes these in parallel where possible, feeding structured results back into the context window.
The UI layer tracks this progress through components/workbench/chat/tool-presentation.ts and components/workbench/chat/tool-progress.ts, which render intermediate outputs to users while stage 2 completes.
Safety Guardrails and Final Polish
Before returning the response, stage 2 applies content safety filters including LaTeX safety checks and boundary enforcement. These guardrails ensure that integrated tool results conform to formatting standards and policy restrictions. The refineWithTools() method handles the final rewrite, producing a coherent narrative that weaves together the original draft reasoning with factual tool outputs.
Implementation Deep Dive
Several specialized modules manage the pipeline's state and synchronization:
Stage Freshness Tracking
The lib/workbench/stage-freshness.ts module tracks the validity of cached generation stages, determining when environmental changes or context shifts require a fresh second pass rather than reusing stored results.
TTS Synchronization
The lib/workbench/tts-stage-sync.ts file ensures text-to-speech engines receive the final refined content rather than raw drafts, preventing audio streams from speaking placeholder text or unverified tool results.
Testing Infrastructure
End-to-end validation occurs in e2e/tests/generation-flow.spec.ts, which drives the complete two-stage flow through the UI to verify integration between draft generation, tool execution, and final rendering. Unit tests in tests/orchestration/conversation-summary.test.ts validate the orchestration logic that assembles stage outputs into coherent conversation history.
Code Examples
Client-side generation trigger from components/workbench/chat/compose.tsx:
async function generateAnswer(prompt: string) {
// Stage 1 – pure LLM draft
const draft = await fetch('/api/generate', {
method: 'POST',
body: JSON.stringify({ prompt, stage: 1 })
}).then(r => r.json());
// Stage 2 – tool-augmented refinement
const final = await fetch('/api/generate', {
method: 'POST',
body: JSON.stringify({ draft, stage: 2 })
}).then(r => r.json());
return final.text;
}
Server-side orchestration from lib/orchestration/generation.ts:
export async function generate(prompt: string, stage: 1 | 2) {
if (stage === 1) {
return await llm.generateDraft(prompt);
} else {
const toolCalls = extractToolCalls(prompt);
const toolResults = await runTools(toolCalls);
return await llm.refineWithTools(prompt, toolResults);
}
}
Summary
- The two-stage pipeline separates initial draft creation from tool integration, minimizing unnecessary API costs.
- Stage 1 produces reasoning-heavy language output without external dependencies.
- Stage 2 detects tool markers, executes backend services, and polishes results with safety guardrails.
- Implementation files like
stage-freshness.tsandtts-stage-sync.tsmanage state synchronization and output coherence. - Test suites including
generation-flow.spec.tsensure reliable end-to-end orchestration of both stages.
Frequently Asked Questions
What triggers the second stage of the OpenMAIC generation pipeline?
The second stage triggers when the system detects tool-call markers (e.g., <tool>search</tool>) within the draft output or when the query explicitly requires external data not contained in the model's training parameters. The stage-freshness.ts module also evaluates cache validity to determine if a fresh refinement pass is necessary.
How does OpenMAIC handle safety during the two-stage process?
Safety checks occur during stage 2 refinement rather than stage 1, ensuring that integrated tool outputs undergo LaTeX safety validation and content-boundary enforcement before reaching the user. This placement prevents unsafe intermediate content from polluting the final response while maintaining the draft's reasoning structure.
What types of tools can be invoked during the refinement stage?
The pipeline supports diverse backend services including web search engines, code execution environments, mathematical computors, and image generation APIs. The tool-presentation.ts component renders these results contextually, allowing the refineWithTools() method to incorporate structured data, search snippets, or execution outputs into the final narrative.
How is the two-stage pipeline tested in the OpenMAIC repository?
Testing occurs at multiple levels: conversation-summary.test.ts validates the orchestration logic that stitches stage outputs together, while generation-flow.spec.ts provides end-to-end Playwright tests simulating real user interactions through both draft and refinement phases. These tests verify that tool calls execute correctly and that TTS synchronization properly aligns with the final refined text.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →