How the OpenMAIC Generation Pipeline Works: From Prompt to Persistent Media Scene
The OpenMAIC generation pipeline transforms user prompts into rich media scenes through an eight-stage architecture involving Next.js frontend components, a server-side Agent Runtime, provider SDKs, Dexie-backed client stores for streaming results, and PostgreSQL persistence for completed outlines.
The THU-MAIC/OpenMAIC repository implements a modular generation system that converts text prompts into interactive media containing text, images, and video. This technical deep dive examines the end-to-end flow from UI interaction to database persistence, referencing actual source file paths and implementation details from the codebase.
Frontend Request Initialization
The pipeline begins when a user submits a prompt through the interactive mode UI. In components/generation/interactive-mode-button.tsx, the component dispatches a POST request to the /api/generation endpoint with the prompt and selected tool type.
// components/generation/interactive-mode-button.tsx
const startGeneration = async (prompt: string) => {
const res = await fetch('/api/generation', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ prompt, tool: 'image' }),
});
if (!res.ok) throw new Error('Generation start failed');
};
This request payload includes the user prompt and the generation tool name (e.g., text, image, or video), which determines which provider SDK the pipeline will invoke later.
API Route and Runtime Delegation
The request reaches the Next.js API handler defined in pages/api/generation.ts. This handler validates the payload and forwards it to the Agent Runtime, which orchestrates the server-side execution.
// pages/api/generation.ts
export default async function handler(req: NextApiRequest, res: NextApiResponse) {
const { prompt, tool } = req.body;
const jobId = await AgentRuntime.startJob({ prompt, tool });
res.status(202).json({ jobId });
}
The Agent Runtime immediately creates a GenerationJob object and returns a 202 Accepted status with a unique job identifier, allowing the client to proceed while processing occurs asynchronously.
Agent Runtime and Tool Selection
At the heart of the server-side logic, lib/server/agent-runtime/generation-tools.ts maintains the registry of available generation capabilities. The AgentRuntime selects the appropriate tool from GENERATION_TOOL_NAMES and initializes the job execution context.
// lib/server/agent-runtime/generation-tools.ts
export const GENERATION_TOOL_NAMES = ['text', 'image', 'video'] as const;
export async function runGeneration(job: GenerationJob) {
const tool = generationToolMap[job.tool];
return await tool.execute(job);
}
Each tool in the generationToolMap knows how to format prompts for its specific provider and how to handle the resulting media streams.
Provider SDK Integration and Streaming
Once the tool is selected, the pipeline invokes the underlying provider through the appropriate SDK in lib/server/providers/. For example, lib/server/providers/openai.ts handles GPT-based text generation, while image and video tools interface with Stable Diffusion or similar services.
The provider streams partial results back to the runtime as they are generated. This streaming architecture ensures that large media files or lengthy text generations do not block the pipeline, allowing the system to process chunks incrementally.
Client-Side State Synchronization
As partial results arrive from the server, the runtime writes them into the Media Generation Store, implemented in lib/store/media-generation.ts using Dexie (a IndexedDB wrapper). This client-side persistence layer caches generation chunks locally for fast UI responsiveness and offline resilience.
// lib/store/media-generation.ts
export const useMediaGenerationStore = createStore((set, get) => ({
outlines: [] as Outline[],
addChunk: (chunk) => set(state => ({
outlines: [...state.outlines, chunk],
})),
}));
The Stage Store (useStageStore in lib/workbench/use-stage-store.ts), a Zustand-based state manager, tracks boolean flags like generationOpen and generationComplete. Components subscribe to this store to react instantly to state changes without polling the server.
// lib/workbench/use-workbench-session.ts
const { outlines } = useMediaGenerationStore();
useEffect(() => {
if (outlines.length) {
stageStore.setState({ generationOpen: true });
}
}, [outlines]);
UI Rendering and Stage Freshness
The workspace UI components in lib/workbench/stage-freshness.ts listen to the Stage Store updates. When new outlines or scenes appear in the Media Generation Store, the UI re-renders the canvas, displays generated images, and updates the workspace stage.
The generationOpen flag signals that a generation is active, triggering loading states and progress indicators, while generationComplete enables editing and export functionality once the stream finishes.
Persistence and Database Storage
Upon completion, the pipeline persists the final outline—including all generated media URLs and metadata—to the server-side PostgreSQL database via the stageOutlinesPut API. This persistence layer enables resume-on-refresh functionality and supports collaborative editing scenarios where multiple users might access the same scene.
The scripts/generation-node-smoke.mjs file demonstrates how a Node.js worker can pull queued jobs, execute the full provider pipeline, and persist results independently of the Next.js request-response cycle, making the architecture suitable for background job processing.
Post-Generation Workflow Transitions
After the pipeline marks the job as complete, the UI transitions out of generation mode. The lib/workbench/pro-playback-exit.ts module handles the cleanup of temporary generation states and ensures the workspace returns to standard editing mode.
Additional post-generation actions—such as playback, export, or scene swapping—become available through lib/workbench/pro-swap.ts, allowing users to iterate on generated content or replace specific media elements without restarting the entire pipeline.
Implementation Summary
When implementing custom generation tools in OpenMAIC, developers must register new tools in GENERATION_TOOL_NAMES, implement the tool's execute method to handle provider-specific API calls, and ensure the tool streams results compatible with the MediaGenerationStore interface. The modular architecture separates provider logic from UI concerns, making it straightforward to add support for new AI models or media formats while maintaining consistent state management across the application.
Summary
- Prompt Entry: Users initiate generation via
components/generation/interactive-mode-button.tsx, which POSTs to/api/generationwith the prompt and tool type. - Server Processing:
pages/api/generation.tsdelegates toAgentRuntime, which selects the appropriate tool fromlib/server/agent-runtime/generation-tools.tsand executes provider calls. - Streaming Results: Partial results flow through
lib/server/providers/and populatelib/store/media-generation.ts(Dexie/IndexedDB) for immediate client access. - UI Updates:
lib/workbench/stage-freshness.tsanduseStageStoresynchronize generation states (generationOpen,generationComplete) with the React component tree. - Persistence: Completed outlines save to PostgreSQL via
stageOutlinesPut, enabling session recovery and collaborative features.
Frequently Asked Questions
How does the AgentRuntime determine which generation tool to execute?
According to the OpenMAIC source code in lib/server/agent-runtime/generation-tools.ts, the runtime references the GENERATION_TOOL_NAMES constant and a generationToolMap object. When AgentRuntime.startJob() receives a job request, it extracts the tool property from the job payload and uses this key to retrieve the corresponding tool implementation, which must implement an execute(job: GenerationJob) method.
What mechanism handles real-time updates during media generation?
The pipeline uses a dual-store architecture on the client side. The Media Generation Store (lib/store/media-generation.ts) persists streaming chunks to IndexedDB via Dexie, while the Stage Store (lib/workbench/use-stage-store.ts) tracks boolean flags like generationOpen and generationComplete using Zustand. React components subscribe to these stores to render updates without polling the server.
Where does OpenMAIC store completed generation results?
Once a generation job finishes, the system persists the complete outline—including media URLs, text content, and metadata—to a PostgreSQL database via the stageOutlinesPut API endpoint. This server-side storage enables users to refresh their browsers and resume editing without losing generated content, and supports collaborative workflows where multiple users might edit the same scene.
Can the pipeline handle simultaneous text, image, and video generation?
Yes. The GENERATION_TOOL_NAMES array in lib/server/agent-runtime/generation-tools.ts defines distinct tools for text, image, and video generation. Each tool operates as an independent module within the Agent Runtime, and the architecture supports concurrent job execution. The scripts/generation-node-smoke.mjs example demonstrates how background workers can process these jobs asynchronously outside the main Next.js server thread.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →