# How the OpenMAIC Generation Pipeline Works: From Prompt to Persistent Media Scene

> Explore the OpenMAIC generation pipeline from prompt to media scene. Discover its eight-stage architecture including Next.js, Agent Runtime, and PostgreSQL for efficient media creation.

- Repository: [MAIC/OpenMAIC](https://github.com/THU-MAIC/OpenMAIC)
- Tags: how-to-guide
- Published: 2026-09-08

---

**The OpenMAIC generation pipeline transforms user prompts into rich media scenes through an eight-stage architecture involving Next.js frontend components, a server-side Agent Runtime, provider SDKs, Dexie-backed client stores for streaming results, and PostgreSQL persistence for completed outlines.**

The THU-MAIC/OpenMAIC repository implements a modular generation system that converts text prompts into interactive media containing text, images, and video. This technical deep dive examines the end-to-end flow from UI interaction to database persistence, referencing actual source file paths and implementation details from the codebase.

## Frontend Request Initialization

The pipeline begins when a user submits a prompt through the interactive mode UI. In [`components/generation/interactive-mode-button.tsx`](https://github.com/THU-MAIC/OpenMAIC/blob/main/components/generation/interactive-mode-button.tsx), the component dispatches a `POST` request to the `/api/generation` endpoint with the prompt and selected tool type.

```typescript
// components/generation/interactive-mode-button.tsx
const startGeneration = async (prompt: string) => {
  const res = await fetch('/api/generation', {
    method: 'POST',
    headers: { 'Content-Type': 'application/json' },
    body: JSON.stringify({ prompt, tool: 'image' }),
  });
  if (!res.ok) throw new Error('Generation start failed');
};

```

This request payload includes the **user prompt** and the **generation tool name** (e.g., `text`, `image`, or `video`), which determines which provider SDK the pipeline will invoke later.

## API Route and Runtime Delegation

The request reaches the Next.js API handler defined in [`pages/api/generation.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/pages/api/generation.ts). This handler validates the payload and forwards it to the **Agent Runtime**, which orchestrates the server-side execution.

```typescript
// pages/api/generation.ts
export default async function handler(req: NextApiRequest, res: NextApiResponse) {
  const { prompt, tool } = req.body;
  const jobId = await AgentRuntime.startJob({ prompt, tool });
  res.status(202).json({ jobId });
}

```

The Agent Runtime immediately creates a **GenerationJob** object and returns a `202 Accepted` status with a unique job identifier, allowing the client to proceed while processing occurs asynchronously.

## Agent Runtime and Tool Selection

At the heart of the server-side logic, [`lib/server/agent-runtime/generation-tools.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/lib/server/agent-runtime/generation-tools.ts) maintains the registry of available generation capabilities. The `AgentRuntime` selects the appropriate tool from `GENERATION_TOOL_NAMES` and initializes the job execution context.

```typescript
// lib/server/agent-runtime/generation-tools.ts
export const GENERATION_TOOL_NAMES = ['text', 'image', 'video'] as const;

export async function runGeneration(job: GenerationJob) {
  const tool = generationToolMap[job.tool];
  return await tool.execute(job);
}

```

Each tool in the `generationToolMap` knows how to format prompts for its specific provider and how to handle the resulting media streams.

## Provider SDK Integration and Streaming

Once the tool is selected, the pipeline invokes the underlying provider through the appropriate SDK in `lib/server/providers/`. For example, [`lib/server/providers/openai.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/lib/server/providers/openai.ts) handles GPT-based text generation, while image and video tools interface with Stable Diffusion or similar services.

The provider streams partial results back to the runtime as they are generated. This streaming architecture ensures that large media files or lengthy text generations do not block the pipeline, allowing the system to process chunks incrementally.

## Client-Side State Synchronization

As partial results arrive from the server, the runtime writes them into the **Media Generation Store**, implemented in [`lib/store/media-generation.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/lib/store/media-generation.ts) using Dexie (a IndexedDB wrapper). This client-side persistence layer caches generation chunks locally for fast UI responsiveness and offline resilience.

```typescript
// lib/store/media-generation.ts
export const useMediaGenerationStore = createStore((set, get) => ({
  outlines: [] as Outline[],
  addChunk: (chunk) => set(state => ({
    outlines: [...state.outlines, chunk],
  })),
}));

```

The **Stage Store** (`useStageStore` in [`lib/workbench/use-stage-store.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/lib/workbench/use-stage-store.ts)), a Zustand-based state manager, tracks boolean flags like `generationOpen` and `generationComplete`. Components subscribe to this store to react instantly to state changes without polling the server.

```tsx
// lib/workbench/use-workbench-session.ts
const { outlines } = useMediaGenerationStore();
useEffect(() => {
  if (outlines.length) {
    stageStore.setState({ generationOpen: true });
  }
}, [outlines]);

```

## UI Rendering and Stage Freshness

The workspace UI components in [`lib/workbench/stage-freshness.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/lib/workbench/stage-freshness.ts) listen to the Stage Store updates. When new outlines or scenes appear in the Media Generation Store, the UI re-renders the canvas, displays generated images, and updates the workspace stage.

The `generationOpen` flag signals that a generation is active, triggering loading states and progress indicators, while `generationComplete` enables editing and export functionality once the stream finishes.

## Persistence and Database Storage

Upon completion, the pipeline persists the final outline—including all generated media URLs and metadata—to the server-side **PostgreSQL** database via the `stageOutlinesPut` API. This persistence layer enables resume-on-refresh functionality and supports collaborative editing scenarios where multiple users might access the same scene.

The `scripts/generation-node-smoke.mjs` file demonstrates how a Node.js worker can pull queued jobs, execute the full provider pipeline, and persist results independently of the Next.js request-response cycle, making the architecture suitable for background job processing.

## Post-Generation Workflow Transitions

After the pipeline marks the job as complete, the UI transitions out of generation mode. The [`lib/workbench/pro-playback-exit.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/lib/workbench/pro-playback-exit.ts) module handles the cleanup of temporary generation states and ensures the workspace returns to standard editing mode.

Additional post-generation actions—such as **playback**, **export**, or **scene swapping**—become available through [`lib/workbench/pro-swap.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/lib/workbench/pro-swap.ts), allowing users to iterate on generated content or replace specific media elements without restarting the entire pipeline.

## Implementation Summary

When implementing custom generation tools in OpenMAIC, developers must register new tools in `GENERATION_TOOL_NAMES`, implement the tool's `execute` method to handle provider-specific API calls, and ensure the tool streams results compatible with the `MediaGenerationStore` interface. The modular architecture separates provider logic from UI concerns, making it straightforward to add support for new AI models or media formats while maintaining consistent state management across the application.

## Summary

- **Prompt Entry**: Users initiate generation via [`components/generation/interactive-mode-button.tsx`](https://github.com/THU-MAIC/OpenMAIC/blob/main/components/generation/interactive-mode-button.tsx), which POSTs to `/api/generation` with the prompt and tool type.
- **Server Processing**: [`pages/api/generation.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/pages/api/generation.ts) delegates to `AgentRuntime`, which selects the appropriate tool from [`lib/server/agent-runtime/generation-tools.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/lib/server/agent-runtime/generation-tools.ts) and executes provider calls.
- **Streaming Results**: Partial results flow through `lib/server/providers/` and populate [`lib/store/media-generation.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/lib/store/media-generation.ts) (Dexie/IndexedDB) for immediate client access.
- **UI Updates**: [`lib/workbench/stage-freshness.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/lib/workbench/stage-freshness.ts) and `useStageStore` synchronize generation states (`generationOpen`, `generationComplete`) with the React component tree.
- **Persistence**: Completed outlines save to PostgreSQL via `stageOutlinesPut`, enabling session recovery and collaborative features.

## Frequently Asked Questions

### How does the AgentRuntime determine which generation tool to execute?

According to the OpenMAIC source code in [`lib/server/agent-runtime/generation-tools.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/lib/server/agent-runtime/generation-tools.ts), the runtime references the `GENERATION_TOOL_NAMES` constant and a `generationToolMap` object. When `AgentRuntime.startJob()` receives a job request, it extracts the `tool` property from the job payload and uses this key to retrieve the corresponding tool implementation, which must implement an `execute(job: GenerationJob)` method.

### What mechanism handles real-time updates during media generation?

The pipeline uses a dual-store architecture on the client side. The **Media Generation Store** ([`lib/store/media-generation.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/lib/store/media-generation.ts)) persists streaming chunks to IndexedDB via Dexie, while the **Stage Store** ([`lib/workbench/use-stage-store.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/lib/workbench/use-stage-store.ts)) tracks boolean flags like `generationOpen` and `generationComplete` using Zustand. React components subscribe to these stores to render updates without polling the server.

### Where does OpenMAIC store completed generation results?

Once a generation job finishes, the system persists the complete outline—including media URLs, text content, and metadata—to a **PostgreSQL** database via the `stageOutlinesPut` API endpoint. This server-side storage enables users to refresh their browsers and resume editing without losing generated content, and supports collaborative workflows where multiple users might edit the same scene.

### Can the pipeline handle simultaneous text, image, and video generation?

Yes. The `GENERATION_TOOL_NAMES` array in [`lib/server/agent-runtime/generation-tools.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/lib/server/agent-runtime/generation-tools.ts) defines distinct tools for `text`, `image`, and `video` generation. Each tool operates as an independent module within the Agent Runtime, and the architecture supports concurrent job execution. The `scripts/generation-node-smoke.mjs` example demonstrates how background workers can process these jobs asynchronously outside the main Next.js server thread.