# How Multi-Agent Orchestration Works in OpenMAIC: A Technical Deep Dive

> Learn how OpenMAIC implements multi-agent orchestration using a stateless LangGraph state machine. Discover how it routes requests to specialized agents and streams structured JSON output for efficient UI events.

- Repository: [MAIC/OpenMAIC](https://github.com/THU-MAIC/OpenMAIC)
- Tags: deep-dive
- Published: 2026-09-11

---

**OpenMAIC implements multi-agent orchestration through a stateless, single-turn LangGraph state machine where a director node routes requests to specialized LLM-backed agents, streaming structured JSON output parsed into UI events without server-side session persistence.**

The OpenMAIC repository provides a lightweight framework for coordinating multiple AI personas in educational contexts. This article examines the TypeScript implementation that powers the system's collaborative AI teaching experience, focusing on how the orchestration layer manages agent selection, state transitions, and streaming responses.

## Core Architecture

The orchestration system centers on two primary files: **[`lib/orchestration/director-graph.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/lib/orchestration/director-graph.ts)** defines the state machine and routing logic, while **[`lib/orchestration/stateless-generate.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/lib/orchestration/stateless-generate.ts)** handles the streaming execution and output parsing.

### State Definition

All execution context lives in **`OrchestratorState`** (lines 48‑76 of [`director-graph.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/director-graph.ts)). This interface declares both immutable inputs—such as `messages`, `agentIds`, and `model`—and mutable fields that nodes update during execution:

- `currentAgentId`: Tracks which agent is currently active
- `turnCount`: Monitors conversation depth for the current request
- `agentResponses`: Accumulates outputs from completed agents
- `whiteboardLedger`: Stores collaborative actions for the shared workspace
- `shouldEnd`: Boolean flag indicating conversation termination
- `totalActions`: Counter for executed whiteboard operations

This state object passes between graph nodes as a typed annotation, ensuring type safety throughout the execution cycle.

### Graph Topology

The system constructs a directed acyclic graph following the pattern **START → director → (agent_generate) → END**. The `createOrchestrationGraph` function (lines 70‑95) assembles this pipeline, enforcing a strict single-turn constraint. Each HTTP request executes exactly one director-to-agent cycle, preventing infinite loops on the server. The client drives multi-turn conversations by sending successive requests with updated state, keeping the backend truly stateless.

## The Director Node

The **`directorNode`** function (lines 92‑123 and 134‑222) serves as the central router, determining which agent speaks next or whether to prompt the user.

### Single versus Multi-Agent Logic

The director implements conditional routing based on agent count and trigger parameters. In **single-agent scenarios**, the code follows a pure deterministic path: on turn 0 it dispatches the sole agent, and on subsequent turns it cues the user to respond. 

For **multi-agent scenarios**, the system supports a fast-path optimization. When a `triggerAgentId` is supplied and the turn count equals zero, the director skips LLM inference entirely and immediately dispatches the specified agent. This avoids unnecessary latency for predetermined opening moves.

### LLM Decision Making

When routing requires intelligence, the director builds a structured prompt via `buildDirectorPrompt`, injecting the conversation summary, whiteboard ledger, and available agent configurations. It calls the language model through `AISdkLangGraphAdapter`, then parses the response using `parseDirectorDecision` to extract:

- `nextAgentId`: The UUID of the selected agent
- `shouldEnd`: Termination signal for the conversation
- "USER" cue: Instruction to yield control to the human participant

This decision flow ensures the director can dynamically adapt conversation flow based on content, context, and pedagogical goals.

## Agent Generation and Parsing

Once the director selects an agent, the system transitions to the generation phase.

### Streaming Agent Responses

The **`runAgentGeneration`** function (lines 44‑124) handles the actual LLM interaction. It constructs a system prompt using `buildStructuredPrompt`, converts the message history to OpenAI format via `convertMessagesToOpenAI`, and initiates streaming through `adapter.streamGenerate`. The raw delta chunks flow into the structured output parser rather than passing directly to the client.

### Structured Output Parser

Located in **[`lib/orchestration/stateless-generate.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/lib/orchestration/stateless-generate.ts)**, the **`parseStructuredChunk`** function (lines 36‑55) processes partial JSON arrays emitted by constraint-guided LLM outputs. The parser accumulates incomplete chunks and emits concrete events only for well-formed objects matching `{type:"text", …}` or `{type:"action", …}` schemas.

The implementation guarantees event ordering through the `looksLikeStructuredFragment` and `finalizeParser` utilities, ensuring stray JSON fragments or malformed deltas never reach the UI layer. This allows agents to interleave natural language explanations with executable whiteboard actions in a single streaming response.

## Stateless Execution Flow

The **`statelessGenerate`** function (lines 92‑106 and 124‑131) serves as the primary entry point. It instantiates the graph via `createOrchestrationGraph`, initializes state through `buildInitialState`, and iterates over the async generator, yielding **`StatelessEvent`** objects to the caller:

- `agent_start`: Signals agent initialization with metadata
- `text_delta`: Streams partial natural language content
- `action`: Dispatches whiteboard operations like spotlight or slide navigation
- `agent_end`: Marks completion of the current agent's turn
- `done`: Terminates the request, returning `directorState` for the next turn

Between events, the function aggregates metadata including total action counts and content previews, updating the director state that clients must return in subsequent requests.

## Agent Registry and Resolution

Agent definitions live in **`useAgentRegistry`**, imported from `lib/orchestration/registry/store`. The **`resolveAgent`** function (lines 81‑87) implements a hierarchical lookup strategy: it first checks request-scoped overrides (agents generated dynamically for the current session), then falls back to the global registry. This dual-tier approach maintains the server's stateless nature while supporting dynamic agent creation.

## Implementation Example

The following pattern demonstrates integrating the orchestration layer into a client application:

```typescript
import { statelessGenerate } from '@/lib/orchestration/stateless-generate'
import { useAgentRegistry } from '@/lib/orchestration/registry/store'
import { AIModel } from '@/lib/types/provider'

// Construct the request payload
const request = {
  messages: [{ role: 'user', content: 'Explain quantum tunneling' }],
  storeState: { /* current slide / whiteboard state */ },
  config: {
    agentIds: ['teacher', 'assistant'],
    triggerAgentId: 'teacher',
    // Optional: inject per-request generated agents here
  },
  directorState: undefined, // Resumes previous context when provided
  userProfile: { nickname: 'Alice' },
}

// Initialize the model adapter
const model: AIModel = /* e.g., new OpenAI({ apiKey, model }) */

// Execute the orchestration loop
for await (const ev of statelessGenerate(request, abortSignal, model)) {
  switch (ev.type) {
    case 'agent_start':
      console.log(`🟢 ${ev.data.agentName} started`)
      break
    case 'text_delta':
      appendText(ev.data.content) // Stream to chat UI
      break
    case 'action':
      dispatchWhiteboardAction(ev.data) // Execute slide commands
      break
    case 'cue_user':
      showUserInputBox() // Yield to human
      break
    case 'done':
      // Persist directorState for the next turn
      saveStateForNextRequest(ev.data.directorState)
      break
  }
}

```

After receiving the `done` event, clients should store `ev.data.directorState` and include it in the subsequent request's `directorState` field to maintain conversation continuity across HTTP boundaries.

## Summary

- **Stateless Architecture**: OpenMAIC uses single-turn LangGraph execution where the client manages session state, eliminating server-side memory requirements.
- **Director Pattern**: The `directorNode` routes traffic between agents using either deterministic logic or LLM-based decision making, supporting both single and multi-agent configurations.
- **Structured Streaming**: The `parseStructuredChunk` parser converts raw LLM deltas into typed events, separating text content from executable actions.
- **Registry Flexibility**: The `resolveAgent` hierarchy allows dynamic agent injection while maintaining global definitions in `useAgentRegistry`.
- **Event-Driven API**: The `statelessGenerate` generator yields discrete event types that map directly to UI updates and whiteboard operations.

## Frequently Asked Questions

### How does OpenMAIC maintain conversation state across multiple turns?

The system remains stateless by requiring the client to pass `directorState` (received from the previous turn's `done` event) into the next request. This object contains the `turnCount`, `whiteboardLedger`, and accumulated metadata, allowing the `directorNode` to resume context without server-side session storage.

### What determines which agent speaks next in a multi-agent scenario?

The `directorNode` evaluates three conditions in order: first, it checks for a `triggerAgentId` on turn 0 (fast-path); second, it queries the LLM using `buildDirectorPrompt` and `parseDirectorDecision` to analyze conversation context; third, if no agents remain relevant, it returns a "USER" cue to prompt human input.

### Can agents execute actions while speaking, or only after completing text generation?

Agents stream actions interleaved with text through **structured JSON output**. The `parseStructuredChunk` parser processes partial arrays in real-time, emitting `text_delta` events for narrative content and `action` events for whiteboard operations as soon as complete JSON objects arrive, rather than waiting for the full response.

### Is it possible to add temporary agents for a single conversation?

Yes. While `useAgentRegistry` holds persistent agent definitions, the `resolveAgent` function checks request-scoped overrides before consulting the global registry. You can inject dynamically generated agent configurations into the request's `config` object, and they will take precedence for that specific turn.