# How Does the `generateSceneActions` Function Work in OpenMAIC?

> Discover how OpenMAIC's generateSceneActions function transforms scenes into interactive actions. Learn about its LLM pipeline, retry logic, and validation.

- Repository: [MAIC/OpenMAIC](https://github.com/THU-MAIC/OpenMAIC)
- Tags: deep-dive
- Published: 2026-09-06

---

**`generateSceneActions`** converts a scene outline and generated content into executable actions like speech prompts, interactive widgets, and media cues through a structured LLM pipeline with built-in retry logic and validation.

The `generateSceneActions` function in THU-MAIC/OpenMAIC serves as the final transformation step in the AI-driven content generation pipeline. After scene content has been generated, this function determines *what actions* should accompany that content—whether that's a voiceover reading text aloud, a clickable quiz widget, or a video that auto-plays. Understanding how it orchestrates LLM calls, handles failures, and validates outputs is essential for anyone extending OpenMAIC's generation capabilities.

## Core Function Signature

`generateSceneActions` is exported from **`packages/@openmaic/generation/src/scene-generator.ts`** with the following signature:

```typescript
export async function generateSceneActions(
    outline: SceneOutline,
    content: SceneContent | null,
    aiCall: AICallFn,
    opts?: { logger?: Logger; retryPolicy?: RetryPolicy }
): Promise<SceneAction[]>

```

- **`outline`** – The abstract scene description containing type, title, and required elements
- **`content`** – Rich content from `generateSceneContent` (may be `null` if generation failed)
- **`aiCall`** – Injectable LLM caller function for testability and provider flexibility
- **`opts`** – Optional utilities for logging and retry configuration

## Step-by-Step Execution Flow

### 1. Guard Clause for Missing Content

If `content` is `null`, the function returns the original outline unchanged with no actions attached. This defensive design prevents downstream failures when upstream generation fails:

```typescript
// Simplified from scene-generator.ts
if (!content) {
  logger?.warn(`Skipping actions for ${outline.id}: no content generated`);
  return [];
}

```

### 2. Structured Prompt Construction

The function assembles a two-part prompt:

- **System prompt**: Describes the scene type (slide, quiz, interactive) and action constraints
- **User prompt**: Includes scene metadata, content body, and a strict JSON schema requirement

The prompt specifically requests a `SceneAction[]` array with fields for `type`, `targetId`, and payload data. Optional guidelines prevent redundant widget generation for elements that already have interactive components.

### 3. LLM Invocation with Resilient Retry Logic

```typescript
// Conceptual flow from the implementation
const rawResponse = await withRetry(
  () => aiCall(prompt),
  opts?.retryPolicy ?? defaultRetryPolicy
);

```

The `retryPolicy` defaults to exponential back-off with jitter. On repeated failure, the function logs the error and returns an empty action list rather than crashing the pipeline.

### 4. Response Parsing and Validation

```typescript
const parsed = JSON.parse(rawResponse) as unknown[];
const validActions = parsed
  .filter((item): item is SceneAction => 
    validateSceneAction(item) // checks required fields
  );

```

Invalid entries are filtered with warnings logged. This lenient parsing ensures partial LLM outputs don't corrupt the entire scene.

### 5. Action Post-Processing

Before returning, actions undergo:

- **Deduplication** by `targetId` + `type` combination
- **TTS enrichment** for speech actions (generates audio assets and attaches `mediaRef`)
- **Widget configuration** for interactive types (injects default scoring rules, timeouts, etc.)

## Integration in the Generation Pipeline

`generateSceneActions` sits at the end of a three-stage generation sequence:

```

SceneOutline ──► generateSceneContent ──► SceneContent ──► generateSceneActions ──► SceneAction[]

```

The function is **idempotent by design**—re-running with identical inputs produces consistent outputs (LLM temperature is pinned for reproducibility).

## Practical Code Examples

### Direct Invocation (Test Pattern)

```typescript
import {
  generateSceneActions,
  generateSceneContent,
  type AICallFn,
} from '@openmaic/generation';

const aiCall: AICallFn = async (prompt) => JSON.stringify([
  { 
    type: 'speech', 
    targetId: 'text_1', 
    payload: { text: 'Welcome to this lesson' } 
  },
  {
    type: 'widget',
    targetId: 'quiz_1',
    payload: { kind: 'multiple-choice', options: ['A', 'B', 'C'] }
  }
]);

const outline = { id: 'scene-1', type: 'interactive', title: 'Intro' };
const content = await generateSceneContent(outline, aiCall);

const actions = await generateSceneActions(outline, content, aiCall, {
  logger: console,
});

console.log(actions.length); // 2

```

### Server-Side Batch Generation

```typescript
import { generateSceneContent, generateSceneActions } from '@openmaic/generation';

async function generateLesson(requirements: Requirements) {
  const outlines = await generateSceneOutlinesFromRequirements(requirements, aiCall);
  
  return Promise.all(
    outlines.map(async (outline) => {
      const content = await generateSceneContent(outline, aiCall);
      const actions = await generateSceneActions(outline, content, aiCall);
      
      return { ...outline, content, actions };
    })
  );
}

```

## Key Source Files

| File | Purpose | Location |
|------|---------|----------|
| [`scene-generator.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/scene-generator.ts) | **`generateSceneActions` implementation** | `packages/@openmaic/generation/src/scene-generator.ts` |
| [`scene-generator.test.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/scene-generator.test.ts) | Unit tests including retry scenarios | `packages/@openmaic/generation/test/` |
| `generation-node-smoke.mjs` | CLI demonstration of full pipeline | `scripts/generation-node-smoke.mjs` |
| [`types/stage.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/types/stage.ts) | `SceneAction` type definitions | [`lib/types/stage.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/lib/types/stage.ts) |

## Summary

- **`generateSceneActions`** transforms `SceneContent` into executable `SceneAction[]` via structured LLM prompts
- Located in **`packages/@openmaic/generation/src/scene-generator.ts`** with extensive test coverage
- Features **defensive null handling**, configurable **retry policies**, and **lenient validation** for production reliability
- Supports **idempotent execution** through deterministic temperature settings
- Extensible architecture allows custom `aiCall` injection for different LLM providers or mocking in tests

## Frequently Asked Questions

### What happens if the LLM returns malformed JSON?

The function wraps `JSON.parse` in a try-catch and applies the configured `retryPolicy`. Persistent failures result in an empty action array with error logging, allowing pipeline continuation. Partially valid arrays are filtered to keep only schema-compliant actions.

### Can I use `generateSceneActions` without first running `generateSceneContent`?

Technically yes, but passing `null` for `content` triggers the early-exit guard and returns an empty array. The function expects rich content to inform action generation—running it standalone defeats its purpose in the OpenMAIC architecture.

### How does retry behavior work for rate-limited LLM providers?

The `retryPolicy` parameter supports exponential back-off with configurable `maxAttempts`, `baseDelayMs`, and `maxDelayMs`. According to the source implementation, defaults use jitter to prevent thundering-herd issues when multiple scenes generate simultaneously.

### Where are the generated speech audio files stored?

The function delegates to a TTS helper that writes assets to the configured media store and returns a `mediaRef` attached to the speech action. The storage backend is injected through the broader generation context, not hardcoded in `generateSceneActions` itself.