How the Materials Pipeline Works in OpenMAIC: From Upload to AI Extraction
OpenMAIC processes user-uploaded files through a multi-stage materials pipeline that stores them in an owner-scoped library, enforces quotas and concurrency limits, and extracts structured content to feed the AI-driven lesson generation engine.
The OpenMAIC repository implements a robust materials pipeline that transforms raw documents—PDFs, Word files, PowerPoints, images, audio, and video—into structured lesson content. This pipeline handles everything from client-side upload scheduling to server-side storage, quota enforcement, and asynchronous content extraction. Understanding this flow is essential for developers integrating with the platform or extending its functionality.
Overview of the Materials Pipeline Architecture
The materials pipeline follows a clear path from the workbench UI to the generation engine. When a user uploads a file, the system creates an owner-material record tied to their account, streams the bytes to durable storage, and optionally triggers extraction jobs that parse the content for AI consumption.
According to the OpenMAIC source code, the pipeline consists of five distinct phases: client upload initiation, server-side storage with quota checks, upload scheduling with concurrency control, content extraction via events, and final integration into the two-stage lesson generation process (Outline → Scenes).
Step 1: Client-Side Upload Initiation
The entry point for all material uploads is the uploadWorkbenchMaterial function defined in lib/workbench/session-store.ts (lines 36–72). This client-side API accepts a File object and sends a POST request to /api/materials.
The function handles the initial HTTP request and returns a promise that resolves with the material metadata, including a unique materialId. The client supports a wide range of MIME types, including PDF, Word documents, PowerPoint presentations, spreadsheets, plain text, images, audio, and video files.
import { uploadWorkbenchMaterial } from '@/lib/workbench/session-store';
const file = new File(['...'], 'lecture.pdf', { type: 'application/pdf' });
try {
const material = await uploadWorkbenchMaterial(file);
console.log('Uploaded material ID:', material.materialId);
} catch (e) {
console.error('Upload failed:', e);
}
Step 2: Server-Side Handling and Storage
The server receives the upload via the Next.js API route in app/api/materials/route.ts. This route performs three critical operations: it validates the request, creates an owner-material row with status uploading, and streams the file bytes into a neutral byte store.
The schema for owner materials is defined in lib/persistence/owner-materials.ts (lines 14–24). Each row stores the owner ID, MIME type, file size, original filename, extraction status, and the object key (ossKey) referencing the stored bytes. Once the upload stream completes successfully, the status transitions to ready, making the material available for extraction.
Step 3: Quota Enforcement and Slot Management
Before persisting any new material, the server validates storage limits in lib/persistence/owner-materials.ts (lines 71–82). The system checks both the owner’s material count and total byte quota, throwing a MaterialQuotaExceededError if limits are exceeded.
Additionally, the MaterialSlotLedger class in lib/workbench/material-upload-scheduling.ts (lines 1–34) maintains a client-side count of occupied, pending, and completed material slots. This ledger enforces a hard limit of 20 materials per lesson generation request, ensuring the composer never receives more source documents than it can effectively process.
Step 4: Upload Scheduling and Concurrency Control
OpenMAIC implements sophisticated upload scheduling to manage network resources and authentication state. The scheduleMaterialUploadBatch function in lib/workbench/material-upload-scheduling.ts (lines 36–56) coordinates uploads with a concurrency limit of three simultaneous requests (MATERIAL_UPLOAD_CONCURRENCY = 3).
The scheduling system uses createMaterialUploadIdentityGate() to manage authentication flow. The first successful upload establishes an HttpOnly owner cookie; subsequent uploads in the batch execute in parallel once this identity is confirmed. This pattern prevents race conditions during session initialization while maximizing throughput.
import {
createMaterialUploadIdentityGate,
scheduleMaterialUploadBatch,
} from '@/lib/workbench/material-upload-scheduling';
const gate = createMaterialUploadIdentityGate();
const files = [file1, file2, file3]; // array of File objects
await Promise.all(
files.map((file) =>
scheduleMaterialUploadBatch(gate, [file], async (f) => {
await uploadWorkbenchMaterial(f);
return true; // indicate success for identity establishment
})
)
);
Step 5: Content Extraction and Event Handling
Once a material reaches ready status, the server may launch asynchronous extraction jobs—such as PDF-to-text conversion, OCR for images, or audio transcription. These jobs emit events of kind material_extraction that flow back to the client.
In lib/workbench/session-store.ts (lines 74–82), the session store reducer handles these events by advancing the internal sequence ID. While the current implementation folds these events to nothing in the UI (producing no visible lines), the extraction results are later attached to the agent's reply and fed into the generation pipeline.
Integration with Lesson Generation
After extraction completes, the material content becomes available to the AI generation engine. As documented in README.md (lines 89–98), OpenMAIC uses a two-stage generation process: first producing a Lesson Outline, then fleshing out each outline item into Scenes (interactive slides, quizzes, or simulations). The extracted material content provides the knowledge base for both stages, grounding the AI's output in the user's source documents.
Key Source Files
lib/workbench/session-store.ts: Client-side API foruploadWorkbenchMaterialand event handling formaterial_extraction.lib/persistence/owner-materials.ts: Schema definition for theowner_materialtable and quota enforcement logic.lib/workbench/material-upload-scheduling.ts: Concurrency control,MaterialSlotLedger, and batch upload scheduling.app/api/materials/route.ts: Server-side API route handling file streams and database insertion.README.md: High-level documentation of the generation pipeline stages.
Summary
- The materials pipeline in OpenMAIC moves files from client upload through server storage, quota validation, and content extraction before feeding the AI generation engine.
- Quota enforcement occurs at the persistence layer via
MaterialQuotaExceededError, while theMaterialSlotLedgerenforces a 20-material limit per lesson. - Concurrency control limits simultaneous uploads to three requests, with the first upload establishing the owner identity cookie.
- Content extraction runs asynchronously, emitting
material_extractionevents that the client session store processes to track progress. - The pipeline supports diverse file types and integrates with a two-stage generation process that produces structured lesson outlines and interactive scenes.
Frequently Asked Questions
What file types does the OpenMAIC materials pipeline support?
The pipeline accepts PDF documents, Microsoft Word files, PowerPoint presentations, spreadsheets, plain text files, images, audio, and video formats. The system stores the MIME type in the owner_material table and routes the file to appropriate extraction handlers based on this metadata.
How does OpenMAIC enforce storage quotas?
Before creating a new owner-material record, the server checks the owner's existing material count and total byte usage against predefined limits in lib/persistence/owner-materials.ts (lines 71–82). If either limit is exceeded, the system throws a MaterialQuotaExceededError and rejects the upload before any bytes are stored.
What is the maximum number of materials per lesson?
The MaterialSlotLedger class enforces a hard limit of 20 materials per lesson generation request. This limit exists in lib/workbench/material-upload-scheduling.ts and ensures the AI composer receives a manageable number of source documents to process during outline and scene generation.
How does the extraction process integrate with lesson generation?
After a material reaches ready status, asynchronous extraction jobs parse the file content and emit material_extraction events. The extracted text or structured data is then attached to the agent's reply context and used as input for the two-stage generation process—first creating a Lesson Outline, then expanding each point into detailed Scenes with interactive elements.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →