How OpenMAIC Handles Multi-Format Document Parsing: PDF, Audio, and Video Extraction
OpenMAIC ingests PDFs through a pluggable provider pipeline that extracts text and images while forwarding embedded audio and video streams to dedicated media resolvers, storing all assets in a unified material layer.
THU-MAIC/OpenMAIC treats every supported document type as a material that can be ingested, parsed, and stored in a unified media layer. When performing multi-format document parsing, the system normalizes content from PDFs—including embedded multimedia—into a consistent data model that downstream AI features can consume. The architecture centers on provider-agnostic extraction and centralized asset lifecycle management.
The PDF Parsing Pipeline
OpenMAIC implements a three-stage pipeline for PDF ingestion that handles both static content and binary media streams.
Provider Selection and Configuration
The system uses a pluggable provider interface defined in lib/pdf/constants.ts. This file exports PDF_PROVIDERS and a lookup helper that selects a provider ID—such as unpdf or mineru-cloud—based on server configuration. The PDFProvider interface abstracts extraction logic, ensuring the API layer remains agnostic to the underlying parsing library.
Text and Image Extraction
The API route app/api/parse-pdf/route.ts receives uploaded PDFs and delegates extraction to the configured provider. The provider returns the document's plain text along with an array of image descriptors (PdfImage[]) containing binary data, page numbers, and dimensions. The test suite tests/providers/provider-neutrality-guard.test.ts verifies that this behavior remains consistent regardless of which provider backend is active.
Embedded Audio and Video Resolution
When a PDF contains embedded media streams—such as annotation-referenced audio clips—the parser forwards those streams to specialized resolvers. The resolve-audio-bytes.ts module handles audio extraction, inferring MIME types like audio/mpeg and audio/x-m4a, while resolve-video-bytes.ts processes video streams for formats including video/mp4. These resolvers persist binary blobs via the media-lifecycle.ts manager, which tracks assets through the workbench's storage layer.
Unified Material Storage Architecture
All extracted content—text, images, audio, and video—converges in the same material table used for generic uploads. The resulting data structure adheres to the ParsedPdf interface:
interface ParsedPdf {
text: string; // Full textual content
images: PdfImage[]; // [{id, src, pageNumber, width, height}]
audioAssets?: MediaAsset[]; // Optional audio tracks from annotations
videoAssets?: MediaAsset[]; // Optional video tracks from annotations
}
As demonstrated in tests/document/session-sources.test.ts, the storage layer attaches asset IDs and MIME types to session metadata. Downstream features—such as the classroom player, video export, or outline generation—query the media layer using these IDs, treating PDF-derived assets identically to standalone uploads.
Implementation Examples
Parsing a PDF via the Public API
Client-side code uploads a file to the parsing endpoint and receives a normalized payload:
// client-side example (React)
async function parsePdf(file: File) {
const form = new FormData();
form.append('file', file);
const resp = await fetch('/api/parse-pdf', { method: 'POST', body: form });
if (!resp.ok) throw new Error('PDF parsing failed');
const parsed: ParsedPdf = await resp.json();
console.log('PDF text:', parsed.text);
console.log('Found images:', parsed.images.length);
console.log('Embedded audio assets:', parsed.audioAssets?.length ?? 0);
}
Accessing Extracted Media in a Lesson
Server-side components retrieve material records that include embedded media references:
// server-side example (Node)
import { getMaterialById } from '@/lib/workbench/material-client';
async function getLessonMedia(lessonId: string) {
const material = await getMaterialById(lessonId);
// `material` contains `audioAssets` and `videoAssets` populated by the parser
return {
audioUrls: material.audioAssets?.map(a => a.url) ?? [],
videoUrls: material.videoAssets?.map(v => v.url) ?? [],
};
}
Manual Provider Invocation
For advanced use cases, you can bypass the API route and invoke a specific provider directly:
import { pdfProviders, getPdfProvider } from '@/lib/pdf/constants';
async function extractWithProvider(pdfKey: string, providerId = 'unpdf') {
const provider = getPdfProvider(providerId);
const { text, images } = await provider.extract(pdfKey);
// Certain providers expose `extractMedia()` for embedded audio/video
const { audio, video } = await provider.extractMedia?.(pdfKey) ?? {};
return { text, images, audio, video };
}
Summary
- OpenMAIC uses a provider-agnostic pipeline defined in
lib/pdf/constants.tsto parse PDFs, supporting backends likeunpdfandmineru-cloud. - The
/api/parse-pdfroute extracts text and images while theParsedPdfinterface standardizes output across all providers. - Embedded audio and video streams are resolved by
resolve-audio-bytes.tsandresolve-video-bytes.ts, then persisted throughmedia-lifecycle.ts. - All assets share a unified storage layer, enabling AI-driven features to treat PDF-derived media identically to standalone uploads.
Frequently Asked Questions
How does OpenMAIC determine which PDF provider to use?
The system reads server configuration to select a provider ID from the PDF_PROVIDERS registry defined in lib/pdf/constants.ts. The getPdfProvider() helper returns the corresponding implementation, allowing the API route to call provider.extract() without hardcoding library-specific logic.
What audio and video formats are supported for extraction?
The resolvers in lib/media/resolve-audio-bytes.ts and lib/media/resolve-video-bytes.ts infer MIME types including audio/mpeg, audio/x-m4a, and video/mp4. These modules process raw binary blobs from PDF annotations and store them as trackable media assets.
Can the parsing pipeline handle PDFs without embedded media?
Yes. The ParsedPdf interface defines audioAssets and videoAssets as optional properties. If a PDF contains only text and static images, the pipeline populates only the text and images fields, leaving media arrays undefined. The provider neutrality tests verify this behavior for document-only inputs.
Where are extracted assets stored after parsing?
Extracted binaries are persisted through the workbench's media-lifecycle.ts manager, which stores them in the material table alongside generic uploads. Asset IDs and MIME types attach to session metadata as shown in tests/document/session-sources.test.ts, allowing downstream services to query URLs via the standard material client.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →