# How OpenMAIC Handles Multi-Format Document Parsing: PDF, Audio, and Video Extraction

> OpenMAIC seamlessly parses multi-format documents, extracting text, images, audio, and video from PDFs and media streams into a unified material layer for efficient data handling.

- Repository: [MAIC/OpenMAIC](https://github.com/THU-MAIC/OpenMAIC)
- Tags: how-to-guide
- Published: 2026-09-13

---

**OpenMAIC ingests PDFs through a pluggable provider pipeline that extracts text and images while forwarding embedded audio and video streams to dedicated media resolvers, storing all assets in a unified material layer.**

THU-MAIC/OpenMAIC treats every supported document type as a *material* that can be ingested, parsed, and stored in a unified media layer. When performing **multi-format document parsing**, the system normalizes content from PDFs—including embedded multimedia—into a consistent data model that downstream AI features can consume. The architecture centers on provider-agnostic extraction and centralized asset lifecycle management.

## The PDF Parsing Pipeline

OpenMAIC implements a three-stage pipeline for PDF ingestion that handles both static content and binary media streams.

### Provider Selection and Configuration

The system uses a pluggable provider interface defined in [`lib/pdf/constants.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/lib/pdf/constants.ts). This file exports `PDF_PROVIDERS` and a lookup helper that selects a provider ID—such as `unpdf` or `mineru-cloud`—based on server configuration. The **PDFProvider** interface abstracts extraction logic, ensuring the API layer remains agnostic to the underlying parsing library.

### Text and Image Extraction

The API route [`app/api/parse-pdf/route.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/app/api/parse-pdf/route.ts) receives uploaded PDFs and delegates extraction to the configured provider. The provider returns the document's plain text along with an array of image descriptors (`PdfImage[]`) containing binary data, page numbers, and dimensions. The test suite [`tests/providers/provider-neutrality-guard.test.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/tests/providers/provider-neutrality-guard.test.ts) verifies that this behavior remains consistent regardless of which provider backend is active.

### Embedded Audio and Video Resolution

When a PDF contains embedded media streams—such as annotation-referenced audio clips—the parser forwards those streams to specialized resolvers. The [`resolve-audio-bytes.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/resolve-audio-bytes.ts) module handles audio extraction, inferring MIME types like `audio/mpeg` and `audio/x-m4a`, while [`resolve-video-bytes.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/resolve-video-bytes.ts) processes video streams for formats including `video/mp4`. These resolvers persist binary blobs via the [`media-lifecycle.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/media-lifecycle.ts) manager, which tracks assets through the workbench's storage layer.

## Unified Material Storage Architecture

All extracted content—text, images, audio, and video—converges in the same material table used for generic uploads. The resulting data structure adheres to the `ParsedPdf` interface:

```typescript
interface ParsedPdf {
  text: string;                     // Full textual content
  images: PdfImage[];               // [{id, src, pageNumber, width, height}]
  audioAssets?: MediaAsset[];       // Optional audio tracks from annotations
  videoAssets?: MediaAsset[];       // Optional video tracks from annotations
}

```

As demonstrated in [`tests/document/session-sources.test.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/tests/document/session-sources.test.ts), the storage layer attaches asset IDs and MIME types to session metadata. Downstream features—such as the classroom player, video export, or outline generation—query the media layer using these IDs, treating PDF-derived assets identically to standalone uploads.

## Implementation Examples

### Parsing a PDF via the Public API

Client-side code uploads a file to the parsing endpoint and receives a normalized payload:

```typescript
// client-side example (React)
async function parsePdf(file: File) {
  const form = new FormData();
  form.append('file', file);
  const resp = await fetch('/api/parse-pdf', { method: 'POST', body: form });
  if (!resp.ok) throw new Error('PDF parsing failed');
  const parsed: ParsedPdf = await resp.json();
  console.log('PDF text:', parsed.text);
  console.log('Found images:', parsed.images.length);
  console.log('Embedded audio assets:', parsed.audioAssets?.length ?? 0);
}

```

### Accessing Extracted Media in a Lesson

Server-side components retrieve material records that include embedded media references:

```typescript
// server-side example (Node)
import { getMaterialById } from '@/lib/workbench/material-client';

async function getLessonMedia(lessonId: string) {
  const material = await getMaterialById(lessonId);
  // `material` contains `audioAssets` and `videoAssets` populated by the parser
  return {
    audioUrls: material.audioAssets?.map(a => a.url) ?? [],
    videoUrls: material.videoAssets?.map(v => v.url) ?? [],
  };
}

```

### Manual Provider Invocation

For advanced use cases, you can bypass the API route and invoke a specific provider directly:

```typescript
import { pdfProviders, getPdfProvider } from '@/lib/pdf/constants';

async function extractWithProvider(pdfKey: string, providerId = 'unpdf') {
  const provider = getPdfProvider(providerId);
  const { text, images } = await provider.extract(pdfKey);
  // Certain providers expose `extractMedia()` for embedded audio/video
  const { audio, video } = await provider.extractMedia?.(pdfKey) ?? {};
  return { text, images, audio, video };
}

```

## Summary

- OpenMAIC uses a **provider-agnostic pipeline** defined in [`lib/pdf/constants.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/lib/pdf/constants.ts) to parse PDFs, supporting backends like `unpdf` and `mineru-cloud`.
- The `/api/parse-pdf` route extracts text and images while the `ParsedPdf` interface standardizes output across all providers.
- Embedded audio and video streams are resolved by [`resolve-audio-bytes.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/resolve-audio-bytes.ts) and [`resolve-video-bytes.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/resolve-video-bytes.ts), then persisted through [`media-lifecycle.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/media-lifecycle.ts).
- All assets share a unified storage layer, enabling AI-driven features to treat PDF-derived media identically to standalone uploads.

## Frequently Asked Questions

### How does OpenMAIC determine which PDF provider to use?

The system reads server configuration to select a provider ID from the `PDF_PROVIDERS` registry defined in [`lib/pdf/constants.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/lib/pdf/constants.ts). The `getPdfProvider()` helper returns the corresponding implementation, allowing the API route to call `provider.extract()` without hardcoding library-specific logic.

### What audio and video formats are supported for extraction?

The resolvers in [`lib/media/resolve-audio-bytes.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/lib/media/resolve-audio-bytes.ts) and [`lib/media/resolve-video-bytes.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/lib/media/resolve-video-bytes.ts) infer MIME types including `audio/mpeg`, `audio/x-m4a`, and `video/mp4`. These modules process raw binary blobs from PDF annotations and store them as trackable media assets.

### Can the parsing pipeline handle PDFs without embedded media?

Yes. The `ParsedPdf` interface defines `audioAssets` and `videoAssets` as optional properties. If a PDF contains only text and static images, the pipeline populates only the `text` and `images` fields, leaving media arrays undefined. The provider neutrality tests verify this behavior for document-only inputs.

### Where are extracted assets stored after parsing?

Extracted binaries are persisted through the workbench's [`media-lifecycle.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/media-lifecycle.ts) manager, which stores them in the material table alongside generic uploads. Asset IDs and MIME types attach to session metadata as shown in [`tests/document/session-sources.test.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/tests/document/session-sources.test.ts), allowing downstream services to query URLs via the standard material client.