# How Document Loaders Parse PDF, Office, and EPUB Files in 5ire

> Discover how 5ire's document loaders parse PDF, Office, and EPUB files. Learn about the modular subsystem converting binary buffers to plain text ContentParts.

- Repository: [Ironben/5ire](https://github.com/nanbingxyz/5ire)
- Tags: internals
- Published: 2026-03-07

---

**The 5ire repository implements a modular document-loader subsystem where specialized loader classes convert binary file buffers into plain-text ContentPart objects, using pdf-parse for PDFs, officeparser for Office documents, and a placeholder implementation for EPUB files.**

The `nanbingxyz/5ire` project provides an extensible document ingestion pipeline that transforms various file formats into structured text for downstream AI services. Understanding how these document loaders parse PDF, Office, and EPUB files reveals the architecture's flexibility and current capabilities.

## The Document Loader Architecture

All loaders implement a common interface defined in the document-loader module. The contract requires a `load` method that accepts a `Uint8Array` buffer and optional MIME type, returning a promise of `ContentPart` objects.

The `ContentPart` type structure:

```typescript
export type ContentPart = {
  type: "text";
  text: string;
};

```

Registration occurs centrally in [`DocumentLoader.ts`](https://github.com/nanbingxyz/5ire/blob/main/DocumentLoader.ts) using the `registerLoader` method, which maps MIME types to their respective handler instances.

## Parsing PDF Files with PDFLoader

The `PDFLoader` class in [`src/main/next/document-loader/PDFLoader.ts`](https://github.com/nanbingxyz/5ire/blob/main/src/main/next/document-loader/PDFLoader.ts) handles `application/pdf` files. It leverages the **pdf-parse** library to extract text content from binary PDF data.

Implementation details:

- Uses dynamic import `import("pdf-parse")` to load the library on-demand
- Instantiates `PDFParse` with a custom `CanvasFactory` to satisfy the parser's rendering requirements
- Calls `getText()` to retrieve the full document text as a single string
- Wraps the result in a `ContentPart` object with type `"text"`

Error handling transforms parsing failures into descriptive messages indicating the PDF could not be processed.

## Extracting Text from Office Documents

The `OfficeLoader` class in [`src/main/next/document-loader/OfficeLoader.ts`](https://github.com/nanbingxyz/5ire/blob/main/src/main/next/document-loader/OfficeLoader.ts) processes Microsoft Office and OpenDocument formats. It supports `.docx`, `.pptx`, `.xlsx`, `.odt`, `.odp`, and `.ods` files through the **officeparser** package.

Parsing strategy:

- Converts the incoming `Uint8Array` to a Node.js `Buffer`
- Invokes `officeParser.parseOfficeAsync(buffer)` which auto-detects the document format
- Receives extracted plain text regardless of the specific Office format
- Returns the text wrapped in a `ContentPart` object

This unified approach eliminates the need for format-specific parsing logic, handling all supported Office types through a single API call.

## EPUB Support Status

The `EpubLoader` class in [`src/main/next/document-loader/EpubLoader.ts`](https://github.com/nanbingxyz/5ire/blob/main/src/main/next/document-loader/EpubLoader.ts) currently exists as a **placeholder implementation**. The file contains commented-out skeleton code indicating the intended structure:

```typescript
// export class EpubLoader implements Loader { … }
// getSupportedFileExtensions = () => ({ epub: 'application/epub+zip' });

```

No active parsing logic is implemented. Adding EPUB support would require integrating a library such as [`epub.js`](https://github.com/nanbingxyz/5ire/blob/main/epub.js) or `node-epub` and implementing the `load` method to extract text content from the EPUB archive structure.

## Using the DocumentLoader Registry

While individual loaders can be instantiated directly, the recommended approach uses the central `DocumentLoader` registry for automatic MIME type detection and dispatch.

Registration pattern from [`DocumentLoader.ts`](https://github.com/nanbingxyz/5ire/blob/main/DocumentLoader.ts):

```typescript
DocumentLoader.registerLoader(new PDFLoader());
DocumentLoader.registerLoader(new OfficeLoader());
// Future: DocumentLoader.registerLoader(new EpubLoader());

```

Generic usage example:

```typescript
import { DocumentLoader } from '@/main/next/document-loader/DocumentLoader';
import { readFile } from 'fs/promises';
import mime from 'mime-types';

async function loadAnyFile(path: string) {
  const buffer = await readFile(path);
  const mimeType = mime.lookup(path) || 'application/octet-stream';
  const parts = await DocumentLoader.load(buffer, mimeType);
  console.log(parts.map(p => p.text).join('\n---\n'));
}

```

This pattern enables seamless handling of multiple file formats without hardcoding specific loader classes in consuming code.

## Summary

- The 5ire document-loader subsystem uses a unified `Loader` interface to convert binary files into structured `ContentPart` objects.
- **PDFLoader** leverages `pdf-parse` with a custom `CanvasFactory` to extract text from PDF documents.
- **OfficeLoader** utilizes `officeparser` to handle Microsoft Office and OpenDocument formats through a single API.
- **EpubLoader** remains unimplemented, containing only placeholder code for future EPUB support.
- The central `DocumentLoader` registry enables automatic MIME type detection and dispatch, simplifying integration for downstream services.

## Frequently Asked Questions

### What file formats does the 5ire document loader currently support?

The 5ire document loader currently supports PDF files (`application/pdf`) through **PDFLoader**, and Microsoft Office formats (`.docx`, `.pptx`, `.xlsx`) along with OpenDocument formats (`.odt`, `.odp`, `.ods`) through **OfficeLoader**. EPUB support is planned but not yet implemented.

### How does the PDFLoader extract text without native canvas support?

**PDFLoader** dynamically imports the `pdf-parse` library and instantiates it with a custom `CanvasFactory` shim. This factory satisfies the PDF parser's rendering requirements without relying on a native browser or Node.js canvas implementation, enabling the `getText()` method to extract content in worker contexts.

### Can I use the document loaders outside of the 5ire application?

Yes, each loader implements the standalone `Loader` interface and can be imported and used independently. Instantiate `PDFLoader` or `OfficeLoader` directly and call their `load(buffer, mimeType)` methods with a `Uint8Array` buffer. Alternatively, use the central `DocumentLoader` registry for automatic format detection.

### Why is EPUB support not implemented in the current codebase?

The **EpubLoader** class exists as a placeholder with commented-out skeleton code indicating the intended structure. The implementation requires selecting and integrating an appropriate EPUB parsing library (such as [`epub.js`](https://github.com/nanbingxyz/5ire/blob/main/epub.js) or `node-epub`) and implementing the text extraction logic to comply with the `Loader` interface contract.