# Harvey-Labs Document Parsing Pipeline: How AI Agents Extract Text from .docx, .pptx, .xlsx, and .pdf Files

> Discover the Harvey-Labs document parsing pipeline that uses AI agents to extract text from docx, pptx, xlsx, and pdf files. Learn how it securely handles binary files in an isolated container.

- Repository: [Harvey/harvey-labs](https://github.com/harveyai/harvey-labs)
- Tags: how-to-guide
- Published: 2026-08-11

---

**Harvey-Labs uses a sandboxed document parsing pipeline that delegates all binary file extraction to an isolated container, separating potentially vulnerable parsing libraries from the host environment.**

This article examines the complete parsing architecture in the harveyai/harvey-labs repository, breaking down how Word documents, PDFs, PowerPoint presentations, and Excel spreadsheets are converted to plain text for AI agent consumption.

## Overview of the Sandboxed Parsing Architecture

The Harvey-Labs system treats document parsing as a security boundary. Rather than importing parsing libraries directly into the agent's execution environment, the codebase routes all binary document processing through [`sandbox/parsers/parse_doc.py`](https://github.com/harveyai/harvey-labs/blob/main/sandbox/parsers/parse_doc.py), which executes inside a purpose-built container.

The entry point is `Tool._read_and_parse()` in [`harness/tools.py`](https://github.com/harveyai/harvey-labs/blob/main/harness/tools.py). When an agent requests file content, this method inspects the file extension. For `.docx`, `.pdf`, `.pptx`, or `.xlsx` files, it immediately redirects to `Tool._parse_in_sandbox()`, which spawns the sandboxed parser.

This design eliminates attack surface: even if a malicious document exploits a vulnerability in `pdfplumber`, `pandas`, or `markitdown`, the compromise remains contained within the sandbox.

## How Each Document Format Is Parsed

The [`parse_doc.py`](https://github.com/harveyai/harvey-labs/blob/main/parse_doc.py) script implements four specialized functions, each leveraging a different tool best suited to the format.

### .docx Files: Pandoc Markdown Conversion

Word documents undergo conversion to Markdown using the **pandoc** CLI tool.

```python

# sandbox/parsers/parse_doc.py - parse_docx implementation

subprocess.run(
    ["pandoc", path, "-t", "markdown", "--wrap=none"],
    capture_output=True,
    text=True,
    check=True
)

```

The `--wrap=none` flag prevents line-wrapping, preserving the original document structure. The Markdown output maintains headings, lists, and basic formatting that help the agent understand document hierarchy.

### .pdf Files: Page-Wise Extraction with pdfplumber

PDF parsing uses **pdfplumber** for robust text and table extraction.

```python

# sandbox/parsers/parse_doc.py - parse_pdf implementation

with pdfplumber.open(path) as pdf:
    for page in pdf.pages:
        text = page.extract_text() or ""
        # Table extraction via page.extract_tables() when present

```

The pipeline iterates through `pdf.pages`, extracting text from each page and attempting table detection. This handles multi-column layouts and tabular data better than simple text dump approaches.

### .pptx Files: MarkItDown PowerPoint Conversion

PowerPoint presentations are processed by **markitdown**, Microsoft's purpose-built conversion library.

```python

# sandbox/parsers/parse_doc.py - parse_pptx implementation

MarkItDown().convert(path).text_content

```

The `MarkItDown` class extracts slide text content in presentation order, producing a linear text representation that preserves the narrative flow of the deck.

### .xlsx Files: pandas Table Stringification

Excel spreadsheets use **pandas** `read_excel` with `sheet_name=None` to load all worksheets.

```python

# sandbox/parsers/parse_doc.py - parse_xlsx implementation

dfs = pd.read_excel(path, sheet_name=None)
for sheet_name, df in dfs.items():
    # Stringify each DataFrame with sheet header delineation

```

Each worksheet becomes a stringified table with explicit sheet name headers, allowing the agent to distinguish between multiple data tabs.

## Complete Execution Flow: From File Path to Extracted Text

Understanding the full call chain clarifies how components interact:

1. **Agent request**: `tool._read("reports/annual_report.pdf")`
2. **Extension detection**: `Tool._read_and_parse()` identifies `.pdf`
3. **Sandbox invocation**: `Tool._parse_in_sandbox("pdf", sb_path)` constructs the command
4. **Container execution**: `parse-doc pdf /sandbox/documents/reports/annual_report.pdf`
5. **Format-specific parsing**: [`parse_doc.py`](https://github.com/harveyai/harvey-labs/blob/main/parse_doc.py) dispatches to `parse_pdf()` → `pdfplumber`
6. **Result return**: Text streams to stdout, captured and returned to the agent

```python

# Example usage from an agent context

content = tool._read("reports/annual_report.pdf", offset=None, limit=None)
print(content)  # Extracted text or structured error message

```

## Error Handling and Security Guarantees

The sandbox design provides clear failure semantics. Successful parsing exits with code 0 and writes extracted text to stdout. Failures—corrupted files, encrypted PDFs, malformed Office documents—write diagnostic messages to stderr and exit non-zero.

In [`harness/tools.py`](https://github.com/harveyai/harvey-labs/blob/main/harness/tools.py) (lines 71-80), `Tool._parse_in_sandbox()` captures these conditions:

- **Timeout enforcement**: Prevents parser hangs on malicious resource-exhaustion documents
- **Exit code inspection**: Distinguishes parsing failures from infrastructure errors
- **Sanitized error surfacing**: Returns concise failure messages to the agent without leaking sandbox internals

## Key Files and Their Responsibilities

| File | Purpose |
|------|---------|
| [`sandbox/parsers/parse_doc.py`](https://github.com/harveyai/harvey-labs/blob/main/sandbox/parsers/parse_doc.py) | Implements all four format parsers and CLI entry point |
| [`harness/tools.py`](https://github.com/harveyai/harvey-labs/blob/main/harness/tools.py) | Host-side orchestration via `_read_and_parse()` and `_parse_in_sandbox()` |
| `sandbox/Dockerfile` | Container definition with `pandoc`, `pdfplumber`, `pandas`, `markitdown` |

The Dockerfile pins specific versions of these parsing dependencies, enabling reproducible builds and controlled security updates.

## Summary

- **Security-first architecture**: All parsing runs in an isolated sandbox, protecting the host from document-borne exploits
- **Specialized tools per format**: Pandoc for Word, pdfplumber for PDFs, markitdown for PowerPoint, pandas for Excel
- **Consistent interface**: `parse-doc <format> <path>` CLI unifies invocation across all file types
- **Robust error handling**: Clear success/failure semantics with structured error propagation

## Frequently Asked Questions

### What happens if a PDF is password-protected or corrupted?

The pdfplumber parser will raise an exception, [`parse_doc.py`](https://github.com/harveyai/harvey-labs/blob/main/parse_doc.py) will exit with a non-zero status, and `Tool._parse_in_sandbox()` will capture the stderr output and return a concise error message to the agent. The host environment remains unaffected.

### Why doesn't Harvey-Labs parse documents directly in the main process?

Direct parsing would require installing `pdfplumber`, `pandas`, and `markitdown` in the agent execution environment. These libraries have extensive dependency trees and historical CVEs. The sandbox containment strategy follows the principle of least privilege—untrusted document content never touches the host interpreter.

### Can the parsing pipeline handle large files or memory constraints?

The sandbox container operates with resource limits defined in its orchestration configuration. The `pdfplumber` and `pandas` implementations stream where possible, though extremely large spreadsheets or high-page-count PDFs may hit container memory boundaries. The `Tool._parse_in_sandbox()` method includes timeout handling to prevent indefinite hangs.

### Is the extracted text formatted or plain?

Output varies by source format. Pandoc produces Markdown with structural markers (headings, lists). pdfplumber returns plain text with attempted table preservation. markitdown extracts linear text from slides. pandas stringifies DataFrames with column alignment. In all cases, the goal is agent-readable text rather than visual fidelity.