Harvey-Labs Document Parsing Pipeline: How AI Agents Extract Text from .docx, .pptx, .xlsx, and .pdf Files
Harvey-Labs uses a sandboxed document parsing pipeline that delegates all binary file extraction to an isolated container, separating potentially vulnerable parsing libraries from the host environment.
This article examines the complete parsing architecture in the harveyai/harvey-labs repository, breaking down how Word documents, PDFs, PowerPoint presentations, and Excel spreadsheets are converted to plain text for AI agent consumption.
Overview of the Sandboxed Parsing Architecture
The Harvey-Labs system treats document parsing as a security boundary. Rather than importing parsing libraries directly into the agent's execution environment, the codebase routes all binary document processing through sandbox/parsers/parse_doc.py, which executes inside a purpose-built container.
The entry point is Tool._read_and_parse() in harness/tools.py. When an agent requests file content, this method inspects the file extension. For .docx, .pdf, .pptx, or .xlsx files, it immediately redirects to Tool._parse_in_sandbox(), which spawns the sandboxed parser.
This design eliminates attack surface: even if a malicious document exploits a vulnerability in pdfplumber, pandas, or markitdown, the compromise remains contained within the sandbox.
How Each Document Format Is Parsed
The parse_doc.py script implements four specialized functions, each leveraging a different tool best suited to the format.
.docx Files: Pandoc Markdown Conversion
Word documents undergo conversion to Markdown using the pandoc CLI tool.
# sandbox/parsers/parse_doc.py - parse_docx implementation
subprocess.run(
["pandoc", path, "-t", "markdown", "--wrap=none"],
capture_output=True,
text=True,
check=True
)
The --wrap=none flag prevents line-wrapping, preserving the original document structure. The Markdown output maintains headings, lists, and basic formatting that help the agent understand document hierarchy.
.pdf Files: Page-Wise Extraction with pdfplumber
PDF parsing uses pdfplumber for robust text and table extraction.
# sandbox/parsers/parse_doc.py - parse_pdf implementation
with pdfplumber.open(path) as pdf:
for page in pdf.pages:
text = page.extract_text() or ""
# Table extraction via page.extract_tables() when present
The pipeline iterates through pdf.pages, extracting text from each page and attempting table detection. This handles multi-column layouts and tabular data better than simple text dump approaches.
.pptx Files: MarkItDown PowerPoint Conversion
PowerPoint presentations are processed by markitdown, Microsoft's purpose-built conversion library.
# sandbox/parsers/parse_doc.py - parse_pptx implementation
MarkItDown().convert(path).text_content
The MarkItDown class extracts slide text content in presentation order, producing a linear text representation that preserves the narrative flow of the deck.
.xlsx Files: pandas Table Stringification
Excel spreadsheets use pandas read_excel with sheet_name=None to load all worksheets.
# sandbox/parsers/parse_doc.py - parse_xlsx implementation
dfs = pd.read_excel(path, sheet_name=None)
for sheet_name, df in dfs.items():
# Stringify each DataFrame with sheet header delineation
Each worksheet becomes a stringified table with explicit sheet name headers, allowing the agent to distinguish between multiple data tabs.
Complete Execution Flow: From File Path to Extracted Text
Understanding the full call chain clarifies how components interact:
- Agent request:
tool._read("reports/annual_report.pdf") - Extension detection:
Tool._read_and_parse()identifies.pdf - Sandbox invocation:
Tool._parse_in_sandbox("pdf", sb_path)constructs the command - Container execution:
parse-doc pdf /sandbox/documents/reports/annual_report.pdf - Format-specific parsing:
parse_doc.pydispatches toparse_pdf()→pdfplumber - Result return: Text streams to stdout, captured and returned to the agent
# Example usage from an agent context
content = tool._read("reports/annual_report.pdf", offset=None, limit=None)
print(content) # Extracted text or structured error message
Error Handling and Security Guarantees
The sandbox design provides clear failure semantics. Successful parsing exits with code 0 and writes extracted text to stdout. Failures—corrupted files, encrypted PDFs, malformed Office documents—write diagnostic messages to stderr and exit non-zero.
In harness/tools.py (lines 71-80), Tool._parse_in_sandbox() captures these conditions:
- Timeout enforcement: Prevents parser hangs on malicious resource-exhaustion documents
- Exit code inspection: Distinguishes parsing failures from infrastructure errors
- Sanitized error surfacing: Returns concise failure messages to the agent without leaking sandbox internals
Key Files and Their Responsibilities
| File | Purpose |
|---|---|
sandbox/parsers/parse_doc.py |
Implements all four format parsers and CLI entry point |
harness/tools.py |
Host-side orchestration via _read_and_parse() and _parse_in_sandbox() |
sandbox/Dockerfile |
Container definition with pandoc, pdfplumber, pandas, markitdown |
The Dockerfile pins specific versions of these parsing dependencies, enabling reproducible builds and controlled security updates.
Summary
- Security-first architecture: All parsing runs in an isolated sandbox, protecting the host from document-borne exploits
- Specialized tools per format: Pandoc for Word, pdfplumber for PDFs, markitdown for PowerPoint, pandas for Excel
- Consistent interface:
parse-doc <format> <path>CLI unifies invocation across all file types - Robust error handling: Clear success/failure semantics with structured error propagation
Frequently Asked Questions
What happens if a PDF is password-protected or corrupted?
The pdfplumber parser will raise an exception, parse_doc.py will exit with a non-zero status, and Tool._parse_in_sandbox() will capture the stderr output and return a concise error message to the agent. The host environment remains unaffected.
Why doesn't Harvey-Labs parse documents directly in the main process?
Direct parsing would require installing pdfplumber, pandas, and markitdown in the agent execution environment. These libraries have extensive dependency trees and historical CVEs. The sandbox containment strategy follows the principle of least privilege—untrusted document content never touches the host interpreter.
Can the parsing pipeline handle large files or memory constraints?
The sandbox container operates with resource limits defined in its orchestration configuration. The pdfplumber and pandas implementations stream where possible, though extremely large spreadsheets or high-page-count PDFs may hit container memory boundaries. The Tool._parse_in_sandbox() method includes timeout handling to prevent indefinite hangs.
Is the extracted text formatted or plain?
Output varies by source format. Pandoc produces Markdown with structural markers (headings, lists). pdfplumber returns plain text with attempted table preservation. markitdown extracts linear text from slides. pandas stringifies DataFrames with column alignment. In all cases, the goal is agent-readable text rather than visual fidelity.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →