# How to Handle Complex PDF Layouts for Extraction in pymupdf_rag.py

> Extract structured Markdown from complex PDFs with pymupdf_rag.py. Learn to handle multi-column text, tables, and graphics for robust data extraction.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: how-to-guide
- Published: 2026-07-07

---

**The [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) script converts complex PDF layouts—including multi‑column text, nested tables, and mixed graphics—into structured Markdown by orchestrating header detection, column analysis, table extraction, and graphic filtering algorithms.**

Handling complex PDF layouts requires more than simple text extraction. The [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) module in the interviewstreet/hiring-agent repository provides a comprehensive pipeline that preserves visual hierarchy while converting challenging documents into clean, GitHub‑flavored Markdown. This guide examines the specific architectural components and implementation strategies that enable robust extraction from sophisticated document structures.

## Core Components for Layout Analysis

The extraction engine relies on specialized components that handle distinct layout challenges. Each component targets specific structural elements to ensure accurate Markdown output.

### Header Detection with IdentifyHeaders and TocHeaders

The script implements two complementary strategies for detecting document headings. **IdentifyHeaders** analyzes font sizes across the document to infer hierarchical levels, mapping detected sizes to Markdown header markers (`#` through `######`). This logic resides in `IdentifyHeaders.__init__` (lines 84‑108) within [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py).

Alternatively, **TocHeaders** reads the PDF’s embedded Table‑Of‑Contents to determine heading levels precisely. The `TocHeaders.get_header_id` method (lines 79‑104) extracts entries directly from the document’s navigation structure, providing more reliable header detection for PDFs with consistent TOC metadata but inconsistent font sizing.

### Multi‑Column Text Processing with column_boxes

Complex layouts often feature multiple text columns that standard extraction reads incorrectly. The `column_boxes` function from [`pymupdf4llm/helpers/multi_column.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf4llm/helpers/multi_column.py) splits pages into logical text columns, even when columns overlap graphics or contain irregular spacing. 

In [`pymupdf_rag.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf_rag.py) (lines 60‑67), the script calls `column_boxes` to generate a list of rectangles representing each column, then processes these rectangles from top‑to‑bottom, left‑to‑right. This ensures the resulting Markdown maintains the correct reading order for multi‑column documents.

### Table Extraction Strategy

Tables are extracted before surrounding text processing to prevent row/column boundaries from being lost. The script uses `page.find_tables` with a configurable strategy—defaulting to `lines_strict`—to identify grid structures. 

In the extraction block (lines 188‑199), tables smaller than 2 rows or 2 columns are automatically discarded as noise. Valid tables are converted to Markdown syntax using `Table.to_markdown()`, preserving cell alignment and content hierarchy even when tables span multiple columns or contain merged cells.

### Graphic Filtering and Vector Drawings

Vector graphics and decorative lines often disrupt text flow. The script clusters vector drawings using `page.cluster_drawings`, then filters each cluster through the `is_significant` function (lines 56‑80). 

This function examines path complexity to distinguish meaningful illustrations from decorative lines or borders. The `refine_boxes` helper (lines 21‑34) further processes these regions to prevent non‑textual elements from interfering with column detection. Only "meaningful" graphics survive filtering and can be optionally saved or embedded in the output.

### Margin Clipping and Link Resolution

To exclude headers and footers from extraction, the script subtracts configurable margins from the page rectangle. The `parms.clip = page.rect + …` assignment (lines 18‑21) defines the active extraction area, ensuring only body content is processed.

For hyperlinks, the `resolve_links` function (lines 32‑66) converts PDF annotations to Markdown link syntax. It handles complex cases where a single text span contains multiple links—such as pipe‑separated or space‑separated references—by splitting the text and matching each segment to the correct URL using content‑based heuristics.

## Implementation Strategies for Complex Layouts

Different PDF architectures require specific parameter configurations. The following strategies address common complex layout scenarios using the `to_markdown()` orchestration function (core loop at lines 31‑48).

### Multi‑Column Document Processing

For research papers and reports with side‑by‑side columns, combine `column_boxes` analysis with margin controls. The script iterates over column rectangles returned by `column_boxes`, feeding each to `write_text()` independently. This maintains logical reading order while preventing text from adjacent columns from merging.

### Nested Tables and Mixed Content

When tables appear within text columns or contain embedded graphics, the extraction engine processes tables first (`output_tables` is called for the rectangle’s top), then handles surrounding text. Tables that intersect column rectangles are extracted and converted to Markdown tabular syntax before the remaining text flows around them.

### Graphics‑Heavy Document Handling

For design specifications or brochures with extensive vector art, adjust the `graphics_limit` parameter to cap processing after a specified number of drawings, preventing performance degradation on pages with thousands of tiny paths. Set `ignore_graphics=False` to retain meaningful illustrations while using `image_size_limit` to filter out small decorative images.

## Code Examples for Complex Layout Extraction

### Basic Extraction with Default Layout Handling

Use this configuration for standard complex layouts with mixed content:

```python
from pymupdf_rag import to_markdown

markdown = to_markdown(
    "reports/annual-report-2024.pdf",
    write_images=True,
    embed_images=False,
    image_path="output/images",
    table_strategy="lines_strict",
    detect_bg_color=True,
    ignore_graphics=False,
)
print(markdown)

```

The `table_strategy="lines_strict"` parameter forces the `page.find_tables` engine (lines 188‑199) to require grid lines, reducing false positives on multi‑column text. The `save_image` routine (lines 65‑95) writes extracted images to the specified `image_path` when `write_images=True`.

### TOC‑Based Header Detection for Structured Documents

For PDFs with reliable Table‑Of‑Contents navigation:

```python
from pymupdf_rag import to_markdown, TocHeaders

pdf = "guides/user-manual.pdf"
hdr_finder = TocHeaders(pdf)
markdown = to_markdown(
    pdf,
    hdr_info=hdr_finder,
    ignore_images=True,
    ignore_graphics=True,
    page_chunks=True,
)
print(markdown)

```

This leverages `TocHeaders.get_header_id` (lines 79‑104) to map TOC entries directly to Markdown headers, bypassing font‑size analysis for documents with explicit navigation structures.

### Handling Graphics‑Intensive PDFs

For design specifications with heavy vector and raster content:

```python
from pymupdf_rag import to_markdown

markdown = to_markdown(
    "specs/design-spec.pdf",
    write_images=True,
    embed_images=True,
    ignore_graphics=False,
    graphics_limit=200,
    image_size_limit=0.08,
    detect_bg_color=False,
)

```

Setting `embed_images=True` triggers base64 encoding within the `save_image` function, creating a single self‑contained Markdown file. The `graphics_limit` parameter prevents excessive processing when `is_significant` (lines 56‑80) evaluates numerous vector clusters.

### Fine‑Grained Column Control for Academic Papers

For two‑column research papers with footnotes:

```python
from pymupdf_rag import to_markdown

markdown = to_markdown(
    "papers/complex-layout.pdf",
    margins=[0.5, 0.5, 0.5, 0.5],
    page_chunks=False,
    ignore_images=True,
    ignore_graphics=True,
    table_strategy="lines_strict",
)

```

The `margins` parameter trims 0.5 inches from each side via the `parms.clip` logic (lines 18‑21), excluding running headers and footers while preserving the column structure detected by `column_boxes`.

## Summary

- **Header Detection**: Choose between `IdentifyHeaders` (font‑size analysis) or `TocHeaders` (TOC‑based) depending on PDF metadata quality.
- **Multi‑Column Processing**: The `column_boxes` function from [`pymupdf4llm/helpers/multi_column.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf4llm/helpers/multi_column.py) isolates column rectangles to maintain correct reading order.
- **Table Extraction**: Use `table_strategy="lines_strict"` with `page.find_tables` (lines 188‑199) for robust grid detection, discarding small tables automatically.
- **Graphic Filtering**: The `is_significant()` function (lines 56‑80) filters decorative vectors while preserving meaningful illustrations.
- **Configuration**: Adjust `margins`, `graphics_limit`, and `image_size_limit` to optimize processing for specific layout complexities.

## Frequently Asked Questions

### How does pymupdf_rag.py detect document headings in complex layouts?

The script implements two strategies: `IdentifyHeaders` analyzes font sizes across the document to infer hierarchy levels (lines 84‑108), while `TocHeaders` reads the PDF’s embedded Table‑Of‑Contents to determine exact header levels (lines 79‑104). Detected headings are mapped to Markdown header markers (`#` through `######`) based on their hierarchical position.

### What is the difference between IdentifyHeaders and TocHeaders?

`IdentifyHeaders` uses heuristics based on font size and formatting to guess heading levels, making it suitable for PDFs without proper TOC metadata. `TocHeaders` extracts heading information directly from the PDF’s navigation structure, providing precise hierarchy mapping for documents with well‑structured bookmarks but inconsistent visual styling.

### How does the script handle multi‑column PDF text extraction?

The `column_boxes` function from [`pymupdf4llm/helpers/multi_column.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pymupdf4llm/helpers/multi_column.py) analyzes the page geometry to identify separate column rectangles. The script processes these rectangles sequentially from top‑to‑bottom, left‑to‑right (lines 60‑67), feeding each to `write_text()` to ensure the Markdown output follows the correct reading order rather than extracting text in raw PDF stream order.

### Can pymupdf_rag.py extract tables that overlap with other content?

Yes, the script uses `page.find_tables` with configurable strategies (defaulting to `lines_strict` at lines 188‑199) to detect tables before processing surrounding text. Tables that intersect column rectangles are extracted and converted to Markdown tabular syntax via `Table.to_markdown()`, preserving row and column boundaries even when tables span multiple columns or contain nested elements.