Extracting Per-Page Markdown from PDFs with pdf-inspector: A Complete Guide

Use pdf2md --json to extract per-page Markdown from PDFs with pdf-inspector, producing page-level JSON output with isolated markdown blocks, tables, and links for each page.

PDF-inspector is a Rust-based tool from the Firecrawl ecosystem that converts PDF files into clean, structured Markdown on a per-page basis. Unlike simple text extractors, it performs sophisticated layout analysis—including column detection, reading order reconstruction, and table recognition—before emitting semantic Markdown. For workflows that need page-granular content, such as LLM prompting or paginated UIs, pdf2md with the --json flag provides exactly that structure.

How pdf-inspector Converts PDFs to Per-Page Markdown

The conversion pipeline in pdf-inspector is organized into three distinct stages, each implemented in dedicated Rust modules:

Stage Purpose Key Module
Detection Classifies PDF type (text-based, scanned, mixed, image-based) src/detector.rs
Extraction & Layout Parses content streams, builds font maps, detects columns and reading order src/extractor/content_stream.rs, src/extractor/layout.rs, src/extractor/fonts.rs
Markdown Generation Converts logical lines to Markdown, classifies elements, cleans output src/markdown/convert.rs, src/markdown/classify.rs, src/markdown/postprocess.rs

The public API exposed in src/lib.rs provides the process_pdf_with_options function, which both CLI binaries and language bindings consume.

CLI Usage for Per-Page Markdown Extraction

Basic JSON Output with Page Boundaries

The pdf2md binary in src/bin/pdf2md.rs handles full extraction. Adding --json triggers per-page segmentation:

pdf2md --json quarterly-report.pdf > report.json

The resulting JSON structure wraps each page as an independent object:

{
  "pages": [
    {
      "page_number": 1,
      "markdown": "# Q3 Financial Results\n\nRevenue increased by 15%...",

      "tables": [],
      "links": [{"url": "https://example.com", "text": "Investor Relations"}]
    },
    {
      "page_number": 2,
      "markdown": "## Detailed Breakdown\n\n| Region | Q3 Sales | Growth |\n|---|---|---|\n| North America | $12.4M | +18% |",

      "tables": [{"rows": 5, "columns": 3, "markdown": "..."}],
      "links": []
    }
  ]
}

This structure is intentionally designed for downstream consumption: each markdown field contains complete, page-isolated content suitable for direct LLM ingestion or UI rendering.

Extracting Specific Page Ranges

For targeted extraction, use the --pages flag:

pdf2md --json --pages 10-15 contract.pdf > pages-10-15.json

Only the specified pages traverse the full pipeline, reducing processing time for large documents.

Programmatic Per-Page Markdown Access

Rust Library API

Call process_pdf_with_options directly from src/lib.rs for embedded use:

use pdf_inspector::{process_pdf_with_options, ProcessOptions};

let opts = ProcessOptions {
    json_output: true,
    page_range: Some("1,3,5-10".to_string()),
    ..Default::default()
};

let result = process_pdf_with_options("document.pdf", opts)?;
// result contains the same JSON structure as CLI --json output

Python Bindings

The PyO3-based Python interface in src/python.rs mirrors the CLI options:

from pdf_inspector import process_pdf

result = process_pdf(
    "document.pdf",
    json_output=True,      # Enables per-page JSON structure

    page_range="1-5"       # Optional: subset of pages

)

# Access individual page markdown

for page in result["pages"]:
    print(f"Page {page['page_number']}: {len(page['markdown'])} chars")
    if page["tables"]:
        print(f"  Contains {len(page['tables'])} table(s)")

Layout Analysis That Preserves Per-Page Integrity

Column Detection and Reading Order

The layout engine in src/extractor/layout.rs uses horizontal projection histograms with pre-masking to identify column structures. It classifies documents into two reading modes:

  • Newspaper layout: Multiple independent columns read sequentially down each column
  • Tabular layout: Y-interleaved rows across columns to preserve logical row order

This analysis runs per-page, ensuring the extracted markdown field reflects the correct visual reading sequence for that specific page.

Table Detection Across Pages

Three detectors execute in priority order for each page (src/tables/detect_rects.rs, src/tables/detect_lines.rs, src/tables/detect_heuristic.rs). When tables span page boundaries, each page's tables array contains only the portion present on that page—maintaining the per-page isolation guarantee.

Handling Complex PDF Types Per Page

The detector in src/detector.rs runs independently on each page to classify:

  • Tiled scans: Multi-image scans assembled into single pages
  • Mixed PDFs: Pages combining text and images requiring selective OCR
  • Low-text-quality PDFs: Content that needs OCR fallback

This per-page classification means a single document can mix processing strategies: page 1 might use direct text extraction while page 2 triggers OCR, with both contributing clean Markdown to their respective JSON entries.

Unicode and Post-Processing Guarantees

Before Markdown reaches the JSON output, several cleanup stages ensure quality:

These steps run per-page, so each markdown string is immediately usable without additional cleanup.

Summary

  • pdf2md --json produces structured JSON with isolated per-page Markdown, tables, and links—ideal for LLM workflows and paginated UIs
  • Page ranges (--pages) enable efficient partial document processing without full pipeline overhead
  • Rust API (process_pdf_with_options in src/lib.rs) and Python bindings (src/python.rs) provide programmatic access with identical output structure
  • Per-page layout analysis (columns, reading order, tables) ensures each page's Markdown reflects correct visual structure
  • Detection and cleanup pipelines operate at page granularity, supporting mixed-content documents

Frequently Asked Questions

How do I get Markdown separated by page instead of one continuous file?

Use pdf2md --json your.pdf. This outputs JSON where each page has its own markdown field, rather than concatenating all pages into a single stream. Access individual pages via result["pages"][n]["markdown"] in Python or equivalent in other languages.

Does pdf-inspector handle multi-column layouts correctly per page?

Yes. The layout engine in src/extractor/layout.rs detects columns via horizontal projection histograms and determines proper reading order for each page independently. Newspaper-style columns are read sequentially; tabular layouts use Y-interleaving to preserve row relationships.

What happens to tables that span multiple pages?

Tables spanning pages are split at page boundaries. Each page's JSON entry contains the table portion present on that page in its tables array. The src/tables/format.rs module handles cell spanning within-page; cross-page table reconstruction requires downstream logic.

Can I extract only specific pages without processing the entire PDF?

Yes. Both CLI (--pages 5-10) and programmatic API (page_range: Some("5-10".to_string())) accept page specifications. The pipeline skips non-selected pages entirely, improving performance for large documents.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →