pdf-inspector Output Formats: Markdown, JSON, and Positioned Text Items

pdf-inspector supports four output formats via the pdf2md binary—default Markdown, raw Markdown, structured JSON, and granular Items-JSON—plus plain text and JSON classification via the detect-pdf binary.

The firecrawl/pdf-inspector repository provides Rust-based command-line tools for PDF content extraction and analysis. Understanding the available pdf-inspector output formats is essential for integrating the tool into document processing pipelines, whether you need human-readable text for immediate consumption or structured data for programmatic manipulation.

pdf2md Output Formats

The pdf2md binary, implemented in src/bin/pdf2md.rs, handles PDF-to-Markdown conversion and supports four distinct output modes controlled via CLI flags.

Default Markdown (Human-Readable)

By default, pdf2md outputs formatted Markdown with optional CLI headers, providing clean, readable text suitable for immediate use or further Markdown processing. This mode prioritizes readability over machine parsing.

Raw Markdown (--raw)

The --raw flag suppresses CLI headers and metadata, outputting only the extracted Markdown content. This format is ideal for shell scripts that pipe content directly into downstream tools without parsing decorative output.

Structured JSON (--json)

When invoked with --json, the tool emits a comprehensive JSON object assembled in the main match block of src/bin/pdf2md.rs (lines 73-106). This payload includes fields such as pdf_type, page_count, processing_time_ms, and the extracted markdown string, plus OCR and layout metadata when available.

Items-JSON (--items-json)

The --items-json flag triggers the format_items_json function, which iterates over TextItem structures to produce granular, position-aware JSON. Each item includes precise coordinates, font information, style attributes, MCID (Marked Content Identifier), and link data. This format is essential for custom post-processing tasks like table reconstruction or aligning text to original PDF coordinates.

detect-pdf Output Formats

The detect-pdf binary, located in src/bin/detect_pdf.rs, classifies PDFs as TextBased, Scanned, Mixed, or ImageBased.

Plain Text Detection

Without additional flags, the binary prints the detected PDF type directly to stdout as simple text. This lightweight mode is suitable for quick classification checks in automation scripts.

JSON Detection and Analysis (--json, --analyze)

The --json flag outputs a structured JSON object containing pdf_type, page_count, and layout details. When combined with --analyze, the JSON payload includes comprehensive layout analysis data, enabling programmatic inspection of document structure without Markdown conversion.

Implementation Architecture

The output formats rely on core extraction functions defined in src/lib.rs, such as process_pdf_with_options and extract_text_with_positions_pages_with_password. The Markdown conversion logic resides in src/markdown/convert.rs, which transforms internal PdfPage structures into Markdown syntax. For Items-JSON, the extractor module (src/extractor/mod.rs) provides the underlying text item data including positions and styling attributes.

Command-Line Examples


# Convert PDF to Markdown with CLI header (default)

pdf2md sample.pdf

# Extract raw Markdown only, no headers

pdf2md sample.pdf --raw > output.md

# Get full extraction as JSON with metadata and timing

pdf2md sample.pdf --json > result.json

# Export positioned text items for custom table reconstruction

pdf2md sample.pdf --items-json > items.json

# Detect PDF type as plain text

detect-pdf sample.pdf

# Detect with structured JSON output

detect-pdf sample.pdf --json > detection.json

# Run layout analysis and output as JSON (no markdown)

detect-pdf sample.pdf --analyze --json > analysis.json

Summary

  • pdf-inspector provides two binaries: pdf2md for content extraction and detect-pdf for type classification and layout analysis.
  • Markdown outputs include default formatted text and --raw mode for header-free content suitable for piping.
  • JSON outputs range from full extraction results (--json) containing pdf_type, page_count, and processing_time_ms, to granular positioned text items (--items-json) with font and coordinate data.
  • Detection outputs support both plain text classification and structured JSON analysis via detect-pdf.
  • Core functionality is implemented across src/bin/pdf2md.rs, src/bin/detect_pdf.rs, and supporting extraction modules.

Frequently Asked Questions

What is the difference between --json and --items-json in pdf2md?

The --json flag outputs a complete extraction result including document metadata, processing statistics, and the full Markdown content as a string. In contrast, --items-json produces a line-by-line array of TextItem objects containing precise positioning data (coordinates), font details, style attributes, and link information, which is essential for reconstructing complex layouts programmatically without parsing Markdown.

Can I use pdf-inspector output formats without installing the CLI?

Yes. While the CLI binaries provide convenient access to the various formats, the core library exposes public APIs in src/lib.rs such as process_pdf_with_options and extract_text_with_positions_pages_with_password. These functions can be integrated directly into Rust applications, allowing programmatic access to the same Markdown, JSON, and item-level data without invoking shell commands.

Does detect-pdf support Markdown output?

No. The detect-pdf binary focuses exclusively on PDF classification and layout analysis. It outputs either plain text (the detected PDF type) or JSON (with --json or --analyze flags). For Markdown conversion, you must use the pdf2md binary, which is specifically designed for content extraction and format conversion.

Which output format preserves the exact position of text in the original PDF?

The Items-JSON format (--items-json) preserves exact positional data. According to the source code in src/bin/pdf2md.rs, the format_items_json function serializes each text item with its coordinates, font information, MCID, and link data, enabling precise alignment with the original PDF layout for applications requiring spatial accuracy or custom layout reconstruction.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →