# pdf-inspector Output Formats: Markdown, JSON, and Positioned Text Items

> Explore pdf-inspector output formats including Markdown, JSON, and positioned text items. Extract data efficiently with its versatile export options.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: api-reference
- Published: 2026-08-04

---

**pdf-inspector supports four output formats via the `pdf2md` binary—default Markdown, raw Markdown, structured JSON, and granular Items-JSON—plus plain text and JSON classification via the `detect-pdf` binary.**

The `firecrawl/pdf-inspector` repository provides Rust-based command-line tools for PDF content extraction and analysis. Understanding the available **pdf-inspector output formats** is essential for integrating the tool into document processing pipelines, whether you need human-readable text for immediate consumption or structured data for programmatic manipulation.

## pdf2md Output Formats

The `pdf2md` binary, implemented in [`src/bin/pdf2md.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs), handles PDF-to-Markdown conversion and supports four distinct output modes controlled via CLI flags.

### Default Markdown (Human-Readable)

By default, `pdf2md` outputs formatted Markdown with optional CLI headers, providing clean, readable text suitable for immediate use or further Markdown processing. This mode prioritizes readability over machine parsing.

### Raw Markdown (--raw)

The `--raw` flag suppresses CLI headers and metadata, outputting only the extracted Markdown content. This format is ideal for shell scripts that pipe content directly into downstream tools without parsing decorative output.

### Structured JSON (--json)

When invoked with `--json`, the tool emits a comprehensive JSON object assembled in the main match block of [`src/bin/pdf2md.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs) (lines 73-106). This payload includes fields such as `pdf_type`, `page_count`, `processing_time_ms`, and the extracted `markdown` string, plus OCR and layout metadata when available.

### Items-JSON (--items-json)

The `--items-json` flag triggers the `format_items_json` function, which iterates over `TextItem` structures to produce granular, position-aware JSON. Each item includes precise coordinates, font information, style attributes, MCID (Marked Content Identifier), and link data. This format is essential for custom post-processing tasks like table reconstruction or aligning text to original PDF coordinates.

## detect-pdf Output Formats

The `detect-pdf` binary, located in [`src/bin/detect_pdf.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/detect_pdf.rs), classifies PDFs as TextBased, Scanned, Mixed, or ImageBased.

### Plain Text Detection

Without additional flags, the binary prints the detected PDF type directly to stdout as simple text. This lightweight mode is suitable for quick classification checks in automation scripts.

### JSON Detection and Analysis (--json, --analyze)

The `--json` flag outputs a structured JSON object containing `pdf_type`, `page_count`, and layout details. When combined with `--analyze`, the JSON payload includes comprehensive layout analysis data, enabling programmatic inspection of document structure without Markdown conversion.

## Implementation Architecture

The output formats rely on core extraction functions defined in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs), such as `process_pdf_with_options` and `extract_text_with_positions_pages_with_password`. The Markdown conversion logic resides in [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs), which transforms internal `PdfPage` structures into Markdown syntax. For Items-JSON, the extractor module ([`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs)) provides the underlying text item data including positions and styling attributes.

## Command-Line Examples

```bash

# Convert PDF to Markdown with CLI header (default)

pdf2md sample.pdf

# Extract raw Markdown only, no headers

pdf2md sample.pdf --raw > output.md

# Get full extraction as JSON with metadata and timing

pdf2md sample.pdf --json > result.json

# Export positioned text items for custom table reconstruction

pdf2md sample.pdf --items-json > items.json

# Detect PDF type as plain text

detect-pdf sample.pdf

# Detect with structured JSON output

detect-pdf sample.pdf --json > detection.json

# Run layout analysis and output as JSON (no markdown)

detect-pdf sample.pdf --analyze --json > analysis.json

```

## Summary

- **pdf-inspector** provides two binaries: `pdf2md` for content extraction and `detect-pdf` for type classification and layout analysis.
- **Markdown outputs** include default formatted text and `--raw` mode for header-free content suitable for piping.
- **JSON outputs** range from full extraction results (`--json`) containing `pdf_type`, `page_count`, and `processing_time_ms`, to granular positioned text items (`--items-json`) with font and coordinate data.
- **Detection outputs** support both plain text classification and structured JSON analysis via `detect-pdf`.
- Core functionality is implemented across [`src/bin/pdf2md.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs), [`src/bin/detect_pdf.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/detect_pdf.rs), and supporting extraction modules.

## Frequently Asked Questions

### What is the difference between --json and --items-json in pdf2md?

The `--json` flag outputs a complete extraction result including document metadata, processing statistics, and the full Markdown content as a string. In contrast, `--items-json` produces a line-by-line array of `TextItem` objects containing precise positioning data (coordinates), font details, style attributes, and link information, which is essential for reconstructing complex layouts programmatically without parsing Markdown.

### Can I use pdf-inspector output formats without installing the CLI?

Yes. While the CLI binaries provide convenient access to the various formats, the core library exposes public APIs in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) such as `process_pdf_with_options` and `extract_text_with_positions_pages_with_password`. These functions can be integrated directly into Rust applications, allowing programmatic access to the same Markdown, JSON, and item-level data without invoking shell commands.

### Does detect-pdf support Markdown output?

No. The `detect-pdf` binary focuses exclusively on PDF classification and layout analysis. It outputs either plain text (the detected PDF type) or JSON (with `--json` or `--analyze` flags). For Markdown conversion, you must use the `pdf2md` binary, which is specifically designed for content extraction and format conversion.

### Which output format preserves the exact position of text in the original PDF?

The **Items-JSON** format (`--items-json`) preserves exact positional data. According to the source code in [`src/bin/pdf2md.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs), the `format_items_json` function serializes each text item with its coordinates, font information, MCID, and link data, enabling precise alignment with the original PDF layout for applications requiring spatial accuracy or custom layout reconstruction.