# Firecrawl PDF-Inspector: 10 Core Features for Rust & Python PDF Analysis

> Explore Firecrawl PDF-Inspector's 10 core features for robust PDF analysis. Get fast text extraction, Markdown conversion, and layout analysis with this Rust and Python library.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: deep-dive
- Published: 2026-08-07

---

**Firecrawl PDF-Inspector is a Rust library with optional Python bindings that provides fast, high-quality PDF type detection, full-text extraction, Markdown conversion, layout analysis, and table detection through a single-pass processing pipeline.**

Firecrawl PDF-Inspector enables developers to analyze, classify, and convert PDF documents to clean Markdown without relying on external OCR services by default. The library processes PDFs in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) using modular components for detection, extraction, and formatting, with each feature exposed through both Rust and Python APIs.

## PDF Type Detection with OCR Signaling

The `detect_pdf` and `detect_pdf_type` functions in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) and [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) classify documents as **TextBased**, **Scanned**, **Mixed**, or other categories. This classification happens rapidly without full text extraction, making it ideal for pipeline routing decisions.

The detector analyzes page content streams and font usage patterns to flag pages requiring external OCR. Detection results include per-page `needs_ocr` flags and machine-readable **OCR reasoning codes** defined as constants in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs):

- `OCR_REASON_SCANNED` — image-based pages with no extractable text
- `suspected_garbled_text` — encoding anomalies detected
- `vector_text` — graphics-based text that may need specialized handling

```rust
use pdf_inspector::detect_pdf;

fn main() -> Result<(), pdf_inspector::PdfError> {
    let info = detect_pdf("scanned.pdf")?;
    println!("Detected type: {:?}, pages: {}", info.pdf_type, info.page_count);
    Ok(())
}

```

The `detect_pdf` function wraps `process_pdf_with_options` with `PdfOptions::detect_only()` for minimal overhead.

## Full-Text Extraction with Font Handling

The extractor module at [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) parses PDF content streams directly, handling **ToUnicode CMaps** for proper character mapping and falling back to **TrueType font analysis** when CMap data is missing or incomplete. This dual-strategy approach maximizes text recovery from malformed or complex PDFs.

Extraction functions in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) return positioned text items with coordinates, enabling downstream layout-aware processing:

```rust
use pdf_inspector::{process_pdf, PdfOptions};

fn main() -> Result<(), pdf_inspector::PdfError> {
    let result = process_pdf("example.pdf")?;
    if let Some(md) = result.markdown {
        println!("{}", md);
    }
    Ok(())
}

```

## Markdown Conversion Pipeline

The `to_markdown*` functions in [`src/markdown/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/mod.rs) transform extracted text items into **token-efficient Markdown** through several preprocessing stages:

1. **Deduplication** — removes overlapping text from multiple content streams
2. **Reading order sorting** — reconstructs logical document flow
3. **Whitespace normalization** — produces clean, compact output

The module supports optional **structured output modes** for downstream consumers requiring parsed document elements rather than raw strings.

## Layout Analysis and Complexity Scoring

The `compute_layout_complexity` functions in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) (calling into `markdown::analysis`) detect:

- **Multi-column layouts** — newspaper and magazine formatting
- **Tabular reading order** — grid-based content organization
- **Complex structural pages** — mixed layouts requiring special handling

This analysis enables conditional processing paths: simple single-column documents proceed through fast extraction, while complex layouts trigger additional structural analysis.

## Three-Stage Table Detection

Firecrawl PDF-Inspector implements a **progressive fallback strategy** for table extraction across three specialized detectors in `src/tables/`:

| Stage | Implementation | Strategy | Fallback Trigger |
|-------|---------------|----------|----------------|
| 1 | [`detect_rects.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/detect_rects.rs) | Rectangle-based geometry detection | No ruling rectangles found |
| 2 | [`detect_lines.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/detect_lines.rs) | Vector line analysis | Insufficient line structure |
| 3 | [`detect_heuristic.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/detect_heuristic.rs) | Text-pattern heuristics | Previous stages fail quality checks |

The entry point `tables::detect_*` in [`src/tables/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/mod.rs) orchestrates this pipeline, producing **pipe-delimited Markdown tables** or flagging regions for OCR when all detection methods fail.

```rust
use pdf_inspector::extract_tables_in_regions_mem;

fn main() -> Result<(), pdf_inspector::PdfError> {
    let regions = vec![(0u32, vec![[100.0, 200.0, 400.0, 500.0]])];
    let tables = extract_tables_in_regions_mem(&std::fs::read("report.pdf")?, &regions)?;
    for page in tables {
        for region in page.regions {
            println!("{}", region.text);
        }
    }
    Ok(())
}

```

## Region-Based Extraction for Hybrid OCR Pipelines

The `extract_text_in_regions_mem` and `extract_tables_in_regions_mem` functions in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) enable **targeted extraction from arbitrary bounding boxes**. This supports hybrid architectures where:

- Firecrawl PDF-Inspector handles structured, extractable regions
- External OCR services process only flagged image areas

Regions are specified as `(page_index, bounding_boxes)` tuples with coordinates in PDF points.

## Vector Grid Detection

The `detect_vector_grid_in_region_mem` function in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) identifies **ruled-line and rectangle grids** within specified regions. This enables **TSR-compatible** (Table Structure Recognition) extraction for financial reports, forms, and technical documents with explicit table ruling.

## Text Quality and Encoding Issue Detection

The [`src/text_quality.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_quality.rs) module provides functions to detect:

- **Broken font encodings** — mismatched character-to-glyph mappings
- **CID garbage** — corrupted character identifier sequences
- **Anomalous text patterns** — statistical outliers indicating extraction failures

These checks feed into the OCR reasoning system, ensuring pages with recoverable quality issues are flagged appropriately rather than silently producing garbage output.

## Python Bindings via N-API

The `napi` crate exposes the full Rust API to Python through [`src/python.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs) and [`napi/src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/napi/src/lib.rs). Python users access identical functionality without Rust compilation:

```python
import pdf_inspector

result = pdf_inspector.process_pdf("sample.pdf")
print(result["markdown"])

```

Bindings are distributed as prebuilt wheels, eliminating the Rust toolchain requirement for Python deployments.

## Single-Pass Processing Architecture

Firecrawl PDF-Inspector optimizes for **single-pass document processing**:

1. Load PDF structure once
2. Classify document type
3. Extract and analyze content
4. Generate Markdown with metadata
5. Return OCR flags and layout information

This architecture minimizes I/O overhead and memory usage compared to multi-tool pipelines.

## Summary

Firecrawl PDF-Inspector provides **10 core capabilities** for PDF analysis:

- **PDF type detection** (`detect_pdf`, `detect_pdf_type`) with fast classification and OCR signaling
- **Full-text extraction** (`extract_text*`) with ToUnicode CMap and TrueType fallback handling
- **Markdown conversion** (`to_markdown*`) producing clean, token-efficient output
- **Layout analysis** (`compute_layout_complexity`) detecting multi-column and complex structures
- **Table detection** via three-stage rectangle → line → heuristic progressive fallback
- **Region-based extraction** (`extract_text_in_regions_mem`, `extract_tables_in_regions_mem`) for hybrid OCR integration
- **Vector grid detection** (`detect_vector_grid_in_region_mem`) for TSR-compatible table extraction
- **OCR reasoning** with machine-readable flags (`OCR_REASON_SCANNED`, `suspected_garbled_text`, `vector_text`)
- **Encoding issue detection** in [`src/text_quality.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_quality.rs) for quality assurance
- **Python bindings** through N-API exposing identical functionality to Python consumers

## Frequently Asked Questions

### What makes Firecrawl PDF-Inspector different from other PDF libraries?

Firecrawl PDF-Inspector focuses on **intelligent classification and OCR reasoning** rather than extraction alone. The [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) module analyzes documents before heavy processing, enabling pipeline decisions that route scanned documents to OCR services while processing text-based PDFs locally. This design minimizes unnecessary external API calls and reduces processing costs.

### Does Firecrawl PDF-Inspector perform OCR itself?

No. The library **detects when OCR is needed** through `detect_pdf` and per-page `needs_ocr` flags, but delegates actual character recognition to external services. The `extract_pages_markdown` function in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) emits structured OCR reasons (`OCR_REASON_SCANNED`, `suspected_garbled_text`, `vector_text`) that downstream systems use to route pages appropriately.

### How accurate is the table detection?

Table detection uses a **three-stage progressive strategy** in `src/tables/`: rectangle-based detection for ruled tables, line-based detection for vector-ruled tables, and heuristic detection for whitespace-delimited tables. Each stage validates output against quality metrics, falling back to the next method or OCR flagging when results are insufficient. This multi-method approach handles diverse table constructions found in financial reports, academic papers, and government documents.

### Can I extract content from specific regions only?

Yes. The `extract_text_in_regions_mem` and `extract_tables_in_regions_mem` functions in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) accept bounding box specifications per page. This enables **targeted extraction** for workflows where only document sections are relevant, or where hybrid OCR pipelines process different regions with different tools.