Firecrawl PDF-Inspector: 10 Core Features for Rust & Python PDF Analysis
Firecrawl PDF-Inspector is a Rust library with optional Python bindings that provides fast, high-quality PDF type detection, full-text extraction, Markdown conversion, layout analysis, and table detection through a single-pass processing pipeline.
Firecrawl PDF-Inspector enables developers to analyze, classify, and convert PDF documents to clean Markdown without relying on external OCR services by default. The library processes PDFs in src/lib.rs using modular components for detection, extraction, and formatting, with each feature exposed through both Rust and Python APIs.
PDF Type Detection with OCR Signaling
The detect_pdf and detect_pdf_type functions in src/lib.rs and src/detector.rs classify documents as TextBased, Scanned, Mixed, or other categories. This classification happens rapidly without full text extraction, making it ideal for pipeline routing decisions.
The detector analyzes page content streams and font usage patterns to flag pages requiring external OCR. Detection results include per-page needs_ocr flags and machine-readable OCR reasoning codes defined as constants in src/lib.rs:
OCR_REASON_SCANNED— image-based pages with no extractable textsuspected_garbled_text— encoding anomalies detectedvector_text— graphics-based text that may need specialized handling
use pdf_inspector::detect_pdf;
fn main() -> Result<(), pdf_inspector::PdfError> {
let info = detect_pdf("scanned.pdf")?;
println!("Detected type: {:?}, pages: {}", info.pdf_type, info.page_count);
Ok(())
}
The detect_pdf function wraps process_pdf_with_options with PdfOptions::detect_only() for minimal overhead.
Full-Text Extraction with Font Handling
The extractor module at src/extractor/mod.rs parses PDF content streams directly, handling ToUnicode CMaps for proper character mapping and falling back to TrueType font analysis when CMap data is missing or incomplete. This dual-strategy approach maximizes text recovery from malformed or complex PDFs.
Extraction functions in src/lib.rs return positioned text items with coordinates, enabling downstream layout-aware processing:
use pdf_inspector::{process_pdf, PdfOptions};
fn main() -> Result<(), pdf_inspector::PdfError> {
let result = process_pdf("example.pdf")?;
if let Some(md) = result.markdown {
println!("{}", md);
}
Ok(())
}
Markdown Conversion Pipeline
The to_markdown* functions in src/markdown/mod.rs transform extracted text items into token-efficient Markdown through several preprocessing stages:
- Deduplication — removes overlapping text from multiple content streams
- Reading order sorting — reconstructs logical document flow
- Whitespace normalization — produces clean, compact output
The module supports optional structured output modes for downstream consumers requiring parsed document elements rather than raw strings.
Layout Analysis and Complexity Scoring
The compute_layout_complexity functions in src/lib.rs (calling into markdown::analysis) detect:
- Multi-column layouts — newspaper and magazine formatting
- Tabular reading order — grid-based content organization
- Complex structural pages — mixed layouts requiring special handling
This analysis enables conditional processing paths: simple single-column documents proceed through fast extraction, while complex layouts trigger additional structural analysis.
Three-Stage Table Detection
Firecrawl PDF-Inspector implements a progressive fallback strategy for table extraction across three specialized detectors in src/tables/:
| Stage | Implementation | Strategy | Fallback Trigger |
|---|---|---|---|
| 1 | detect_rects.rs |
Rectangle-based geometry detection | No ruling rectangles found |
| 2 | detect_lines.rs |
Vector line analysis | Insufficient line structure |
| 3 | detect_heuristic.rs |
Text-pattern heuristics | Previous stages fail quality checks |
The entry point tables::detect_* in src/tables/mod.rs orchestrates this pipeline, producing pipe-delimited Markdown tables or flagging regions for OCR when all detection methods fail.
use pdf_inspector::extract_tables_in_regions_mem;
fn main() -> Result<(), pdf_inspector::PdfError> {
let regions = vec![(0u32, vec![[100.0, 200.0, 400.0, 500.0]])];
let tables = extract_tables_in_regions_mem(&std::fs::read("report.pdf")?, ®ions)?;
for page in tables {
for region in page.regions {
println!("{}", region.text);
}
}
Ok(())
}
Region-Based Extraction for Hybrid OCR Pipelines
The extract_text_in_regions_mem and extract_tables_in_regions_mem functions in src/lib.rs enable targeted extraction from arbitrary bounding boxes. This supports hybrid architectures where:
- Firecrawl PDF-Inspector handles structured, extractable regions
- External OCR services process only flagged image areas
Regions are specified as (page_index, bounding_boxes) tuples with coordinates in PDF points.
Vector Grid Detection
The detect_vector_grid_in_region_mem function in src/lib.rs identifies ruled-line and rectangle grids within specified regions. This enables TSR-compatible (Table Structure Recognition) extraction for financial reports, forms, and technical documents with explicit table ruling.
Text Quality and Encoding Issue Detection
The src/text_quality.rs module provides functions to detect:
- Broken font encodings — mismatched character-to-glyph mappings
- CID garbage — corrupted character identifier sequences
- Anomalous text patterns — statistical outliers indicating extraction failures
These checks feed into the OCR reasoning system, ensuring pages with recoverable quality issues are flagged appropriately rather than silently producing garbage output.
Python Bindings via N-API
The napi crate exposes the full Rust API to Python through src/python.rs and napi/src/lib.rs. Python users access identical functionality without Rust compilation:
import pdf_inspector
result = pdf_inspector.process_pdf("sample.pdf")
print(result["markdown"])
Bindings are distributed as prebuilt wheels, eliminating the Rust toolchain requirement for Python deployments.
Single-Pass Processing Architecture
Firecrawl PDF-Inspector optimizes for single-pass document processing:
- Load PDF structure once
- Classify document type
- Extract and analyze content
- Generate Markdown with metadata
- Return OCR flags and layout information
This architecture minimizes I/O overhead and memory usage compared to multi-tool pipelines.
Summary
Firecrawl PDF-Inspector provides 10 core capabilities for PDF analysis:
- PDF type detection (
detect_pdf,detect_pdf_type) with fast classification and OCR signaling - Full-text extraction (
extract_text*) with ToUnicode CMap and TrueType fallback handling - Markdown conversion (
to_markdown*) producing clean, token-efficient output - Layout analysis (
compute_layout_complexity) detecting multi-column and complex structures - Table detection via three-stage rectangle → line → heuristic progressive fallback
- Region-based extraction (
extract_text_in_regions_mem,extract_tables_in_regions_mem) for hybrid OCR integration - Vector grid detection (
detect_vector_grid_in_region_mem) for TSR-compatible table extraction - OCR reasoning with machine-readable flags (
OCR_REASON_SCANNED,suspected_garbled_text,vector_text) - Encoding issue detection in
src/text_quality.rsfor quality assurance - Python bindings through N-API exposing identical functionality to Python consumers
Frequently Asked Questions
What makes Firecrawl PDF-Inspector different from other PDF libraries?
Firecrawl PDF-Inspector focuses on intelligent classification and OCR reasoning rather than extraction alone. The src/detector.rs module analyzes documents before heavy processing, enabling pipeline decisions that route scanned documents to OCR services while processing text-based PDFs locally. This design minimizes unnecessary external API calls and reduces processing costs.
Does Firecrawl PDF-Inspector perform OCR itself?
No. The library detects when OCR is needed through detect_pdf and per-page needs_ocr flags, but delegates actual character recognition to external services. The extract_pages_markdown function in src/lib.rs emits structured OCR reasons (OCR_REASON_SCANNED, suspected_garbled_text, vector_text) that downstream systems use to route pages appropriately.
How accurate is the table detection?
Table detection uses a three-stage progressive strategy in src/tables/: rectangle-based detection for ruled tables, line-based detection for vector-ruled tables, and heuristic detection for whitespace-delimited tables. Each stage validates output against quality metrics, falling back to the next method or OCR flagging when results are insufficient. This multi-method approach handles diverse table constructions found in financial reports, academic papers, and government documents.
Can I extract content from specific regions only?
Yes. The extract_text_in_regions_mem and extract_tables_in_regions_mem functions in src/lib.rs accept bounding box specifications per page. This enables targeted extraction for workflows where only document sections are relevant, or where hybrid OCR pipelines process different regions with different tools.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →