How to Use Firecrawl pdf‑inspector for Programmatic PDF Analysis
Yes—Firecrawl pdf‑inspector is purpose-built for programmatic PDF analysis, offering native Rust APIs with bindings for Python, Node.js, and WebAssembly.
This open‑source library exposes a consistent extraction and classification engine across all runtimes, making it ideal for automated workflows in scripts, backend services, or browser applications. The same core logic powers every interface, so your choice of language doesn't compromise analysis quality.
Programmatic Interfaces Overview
The firecrawl/pdf-inspector repository structures its capabilities into layered components that you can access through multiple entry points:
| Interface | Package Name | Best For |
|---|---|---|
| Rust (native) | pdf-inspector crate |
High‑performance systems, embedded logic |
| Python | pdf-inspector (PyPI) |
Data pipelines, ML workflows, scripting |
| Node.js | @firecrawl/pdf-inspector |
Serverless functions, API services |
| WebAssembly | @firecrawl/pdf-inspector-wasm |
Client‑side parsing, privacy‑first apps |
Each binding wraps the identical engine implemented in src/lib.rs, where key functions like process_pdf, detect_pdf, and extract_text orchestrate PDF loading, type classification, and content extraction.
Core API Functions
The public API in src/lib.rs provides three primary entry points for programmatic use:
process_pdf– Complete analysis pipeline returning PDF type, extracted text, markdown output, and layout structure.detect_pdf– Lightweight classification without full extraction; returnsPdfClassification(TextBased, Scanned, ImageBased, or Mixed).extract_text– Plain text extraction only, skipping markdown generation for faster processing.
Classification logic lives in src/detector.rs, which applies multiple scan strategies to categorize input files. The extraction pipeline in src/extractor/mod.rs handles content stream parsing, font decoding, and positional text grouping into TextItem and TextLine structures.
Python Integration
Install via pip and use synchronously:
import pdf_inspector
# Full analysis with markdown output
result = pdf_inspector.process_pdf("annual_report.pdf")
print(result.pdf_type) # "text_based", "scanned", etc.
print(result.markdown) # Structured markdown string
# Detection-only for routing logic
classification = pdf_inspector.detect_pdf("unknown.pdf")
if classification.is_text_based:
content = pdf_inspector.extract_text("unknown.pdf")
The Python package includes type stubs (pdf_inspector.pyi) for IDE autocomplete and static checking.
Node.js Integration
The N‑API binding supports both file paths and Buffer objects for serverless environments:
const { processPdf, detectPdf, extractText } = require("@firecrawl/pdf-inspector");
async function analyzeDocument(buffer) {
const result = await processPdf(buffer);
return {
type: result.pdf_type,
markdown: result.markdown,
pageCount: result.pages.length
};
}
See napi/README.md in the repository for async streaming patterns and error handling details.
WebAssembly Browser Usage
For client‑side parsing without server round‑trips:
<script type="module">
import init, { process_pdf } from "@firecrawl/pdf-inspector-wasm";
await init();
const response = await fetch("/document.pdf");
const result = process_pdf(new Uint8Array(await response.arrayBuffer()));
// result.markdown contains the structured output
document.getElementById("output").textContent = result.markdown;
</script>
The WASM build is distributed via CDN and documented in wasm/README.md.
Rust Direct Integration
For systems requiring maximum control or custom extraction pipelines:
use pdf_inspector::{process_pdf, PdfProcessResult};
fn analyze_upload(path: &str) -> Result<String, Box<dyn std::error::Error>> {
let result: PdfProcessResult = process_pdf(path)?;
match result.pdf_type.as_str() {
"text_based" => Ok(result.markdown.unwrap_or_default()),
"scanned" => Err("OCR required for scanned documents".into()),
_ => Ok(result.plain_text.unwrap_or_default())
}
}
Direct crate usage lets you configure extractor behavior and handle TextItem collections before markdown conversion in src/markdown/convert.rs.
Architecture Deep Dive
Understanding the internal structure helps optimize programmatic calls:
| Component | Source File | Responsibility |
|---|---|---|
| Detector | src/detector.rs |
Multi‑strategy PDF classification |
| Extractor | src/extractor/mod.rs |
Content parsing, font decoding, layout analysis |
| Markdown generator | src/markdown/convert.rs |
Token‑efficient markdown with headers, lists, tables |
| Utilities | src/tounicode.rs, src/text_utils.rs |
Unicode, CJK, RTL text handling |
The markdown generator specifically optimizes for LLM context windows by producing clean structural markers rather than visual formatting.
Performance Considerations
For programmatic workflows at scale:
- Use
detect_pdffirst to route documents—it's significantly faster than full extraction. - Prefer
extract_textoverprocess_pdf** when you don't need markdown or layout data. - Buffer reuse in Node.js and WASM reduces memory allocations for batch processing.
- Rust native offers zero‑copy access to internal structures; use it for sub‑millisecond latency requirements.
Summary
- Firecrawl pdf‑inspector exposes identical core functionality through Rust, Python, Node.js, and WebAssembly interfaces.
- Entry functions
process_pdf,detect_pdf, andextract_textinsrc/lib.rsprovide tiered analysis depth for different use cases. - Python and Node.js bindings are thin wrappers; WASM enables browser‑native parsing without server dependencies.
- The layered architecture (detector → extractor → markdown generator) ensures consistent results across all runtimes.
Frequently Asked Questions
Does pdf‑inspector require a running Firecrawl service?
No. The library operates entirely offline with no external API calls. All processing happens locally using the Rust core engine.
Can I use pdf‑inspector in a serverless environment?
Yes. The Node.js package works in AWS Lambda, Vercel Edge Functions, and Cloudflare Workers. The WASM build runs in Deno and WinterCG‑compatible runtimes.
What PDF types can the detector identify?
The detector in src/detector.rs classifies files as TextBased, Scanned, ImageBased, or Mixed. Scanned detection relies on text density analysis rather than OCR, so it flags documents requiring external OCR services without attempting extraction.
Is the Python binding pure Python or does it require Rust compilation?
The PyPI package distributes pre‑compiled wheels for common platforms. You only need a Rust toolchain if installing from source or targeting an unsupported architecture.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →