How to Integrate firecrawl/pdf-inspector into a Node.js Project
Install the @firecrawl/pdf-inspector npm package and use its async functions like processPdf() to extract Markdown, detect PDF types, and parse tables directly in your Node.js application.
The firecrawl/pdf-inspector repository provides a Rust-based PDF processing engine with native Node.js bindings via N-API. This integration lets JavaScript developers leverage high-performance PDF extraction—including layout-aware Markdown generation and table detection—without managing external services.
Overview: How the Node.js Integration Works
The library exposes Rust functions through a N-API layer, compiled to a native .node binary. The workflow follows four layers:
| Layer | Responsibility | Key Source |
|---|---|---|
| Rust Core | PDF parsing, layout analysis, table clustering, Markdown generation | src/lib.rs, src/detector.rs, src/markdown/convert.rs |
| N-API Glue | Bind Rust functions to Node.js, handle type conversions | napi/src/lib.rs |
| Node Wrapper | Load native binary and re-export JavaScript-callable API | npm package entry point |
| CLI Binaries | Standalone debugging tools (pdf2md, detect-pdf) |
src/bin/pdf2md.rs, src/bin/detect_pdf.rs |
According to the firecrawl/pdf-inspector source code, the public API defined in src/lib.rs includes process_pdf, detect_pdf, classify_pdf, extract_text, and table extraction helpers. The N-API implementation in napi/src/lib.rs wraps these with #[napi] macros, converting JavaScript Buffer values to Rust &[u8] and Result types to JS exceptions.
Installation
Install from npm or directly from the GitHub repository:
# Published package
npm install @firecrawl/pdf-inspector
# Latest source
npm install git+https://github.com/firecrawl/pdf-inspector.git
The package automatically loads the appropriate native binary for your platform on require() or import.
Core Methods for Node.js Integration
The library provides five primary async functions. All return Promises and accept file paths (strings) or URLs.
processPdf() — Extract Markdown
processPdf(path, options?) converts PDFs to token-efficient Markdown, preserving structural elements like headings, lists, and tables.
import { processPdf } from '@firecrawl/pdf-inspector';
const pdfPath = './documents/annual-report.pdf';
(async () => {
try {
// Returns Markdown string (set json: true for structured output)
const markdown = await processPdf(pdfPath, { json: false });
console.log(markdown);
} catch (err) {
console.error('PDF processing failed:', err.message);
}
})();
The Markdown generation logic resides in src/markdown/convert.rs, which reconstructs document flow from positional text data extracted by src/extractor/mod.rs.
detectPdf() — Classify PDF Type
detectPdf(path) identifies whether a PDF is text-based, scanned/image-based, or mixed—useful for routing documents to appropriate processing pipelines.
import { detectPdf } from '@firecrawl/pdf-inspector';
(async () => {
const result = await detectPdf('./uploads/document.pdf');
// result.type: 'TextBased' | 'Scanned' | 'Mixed'
console.log('Classification:', result.type);
// Additional metadata available
if (result.hasTextLayer) {
console.log('Extractable text layer detected');
}
})();
The detection algorithm in src/detector.rs analyzes font objects, text operators, and image tiling patterns to determine PDF composition.
extractTextWithPositions() — Raw Positional Data
extractTextWithPositions(path) returns text elements with bounding box coordinates for custom layout processing.
import { extractTextWithPositions } from '@firecrawl/pdf-inspector';
(async () => {
const items = await extractTextWithPositions('./invoices/batch-001.pdf');
// Array of {text, x, y, width, height, page, fontName?}
const headerItems = items.filter(i =>
i.y < 100 && i.page === 1
);
console.log('Header text elements:', headerItems);
})();
extractTables() — Structured Table Extraction
extractTables(path) identifies tabular regions and returns parsed row/column structures with spanning cell detection.
import { extractTables } from '@firecrawl/pdf-inspector';
(async () => {
const tables = await extractTables('./financials/q3-data.pdf');
for (const table of tables) {
console.log(`Table on page ${table.page}: ${table.rows.length} rows`);
// Access structured data: table.cells[row][col] = {text, rowSpan, colSpan}
}
})();
Table extraction uses clustering algorithms on positional text data to identify grid structures without relying on explicit PDF table markup.
classifyPdf() — Content Categorization
classifyPdf(path) provides document-level classification beyond the basic type detection.
import { classifyPdf } from '@firecrawl/pdf-inspector';
(async () => {
const classification = await classifyPdf('./unknown.pdf');
console.log('Document category:', classification.category);
console.log('Confidence:', classification.confidence);
})();
Complete Server Integration Example
Use firecrawl/pdf-inspector in an Express.js API endpoint for PDF-to-Markdown conversion:
import express from 'express';
import { processPdf, detectPdf } from '@firecrawl/pdf-inspector';
import { mkdirSync, writeFileSync } from 'fs';
import { join } from 'path';
import { tmpdir } from 'os';
const app = express();
app.use(express.json({ limit: '50mb' }));
// Convert PDF URL or base64 to Markdown
app.post('/convert', async (req, res) => {
const { url, base64Pdf, options = {} } = req.body;
try {
let pdfPath;
if (base64Pdf) {
// Decode base64 to temp file
const buffer = Buffer.from(base64Pdf, 'base64');
pdfPath = join(tmpdir(), `pdf-${Date.now()}.pdf`);
writeFileSync(pdfPath, buffer);
} else if (url) {
pdfPath = url; // Supports HTTP(S) URLs or local paths
} else {
return res.status(400).json({ error: 'Provide url or base64Pdf' });
}
// Optional: pre-detect PDF type for logging
const detection = await detectPdf(pdfPath);
console.log(`Processing ${detection.type} PDF`);
// Extract Markdown
const markdown = await processPdf(pdfPath, {
json: options.json || false,
preserveLineBreaks: options.preserveLineBreaks ?? true
});
res.json({
success: true,
pdfType: detection.type,
markdown,
charCount: markdown.length
});
} catch (err) {
res.status(500).json({
error: err.message,
code: err.code || 'PROCESSING_FAILED'
});
}
});
app.listen(3000, () => {
console.log('PDF inspector API on http://localhost:3000');
});
Error Handling Patterns
Rust errors propagate as JavaScript Error objects with descriptive messages. Handle specific failure modes:
import { processPdf } from '@firecrawl/pdf-inspector';
async function safeProcessPdf(path) {
try {
return await processPdf(path);
} catch (err) {
// Common error patterns from src/lib.rs Result handling
if (err.message.includes('password')) {
throw new Error('PDF requires password - encrypted documents not supported');
}
if (err.message.includes('malformed')) {
throw new Error('Corrupted PDF file - cannot parse structure');
}
if (err.message.includes('not found')) {
throw new Error(`File not found: ${path}`);
}
throw err; // Re-throw unexpected errors
}
}
Source File Reference
| File | Purpose | Implementation |
|---|---|---|
src/lib.rs |
Public API entry points: process_pdf, detect_pdf, extract_text, extract_tables |
Core orchestration |
src/detector.rs |
PDF type classification, tiled-scan heuristics | Detection logic |
src/extractor/mod.rs |
Text extraction pipeline, font analysis | Extraction engine |
src/markdown/convert.rs |
Markdown serialization from extracted elements | Output formatting |
napi/src/lib.rs |
N-API bindings with #[napi] macros |
Node.js interface |
src/bin/pdf2md.rs |
CLI tool for manual testing | Standalone utility |
Summary
- Install with
npm install @firecrawl/pdf-inspectorto get pre-built native binaries - Import async functions:
processPdf,detectPdf,extractTextWithPositions,extractTables,classifyPdf - Handle all operations with
async/await—the N-API layer automatically converts Rust Results to JS Promises - Deploy in server environments with standard error handling; the Rust core handles PDF security sandboxing
- Reference
src/lib.rsandnapi/src/lib.rsfor API contracts and type conversions
Frequently Asked Questions
Where does the actual PDF processing happen in the Node.js integration?
The heavy processing occurs in the Rust core (src/lib.rs, src/extractor/mod.rs), compiled to a native Node addon. The JavaScript layer in napi/src/lib.rs only handles data marshalling—converting JS strings/Buffers to Rust types and back—so performance matches native Rust speeds.
Can I use firecrawl/pdf-inspector with CommonJS require()?
Yes. The npm package exports both ESM (import) and CommonJS (require) targets. Use const { processPdf } = require('@firecrawl/pdf-inspector') if your project doesn't support ESM, though ESM is recommended for better tree-shaking.
What PDF features are not supported by the Node.js binding?
Per the source in src/lib.rs, password-encrypted PDFs require decryption before processing—the library does not handle password prompts. Additionally, JavaScript-executing PDFs (XFA forms, JavaScript actions) are executed; the library extracts static content only.
How do I debug extraction quality issues?
Use the CLI binaries built from src/bin/pdf2md.rs and src/bin/detect_pdf.rs to test files outside your Node application. Run cargo run --bin pdf2md -- /path/to/file.pdf from the repository root to see raw Rust output, isolating whether issues stem from the core engine or the Node.js binding layer.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →