How to Integrate firecrawl/pdf-inspector into a Node.js Project

Install the @firecrawl/pdf-inspector npm package and use its async functions like processPdf() to extract Markdown, detect PDF types, and parse tables directly in your Node.js application.

The firecrawl/pdf-inspector repository provides a Rust-based PDF processing engine with native Node.js bindings via N-API. This integration lets JavaScript developers leverage high-performance PDF extraction—including layout-aware Markdown generation and table detection—without managing external services.

Overview: How the Node.js Integration Works

The library exposes Rust functions through a N-API layer, compiled to a native .node binary. The workflow follows four layers:

Layer Responsibility Key Source
Rust Core PDF parsing, layout analysis, table clustering, Markdown generation src/lib.rs, src/detector.rs, src/markdown/convert.rs
N-API Glue Bind Rust functions to Node.js, handle type conversions napi/src/lib.rs
Node Wrapper Load native binary and re-export JavaScript-callable API npm package entry point
CLI Binaries Standalone debugging tools (pdf2md, detect-pdf) src/bin/pdf2md.rs, src/bin/detect_pdf.rs

According to the firecrawl/pdf-inspector source code, the public API defined in src/lib.rs includes process_pdf, detect_pdf, classify_pdf, extract_text, and table extraction helpers. The N-API implementation in napi/src/lib.rs wraps these with #[napi] macros, converting JavaScript Buffer values to Rust &[u8] and Result types to JS exceptions.

Installation

Install from npm or directly from the GitHub repository:


# Published package

npm install @firecrawl/pdf-inspector

# Latest source

npm install git+https://github.com/firecrawl/pdf-inspector.git

The package automatically loads the appropriate native binary for your platform on require() or import.

Core Methods for Node.js Integration

The library provides five primary async functions. All return Promises and accept file paths (strings) or URLs.

processPdf() — Extract Markdown

processPdf(path, options?) converts PDFs to token-efficient Markdown, preserving structural elements like headings, lists, and tables.

import { processPdf } from '@firecrawl/pdf-inspector';

const pdfPath = './documents/annual-report.pdf';

(async () => {
  try {
    // Returns Markdown string (set json: true for structured output)
    const markdown = await processPdf(pdfPath, { json: false });
    console.log(markdown);
  } catch (err) {
    console.error('PDF processing failed:', err.message);
  }
})();

The Markdown generation logic resides in src/markdown/convert.rs, which reconstructs document flow from positional text data extracted by src/extractor/mod.rs.

detectPdf() — Classify PDF Type

detectPdf(path) identifies whether a PDF is text-based, scanned/image-based, or mixed—useful for routing documents to appropriate processing pipelines.

import { detectPdf } from '@firecrawl/pdf-inspector';

(async () => {
  const result = await detectPdf('./uploads/document.pdf');
  
  // result.type: 'TextBased' | 'Scanned' | 'Mixed'
  console.log('Classification:', result.type);
  
  // Additional metadata available
  if (result.hasTextLayer) {
    console.log('Extractable text layer detected');
  }
})();

The detection algorithm in src/detector.rs analyzes font objects, text operators, and image tiling patterns to determine PDF composition.

extractTextWithPositions() — Raw Positional Data

extractTextWithPositions(path) returns text elements with bounding box coordinates for custom layout processing.

import { extractTextWithPositions } from '@firecrawl/pdf-inspector';

(async () => {
  const items = await extractTextWithPositions('./invoices/batch-001.pdf');
  
  // Array of {text, x, y, width, height, page, fontName?}
  const headerItems = items.filter(i => 
    i.y < 100 && i.page === 1
  );
  
  console.log('Header text elements:', headerItems);
})();

extractTables() — Structured Table Extraction

extractTables(path) identifies tabular regions and returns parsed row/column structures with spanning cell detection.

import { extractTables } from '@firecrawl/pdf-inspector';

(async () => {
  const tables = await extractTables('./financials/q3-data.pdf');
  
  for (const table of tables) {
    console.log(`Table on page ${table.page}: ${table.rows.length} rows`);
    // Access structured data: table.cells[row][col] = {text, rowSpan, colSpan}
  }
})();

Table extraction uses clustering algorithms on positional text data to identify grid structures without relying on explicit PDF table markup.

classifyPdf() — Content Categorization

classifyPdf(path) provides document-level classification beyond the basic type detection.

import { classifyPdf } from '@firecrawl/pdf-inspector';

(async () => {
  const classification = await classifyPdf('./unknown.pdf');
  console.log('Document category:', classification.category);
  console.log('Confidence:', classification.confidence);
})();

Complete Server Integration Example

Use firecrawl/pdf-inspector in an Express.js API endpoint for PDF-to-Markdown conversion:

import express from 'express';
import { processPdf, detectPdf } from '@firecrawl/pdf-inspector';
import { mkdirSync, writeFileSync } from 'fs';
import { join } from 'path';
import { tmpdir } from 'os';

const app = express();
app.use(express.json({ limit: '50mb' }));

// Convert PDF URL or base64 to Markdown
app.post('/convert', async (req, res) => {
  const { url, base64Pdf, options = {} } = req.body;
  
  try {
    let pdfPath;
    
    if (base64Pdf) {
      // Decode base64 to temp file
      const buffer = Buffer.from(base64Pdf, 'base64');
      pdfPath = join(tmpdir(), `pdf-${Date.now()}.pdf`);
      writeFileSync(pdfPath, buffer);
    } else if (url) {
      pdfPath = url; // Supports HTTP(S) URLs or local paths
    } else {
      return res.status(400).json({ error: 'Provide url or base64Pdf' });
    }
    
    // Optional: pre-detect PDF type for logging
    const detection = await detectPdf(pdfPath);
    console.log(`Processing ${detection.type} PDF`);
    
    // Extract Markdown
    const markdown = await processPdf(pdfPath, {
      json: options.json || false,
      preserveLineBreaks: options.preserveLineBreaks ?? true
    });
    
    res.json({
      success: true,
      pdfType: detection.type,
      markdown,
      charCount: markdown.length
    });
    
  } catch (err) {
    res.status(500).json({ 
      error: err.message,
      code: err.code || 'PROCESSING_FAILED'
    });
  }
});

app.listen(3000, () => {
  console.log('PDF inspector API on http://localhost:3000');
});

Error Handling Patterns

Rust errors propagate as JavaScript Error objects with descriptive messages. Handle specific failure modes:

import { processPdf } from '@firecrawl/pdf-inspector';

async function safeProcessPdf(path) {
  try {
    return await processPdf(path);
  } catch (err) {
    // Common error patterns from src/lib.rs Result handling
    if (err.message.includes('password')) {
      throw new Error('PDF requires password - encrypted documents not supported');
    }
    if (err.message.includes('malformed')) {
      throw new Error('Corrupted PDF file - cannot parse structure');
    }
    if (err.message.includes('not found')) {
      throw new Error(`File not found: ${path}`);
    }
    throw err; // Re-throw unexpected errors
  }
}

Source File Reference

File Purpose Implementation
src/lib.rs Public API entry points: process_pdf, detect_pdf, extract_text, extract_tables Core orchestration
src/detector.rs PDF type classification, tiled-scan heuristics Detection logic
src/extractor/mod.rs Text extraction pipeline, font analysis Extraction engine
src/markdown/convert.rs Markdown serialization from extracted elements Output formatting
napi/src/lib.rs N-API bindings with #[napi] macros Node.js interface
src/bin/pdf2md.rs CLI tool for manual testing Standalone utility

Summary

  • Install with npm install @firecrawl/pdf-inspector to get pre-built native binaries
  • Import async functions: processPdf, detectPdf, extractTextWithPositions, extractTables, classifyPdf
  • Handle all operations with async/await—the N-API layer automatically converts Rust Results to JS Promises
  • Deploy in server environments with standard error handling; the Rust core handles PDF security sandboxing
  • Reference src/lib.rs and napi/src/lib.rs for API contracts and type conversions

Frequently Asked Questions

Where does the actual PDF processing happen in the Node.js integration?

The heavy processing occurs in the Rust core (src/lib.rs, src/extractor/mod.rs), compiled to a native Node addon. The JavaScript layer in napi/src/lib.rs only handles data marshalling—converting JS strings/Buffers to Rust types and back—so performance matches native Rust speeds.

Can I use firecrawl/pdf-inspector with CommonJS require()?

Yes. The npm package exports both ESM (import) and CommonJS (require) targets. Use const { processPdf } = require('@firecrawl/pdf-inspector') if your project doesn't support ESM, though ESM is recommended for better tree-shaking.

What PDF features are not supported by the Node.js binding?

Per the source in src/lib.rs, password-encrypted PDFs require decryption before processing—the library does not handle password prompts. Additionally, JavaScript-executing PDFs (XFA forms, JavaScript actions) are executed; the library extracts static content only.

How do I debug extraction quality issues?

Use the CLI binaries built from src/bin/pdf2md.rs and src/bin/detect_pdf.rs to test files outside your Node application. Run cargo run --bin pdf2md -- /path/to/file.pdf from the repository root to see raw Rust output, isolating whether issues stem from the core engine or the Node.js binding layer.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →