# How to Integrate firecrawl/pdf-inspector into a Node.js Project

> Integrate firecrawl pdf-inspector into your Node.js app. Extract Markdown, detect PDF types, and parse tables easily with our async functions. Get started today.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-07

---

**Install the `@firecrawl/pdf-inspector` npm package and use its async functions like `processPdf()` to extract Markdown, detect PDF types, and parse tables directly in your Node.js application.**

The **firecrawl/pdf-inspector** repository provides a Rust-based PDF processing engine with native Node.js bindings via N-API. This integration lets JavaScript developers leverage high-performance PDF extraction—including layout-aware Markdown generation and table detection—without managing external services.

## Overview: How the Node.js Integration Works

The library exposes Rust functions through a N-API layer, compiled to a native `.node` binary. The workflow follows four layers:

| Layer | Responsibility | Key Source |
|-------|----------------|------------|
| **Rust Core** | PDF parsing, layout analysis, table clustering, Markdown generation | [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs), [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs), [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs) |
| **N-API Glue** | Bind Rust functions to Node.js, handle type conversions | [`napi/src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/napi/src/lib.rs) |
| **Node Wrapper** | Load native binary and re-export JavaScript-callable API | npm package entry point |
| **CLI Binaries** | Standalone debugging tools (`pdf2md`, `detect-pdf`) | [`src/bin/pdf2md.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs), [`src/bin/detect_pdf.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/detect_pdf.rs) |

According to the firecrawl/pdf-inspector source code, the public API defined in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) includes `process_pdf`, `detect_pdf`, `classify_pdf`, `extract_text`, and table extraction helpers. The N-API implementation in [`napi/src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/napi/src/lib.rs) wraps these with `#[napi]` macros, converting JavaScript `Buffer` values to Rust `&[u8]` and `Result` types to JS exceptions.

## Installation

Install from npm or directly from the GitHub repository:

```bash

# Published package

npm install @firecrawl/pdf-inspector

# Latest source

npm install git+https://github.com/firecrawl/pdf-inspector.git

```

The package automatically loads the appropriate native binary for your platform on `require()` or `import`.

## Core Methods for Node.js Integration

The library provides five primary async functions. All return **Promises** and accept file paths (strings) or URLs.

### processPdf() — Extract Markdown

**`processPdf(path, options?)`** converts PDFs to token-efficient Markdown, preserving structural elements like headings, lists, and tables.

```javascript
import { processPdf } from '@firecrawl/pdf-inspector';

const pdfPath = './documents/annual-report.pdf';

(async () => {
  try {
    // Returns Markdown string (set json: true for structured output)
    const markdown = await processPdf(pdfPath, { json: false });
    console.log(markdown);
  } catch (err) {
    console.error('PDF processing failed:', err.message);
  }
})();

```

The Markdown generation logic resides in [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs), which reconstructs document flow from positional text data extracted by [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs).

### detectPdf() — Classify PDF Type

**`detectPdf(path)`** identifies whether a PDF is text-based, scanned/image-based, or mixed—useful for routing documents to appropriate processing pipelines.

```javascript
import { detectPdf } from '@firecrawl/pdf-inspector';

(async () => {
  const result = await detectPdf('./uploads/document.pdf');
  
  // result.type: 'TextBased' | 'Scanned' | 'Mixed'
  console.log('Classification:', result.type);
  
  // Additional metadata available
  if (result.hasTextLayer) {
    console.log('Extractable text layer detected');
  }
})();

```

The detection algorithm in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) analyzes font objects, text operators, and image tiling patterns to determine PDF composition.

### extractTextWithPositions() — Raw Positional Data

**`extractTextWithPositions(path)`** returns text elements with bounding box coordinates for custom layout processing.

```javascript
import { extractTextWithPositions } from '@firecrawl/pdf-inspector';

(async () => {
  const items = await extractTextWithPositions('./invoices/batch-001.pdf');
  
  // Array of {text, x, y, width, height, page, fontName?}
  const headerItems = items.filter(i => 
    i.y < 100 && i.page === 1
  );
  
  console.log('Header text elements:', headerItems);
})();

```

### extractTables() — Structured Table Extraction

**`extractTables(path)`** identifies tabular regions and returns parsed row/column structures with spanning cell detection.

```javascript
import { extractTables } from '@firecrawl/pdf-inspector';

(async () => {
  const tables = await extractTables('./financials/q3-data.pdf');
  
  for (const table of tables) {
    console.log(`Table on page ${table.page}: ${table.rows.length} rows`);
    // Access structured data: table.cells[row][col] = {text, rowSpan, colSpan}
  }
})();

```

Table extraction uses clustering algorithms on positional text data to identify grid structures without relying on explicit PDF table markup.

### classifyPdf() — Content Categorization

**`classifyPdf(path)`** provides document-level classification beyond the basic type detection.

```javascript
import { classifyPdf } from '@firecrawl/pdf-inspector';

(async () => {
  const classification = await classifyPdf('./unknown.pdf');
  console.log('Document category:', classification.category);
  console.log('Confidence:', classification.confidence);
})();

```

## Complete Server Integration Example

Use firecrawl/pdf-inspector in an Express.js API endpoint for PDF-to-Markdown conversion:

```javascript
import express from 'express';
import { processPdf, detectPdf } from '@firecrawl/pdf-inspector';
import { mkdirSync, writeFileSync } from 'fs';
import { join } from 'path';
import { tmpdir } from 'os';

const app = express();
app.use(express.json({ limit: '50mb' }));

// Convert PDF URL or base64 to Markdown
app.post('/convert', async (req, res) => {
  const { url, base64Pdf, options = {} } = req.body;
  
  try {
    let pdfPath;
    
    if (base64Pdf) {
      // Decode base64 to temp file
      const buffer = Buffer.from(base64Pdf, 'base64');
      pdfPath = join(tmpdir(), `pdf-${Date.now()}.pdf`);
      writeFileSync(pdfPath, buffer);
    } else if (url) {
      pdfPath = url; // Supports HTTP(S) URLs or local paths
    } else {
      return res.status(400).json({ error: 'Provide url or base64Pdf' });
    }
    
    // Optional: pre-detect PDF type for logging
    const detection = await detectPdf(pdfPath);
    console.log(`Processing ${detection.type} PDF`);
    
    // Extract Markdown
    const markdown = await processPdf(pdfPath, {
      json: options.json || false,
      preserveLineBreaks: options.preserveLineBreaks ?? true
    });
    
    res.json({
      success: true,
      pdfType: detection.type,
      markdown,
      charCount: markdown.length
    });
    
  } catch (err) {
    res.status(500).json({ 
      error: err.message,
      code: err.code || 'PROCESSING_FAILED'
    });
  }
});

app.listen(3000, () => {
  console.log('PDF inspector API on http://localhost:3000');
});

```

## Error Handling Patterns

Rust errors propagate as JavaScript `Error` objects with descriptive messages. Handle specific failure modes:

```javascript
import { processPdf } from '@firecrawl/pdf-inspector';

async function safeProcessPdf(path) {
  try {
    return await processPdf(path);
  } catch (err) {
    // Common error patterns from src/lib.rs Result handling
    if (err.message.includes('password')) {
      throw new Error('PDF requires password - encrypted documents not supported');
    }
    if (err.message.includes('malformed')) {
      throw new Error('Corrupted PDF file - cannot parse structure');
    }
    if (err.message.includes('not found')) {
      throw new Error(`File not found: ${path}`);
    }
    throw err; // Re-throw unexpected errors
  }
}

```

## Source File Reference

| File | Purpose | Implementation |
|------|---------|----------------|
| [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) | Public API entry points: `process_pdf`, `detect_pdf`, `extract_text`, `extract_tables` | Core orchestration |
| [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) | PDF type classification, tiled-scan heuristics | Detection logic |
| [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) | Text extraction pipeline, font analysis | Extraction engine |
| [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs) | Markdown serialization from extracted elements | Output formatting |
| [`napi/src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/napi/src/lib.rs) | N-API bindings with `#[napi]` macros | Node.js interface |
| [`src/bin/pdf2md.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs) | CLI tool for manual testing | Standalone utility |

## Summary

- **Install** with `npm install @firecrawl/pdf-inspector` to get pre-built native binaries
- **Import** async functions: `processPdf`, `detectPdf`, `extractTextWithPositions`, `extractTables`, `classifyPdf`
- **Handle** all operations with `async/await`—the N-API layer automatically converts Rust Results to JS Promises
- **Deploy** in server environments with standard error handling; the Rust core handles PDF security sandboxing
- **Reference** [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) and [`napi/src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/napi/src/lib.rs) for API contracts and type conversions

## Frequently Asked Questions

### Where does the actual PDF processing happen in the Node.js integration?

The heavy processing occurs in the Rust core ([`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs), [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs)), compiled to a native Node addon. The JavaScript layer in [`napi/src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/napi/src/lib.rs) only handles data marshalling—converting JS strings/Buffers to Rust types and back—so performance matches native Rust speeds.

### Can I use firecrawl/pdf-inspector with CommonJS require()?

Yes. The npm package exports both ESM (`import`) and CommonJS (`require`) targets. Use `const { processPdf } = require('@firecrawl/pdf-inspector')` if your project doesn't support ESM, though ESM is recommended for better tree-shaking.

### What PDF features are not supported by the Node.js binding?

Per the source in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs), password-encrypted PDFs require decryption before processing—the library does not handle password prompts. Additionally, JavaScript-executing PDFs (XFA forms, JavaScript actions) are executed; the library extracts static content only.

### How do I debug extraction quality issues?

Use the CLI binaries built from [`src/bin/pdf2md.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs) and [`src/bin/detect_pdf.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/detect_pdf.rs) to test files outside your Node application. Run `cargo run --bin pdf2md -- /path/to/file.pdf` from the repository root to see raw Rust output, isolating whether issues stem from the core engine or the Node.js binding layer.