How to Integrate pdf-inspector into a Node.js Project: Complete Setup Guide

You can integrate pdf-inspector into a Node.js project by installing the @firecrawl/pdf-inspector npm package, which provides pre-built N-API binaries that expose Rust-based PDF classification and Markdown extraction functions without requiring a local Rust toolchain.

The firecrawl/pdf-inspector repository delivers a high-performance Rust core for PDF analysis wrapped in Node.js bindings. When you integrate pdf-inspector into a Node.js project, you gain access to sub-50ms PDF classification, full Markdown extraction, and region-specific text extraction while avoiding redundant I/O through single-load parsing.

Installation and Setup

Start by installing the package from npm. The distribution includes pre-built binaries for Linux x64/ARM64, macOS ARM64, and Windows x64, so no Rust compiler is required on the target machine.

npm install @firecrawl/pdf-inspector

After installation, import the module and read your PDF file into a Buffer:

import { readFileSync } from 'fs';
import { classifyPdf, processPdf } from '@firecrawl/pdf-inspector';

const pdf = readFileSync('path/to/document.pdf');

The package automatically selects the correct platform-specific binary during installation, as documented in napi/README.md【/cache/repos/github.com/firecrawl/pdf-inspector/main/napi/README.md†L37-L45】.

Classifying PDF Documents

Use the classifyPdf function to determine document type and OCR requirements before processing. This function, implemented via the detector logic in src/detector.rs, analyzes the PDF in approximately 10–50 milliseconds【/cache/repos/github.com/firecrawl/pdf-inspector/main/README.md†L14-L16】.

const classification = classifyPdf(pdf);

console.log(classification.pdfType);        // "TextBased" | "Scanned" | "Mixed" | "ImageBased"
console.log(classification.pageCount);      // total pages
console.log(classification.pagesNeedingOcr); // Array of page numbers

The PdfClassification object returns a confidence score and a pagesNeedingOcr array, allowing you to route only specific pages to OCR services rather than processing the entire document【/cache/repos/github.com/firecrawl/pdf-inspector/main/napi/README.md†L41-L56】.

Extracting Markdown Content

For complete document conversion, use processPdf. This function orchestrates the full pipeline defined in src/lib.rs: detection, extraction, table detection (via src/tables/detect_rects.rs and src/tables/detect_heuristic.rs), and Markdown generation (via src/markdown/convert.rs)【/cache/repos/github.com/firecrawl/pdf-inspector/main/README.md†L74-L78】【/cache/repos/github.com/firecrawl/pdf-inspector/main/README.md†L165-L173】.

const result = processPdf(pdf);

console.log(result.pdfType);   // e.g., "TextBased"
console.log(result.markdown);  // Clean Markdown string or null

Because the parser loads the PDF once and shares the internal representation between classification and extraction stages, calling processPdf avoids redundant I/O operations【/cache/repos/github.com/firecrawl/pdf-inspector/main/README.md†L65-L73】.

Region-Based Text Extraction

For hybrid OCR pipelines, use extractTextInRegions to extract text from specific rectangular areas. This function accepts an array of page regions and returns text segments with OCR necessity flags.

import { extractTextInRegions } from '@firecrawl/pdf-inspector';

const regions = [
  { page: 0, regions: [[0, 0, 300, 400], [300, 0, 612, 400]] },
];

const regionTexts = extractTextInRegions(pdf, regions);

for (const region of regionTexts[0].regions) {
  if (region.needsOcr) {
    // Send to OCR service based on region.ocrReason
  } else {
    console.log(region.text);
  }
}

Each RegionText object includes a needsOcr boolean and an ocrReason string (such as "suspected_garbled_text"), enabling precise OCR fallback only where the extractor detects unreliable text【/cache/repos/github.com/firecrawl/pdf-inspector/main/napi/README.md†L58-L83】.

Core Architecture and Source Files

Understanding the underlying Rust architecture helps optimize your integration. The public API exposed to Node.js resides in src/lib.rs, which coordinates several specialized modules:

The N-API bindings wrap these Rust modules to provide synchronous JavaScript functions that operate on Node.js Buffer objects【/cache/repos/github.com/firecrawl/pdf-inspector/main/README.md†L165-L173】.

Summary

  • Install @firecrawl/pdf-inspector to add pre-built Rust PDF processing capabilities to Node.js without compiling from source.
  • Use classifyPdf for rapid document classification (10–50ms) to determine OCR requirements before full extraction.
  • Call processPdf to extract complete Markdown representation while avoiding redundant file I/O through single-load parsing.
  • Implement extractTextInRegions for targeted text extraction when building hybrid OCR pipelines that process only specific document areas.
  • Reference source files including src/lib.rs, src/detector.rs, and src/extractor/mod.rs to understand the underlying extraction pipeline.

Frequently Asked Questions

Do I need to install Rust to use pdf-inspector in my Node.js project?

No. The @firecrawl/pdf-inspector package ships with pre-compiled N-API binaries for Linux x64/ARM64, macOS ARM64, and Windows x64. The npm install process automatically downloads the correct platform-specific binary, allowing you to integrate pdf-inspector into a Node.js project without a local Rust toolchain.

How does pdf-inspector determine which pages need OCR?

The classifyPdf function analyzes font widths, encoding tables, and content streams in src/detector.rs to classify each page as TextBased, Scanned, Mixed, or ImageBased. It returns a pagesNeedingOcr array containing page numbers where text extraction would be unreliable, along with a confidence score indicating the classification certainty.

Can I extract text from specific regions of a PDF page?

Yes. Use the extractTextInRegions function to define rectangular coordinates on specific pages. This method returns text segments with needsOcr flags and ocrReason explanations, enabling you to build hybrid pipelines that combine structural extraction with external OCR services only for problematic regions.

What is the performance impact of processing large PDFs?

The Rust core parses the PDF once and shares the internal representation between classification and extraction stages. This single-load architecture means subsequent operations on the same buffer are computationally cheap, with initial classification completing in approximately 10–50 milliseconds depending on document complexity.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →