# How to Integrate pdf-inspector into a Node.js Project: Complete Setup Guide

> Integrate pdf-inspector into Node.js easily. Install the npm package for Rust-based PDF classification and Markdown extraction without a local Rust toolchain. Get started now!

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-04

---

**You can integrate pdf-inspector into a Node.js project by installing the `@firecrawl/pdf-inspector` npm package, which provides pre-built N-API binaries that expose Rust-based PDF classification and Markdown extraction functions without requiring a local Rust toolchain.**

The `firecrawl/pdf-inspector` repository delivers a high-performance Rust core for PDF analysis wrapped in Node.js bindings. When you integrate pdf-inspector into a Node.js project, you gain access to sub-50ms PDF classification, full Markdown extraction, and region-specific text extraction while avoiding redundant I/O through single-load parsing.

## Installation and Setup

Start by installing the package from npm. The distribution includes pre-built binaries for Linux x64/ARM64, macOS ARM64, and Windows x64, so no Rust compiler is required on the target machine.

```bash
npm install @firecrawl/pdf-inspector

```

After installation, import the module and read your PDF file into a Buffer:

```javascript
import { readFileSync } from 'fs';
import { classifyPdf, processPdf } from '@firecrawl/pdf-inspector';

const pdf = readFileSync('path/to/document.pdf');

```

The package automatically selects the correct platform-specific binary during installation, as documented in [`napi/README.md`](https://github.com/firecrawl/pdf-inspector/blob/main/napi/README.md)【/cache/repos/github.com/firecrawl/pdf-inspector/main/napi/README.md†L37-L45】.

## Classifying PDF Documents

Use the `classifyPdf` function to determine document type and OCR requirements before processing. This function, implemented via the detector logic in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs), analyzes the PDF in approximately 10–50 milliseconds【/cache/repos/github.com/firecrawl/pdf-inspector/main/README.md†L14-L16】.

```javascript
const classification = classifyPdf(pdf);

console.log(classification.pdfType);        // "TextBased" | "Scanned" | "Mixed" | "ImageBased"
console.log(classification.pageCount);      // total pages
console.log(classification.pagesNeedingOcr); // Array of page numbers

```

The `PdfClassification` object returns a `confidence` score and a `pagesNeedingOcr` array, allowing you to route only specific pages to OCR services rather than processing the entire document【/cache/repos/github.com/firecrawl/pdf-inspector/main/napi/README.md†L41-L56】.

## Extracting Markdown Content

For complete document conversion, use `processPdf`. This function orchestrates the full pipeline defined in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs): detection, extraction, table detection (via [`src/tables/detect_rects.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_rects.rs) and [`src/tables/detect_heuristic.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_heuristic.rs)), and Markdown generation (via [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs))【/cache/repos/github.com/firecrawl/pdf-inspector/main/README.md†L74-L78】【/cache/repos/github.com/firecrawl/pdf-inspector/main/README.md†L165-L173】.

```javascript
const result = processPdf(pdf);

console.log(result.pdfType);   // e.g., "TextBased"
console.log(result.markdown);  // Clean Markdown string or null

```

Because the parser loads the PDF once and shares the internal representation between classification and extraction stages, calling `processPdf` avoids redundant I/O operations【/cache/repos/github.com/firecrawl/pdf-inspector/main/README.md†L65-L73】.

## Region-Based Text Extraction

For hybrid OCR pipelines, use `extractTextInRegions` to extract text from specific rectangular areas. This function accepts an array of page regions and returns text segments with OCR necessity flags.

```javascript
import { extractTextInRegions } from '@firecrawl/pdf-inspector';

const regions = [
  { page: 0, regions: [[0, 0, 300, 400], [300, 0, 612, 400]] },
];

const regionTexts = extractTextInRegions(pdf, regions);

for (const region of regionTexts[0].regions) {
  if (region.needsOcr) {
    // Send to OCR service based on region.ocrReason
  } else {
    console.log(region.text);
  }
}

```

Each `RegionText` object includes a `needsOcr` boolean and an `ocrReason` string (such as `"suspected_garbled_text"`), enabling precise OCR fallback only where the extractor detects unreliable text【/cache/repos/github.com/firecrawl/pdf-inspector/main/napi/README.md†L58-L83】.

## Core Architecture and Source Files

Understanding the underlying Rust architecture helps optimize your integration. The public API exposed to Node.js resides in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs), which coordinates several specialized modules:

- **[`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs)**: Fast PDF type detection using font and content stream analysis.
- **[`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs)**: Orchestrates font handling, content stream parsing, and layout detection.
- **[`src/tables/detect_rects.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_rects.rs)**: Rectangle-based table detection using union-find algorithms.
- **[`src/tables/detect_heuristic.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_heuristic.rs)**: Heuristic table detection using gap-histogram analysis.
- **[`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs)**: Core Markdown generation loop.

The N-API bindings wrap these Rust modules to provide synchronous JavaScript functions that operate on Node.js Buffer objects【/cache/repos/github.com/firecrawl/pdf-inspector/main/README.md†L165-L173】.

## Summary

- Install `@firecrawl/pdf-inspector` to add pre-built Rust PDF processing capabilities to Node.js without compiling from source.
- Use `classifyPdf` for rapid document classification (10–50ms) to determine OCR requirements before full extraction.
- Call `processPdf` to extract complete Markdown representation while avoiding redundant file I/O through single-load parsing.
- Implement `extractTextInRegions` for targeted text extraction when building hybrid OCR pipelines that process only specific document areas.
- Reference source files including [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs), [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs), and [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) to understand the underlying extraction pipeline.

## Frequently Asked Questions

### Do I need to install Rust to use pdf-inspector in my Node.js project?

No. The `@firecrawl/pdf-inspector` package ships with pre-compiled N-API binaries for Linux x64/ARM64, macOS ARM64, and Windows x64. The npm install process automatically downloads the correct platform-specific binary, allowing you to integrate pdf-inspector into a Node.js project without a local Rust toolchain.

### How does pdf-inspector determine which pages need OCR?

The `classifyPdf` function analyzes font widths, encoding tables, and content streams in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) to classify each page as TextBased, Scanned, Mixed, or ImageBased. It returns a `pagesNeedingOcr` array containing page numbers where text extraction would be unreliable, along with a confidence score indicating the classification certainty.

### Can I extract text from specific regions of a PDF page?

Yes. Use the `extractTextInRegions` function to define rectangular coordinates on specific pages. This method returns text segments with `needsOcr` flags and `ocrReason` explanations, enabling you to build hybrid pipelines that combine structural extraction with external OCR services only for problematic regions.

### What is the performance impact of processing large PDFs?

The Rust core parses the PDF once and shares the internal representation between classification and extraction stages. This single-load architecture means subsequent operations on the same buffer are computationally cheap, with initial classification completing in approximately 10–50 milliseconds depending on document complexity.