Is Firecrawl pdf-inspector Open Source? License, Architecture, and Usage Guide

Yes. Firecrawl pdf-inspector is open source under the MIT License, with all source code publicly available on GitHub for reading, modification, and redistribution.

This article explains the licensing terms, examines the Rust-based architecture, and demonstrates how to use the library from Python, Node.js, WebAssembly, and the command line. Whether you're evaluating pdf-inspector for commercial use or planning to extend its capabilities, the MIT License grants broad permissions with minimal restrictions.

MIT License Terms and Permissions

The firecrawl/pdf-inspector repository includes a LICENSE file at the repository root that explicitly states the MIT terms. According to the source code, this license allows you to:

  • Use the software for any purpose, including commercial applications
  • Modify the source code to suit your needs
  • Distribute original or modified versions
  • Sublicense the software under compatible terms

The only requirement is preserving the copyright notice and license text in any distributed copies. No copyleft provisions apply, making pdf-inspector suitable for both open-source and proprietary projects.

Core Architecture: Detection Before Extraction

Pdf-inspector implements a single-load pipeline that reads a PDF once, then shares the parsed structure between detection and extraction stages. This eliminates redundant I/O and enables fast classification without full text extraction.

PDF Type Detection (src/detector.rs)

The detector quickly categorizes PDFs into four types and identifies which pages might need OCR:

  • TextBased – Native text with extractable content streams
  • Scanned – Image-only pages requiring OCR
  • ImageBased – PDFs containing mostly embedded images
  • Mixed – Hybrid documents with varied page types

The detection logic in src/detector.rs analyzes PDF structure without rendering, making it significantly faster than full parsing.

Extraction Pipeline (src/extractor/)

Once detection completes, the extractor processes fonts and content streams:

Module File Responsibility
Font handling src/extractor/fonts.rs CFF, TrueType, Type1 font parsing
Content streams src/extractor/content_stream.rs Page content parsing and text positioning
Layout analysis src/extractor/layout.rs Column detection and reading-order computation
XObjects src/extractor/xobjects.rs Form objects and external references
Links src/extractor/links.rs Hyperlink extraction and annotation parsing

Table Detection (src/tables/)

The table detection system uses two complementary approaches:

Markdown Conversion (src/markdown/)

The Markdown pipeline analyzes typography and document structure:

Public API Entry Points

The library exposes two primary functions in src/lib.rs:

  • process_pdf – Runs the complete pipeline: detection, extraction, table detection, and Markdown conversion
  • classify_pdf – Runs detection only, returning PDF type and OCR recommendations without full extraction

Both functions accept options for customizing font handling, table detection thresholds, and Markdown output formatting.

Usage Examples by Language

Python (PyO3 Bindings)

import pdf_inspector

# Process a PDF file through the full pipeline

result = pdf_inspector.process_pdf("sample.pdf")

print("PDF type :", result.pdf_type)        # e.g., "text_based"

print("Markdown  :", result.markdown[:200]) # first 200 characters of output

For complete API documentation, see [docs/python.md](https://github.com/firecrawl/pdf-inspector/blob/main/docs/python.md).

Rust

use pdf_inspector::process_pdf;

fn main() -> Result<(), pdf_inspector::Error> {
    let result = process_pdf("sample.pdf")?;
    println!("PDF type: {:?}", result.pdf_type);
    if let Some(md) = result.markdown {
        println!("{}", md);
    }
    Ok(())
}

Reference: [docs/rust-api.md](https://github.com/firecrawl/pdf-inspector/blob/main/docs/rust-api.md)

Node.js (N-API)

import { readFileSync } from "fs";
import { processPdf } from "@firecrawl/pdf-inspector";

const pdfData = readFileSync("sample.pdf");
const result = processPdf(pdfData);

console.log("PDF type :", result.pdfType);   // "TextBased", "Scanned", etc.
console.log("Markdown :", result.markdown?.slice(0, 200));

Reference: [napi/README.md](https://github.com/firecrawl/pdf-inspector/blob/main/napi/README.md)

Command Line


# Convert a PDF to Markdown

pdf2md sample.pdf

# Get JSON with PDF type and positioned text items

pdf2md sample.pdf --json

# Detect PDF type only (faster, no extraction)

detect_pdf sample.pdf

The CLI tools are implemented in src/bin/pdf2md.rs and src/bin/detect_pdf.rs.

Key Implementation Files

File Purpose
src/lib.rs Public API (process_pdf, classify_pdf) and options builder
src/detector.rs Fast PDF-type detection logic
src/extractor/layout.rs Column detection and reading-order computation
src/tables/detect_rects.rs Rectangle-based table detection (union-find)
src/markdown/convert.rs Core Markdown conversion loop
src/markdown/analysis.rs Font statistics and heading tier detection
src/text_utils.rs CJK, RTL, ligature expansion, bold/italic detection
src/bin/pdf2md.rs CLI binary for Markdown conversion
src/bin/detect_pdf.rs CLI for PDF type detection only

Summary

  • Firecrawl pdf-inspector is open source under the MIT License, permitting commercial and private use with minimal attribution requirements
  • The single-load architecture parses PDFs once, then routes through detection or full extraction as needed
  • Multi-language bindings support Python, Node.js, WebAssembly, and Rust-native usage
  • Modular design separates concerns: detection (src/detector.rs), extraction (src/extractor/), table detection (src/tables/), and Markdown conversion (src/markdown/)
  • Two primary APIs (process_pdf and classify_pdf in src/lib.rs) serve different performance and accuracy tradeoffs

Frequently Asked Questions

Can I use Firecrawl pdf-inspector in a commercial product?

Yes. The MIT License explicitly permits commercial use, modification, and distribution. You must include the original copyright notice and license text, but no other restrictions apply. This makes pdf-inspector suitable for both proprietary software and SaaS products.

Does the MIT License require me to open-source my modifications?

No. The MIT License is permissive, not copyleft. You can modify pdf-inspector and keep your changes proprietary. The only requirement is preserving the original copyright notice and license text in any distributed copies of the library itself.

What makes pdf-inspector faster than OCR-based PDF tools?

Pdf-inspector uses detection before extraction: the src/detector.rs module quickly classifies PDF type by analyzing structure without rendering. For text-based PDFs, src/extractor/content_stream.rs parses native text positions directly from PDF content streams, avoiding expensive image rasterization and OCR entirely. OCR is only recommended for scanned pages identified during detection.

How do I extend pdf-inspector with custom Markdown formatting?

Fork the repository and modify the src/markdown/ modules. The pipeline structure in src/markdown/convert.rs processes lines through analysis, preprocessing, conversion, classification, and postprocessing stages. You can customize heading detection logic in src/markdown/analysis.rs or add new block types in src/markdown/classify.rs. Since the project uses MIT licensing, your modifications can remain private or be contributed back via pull request.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →