# Is Firecrawl pdf-inspector Open Source? License, Architecture, and Usage Guide

> Yes Firecrawl pdf-inspector is open source under the MIT License. Explore its GitHub repository for source code, architecture details, and usage.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: getting-started
- Published: 2026-08-07

---

**Yes. Firecrawl pdf-inspector is open source under the MIT License**, with all source code publicly available on GitHub for reading, modification, and redistribution.

This article explains the licensing terms, examines the Rust-based architecture, and demonstrates how to use the library from Python, Node.js, WebAssembly, and the command line. Whether you're evaluating pdf-inspector for commercial use or planning to extend its capabilities, the MIT License grants broad permissions with minimal restrictions.

## MIT License Terms and Permissions

The `firecrawl/pdf-inspector` repository includes a `LICENSE` file at the repository root that explicitly states the MIT terms. According to the source code, this license allows you to:

- **Use** the software for any purpose, including commercial applications
- **Modify** the source code to suit your needs
- **Distribute** original or modified versions
- **Sublicense** the software under compatible terms

The only requirement is preserving the copyright notice and license text in any distributed copies. No copyleft provisions apply, making pdf-inspector suitable for both open-source and proprietary projects.

## Core Architecture: Detection Before Extraction

Pdf-inspector implements a **single-load pipeline** that reads a PDF once, then shares the parsed structure between detection and extraction stages. This eliminates redundant I/O and enables fast classification without full text extraction.

### PDF Type Detection ([`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs))

The detector quickly categorizes PDFs into four types and identifies which pages might need OCR:

- **TextBased** – Native text with extractable content streams
- **Scanned** – Image-only pages requiring OCR
- **ImageBased** – PDFs containing mostly embedded images
- **Mixed** – Hybrid documents with varied page types

The detection logic in [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) analyzes PDF structure without rendering, making it significantly faster than full parsing.

### Extraction Pipeline (`src/extractor/`)

Once detection completes, the extractor processes fonts and content streams:

| Module | File | Responsibility |
|--------|------|--------------|
| Font handling | [`src/extractor/fonts.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/fonts.rs) | CFF, TrueType, Type1 font parsing |
| Content streams | [`src/extractor/content_stream.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs) | Page content parsing and text positioning |
| Layout analysis | [`src/extractor/layout.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs) | Column detection and reading-order computation |
| XObjects | [`src/extractor/xobjects.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/xobjects.rs) | Form objects and external references |
| Links | [`src/extractor/links.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/links.rs) | Hyperlink extraction and annotation parsing |

### Table Detection (`src/tables/`)

The table detection system uses two complementary approaches:

- **[`src/tables/detect_rects.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_rects.rs)** – Rectangle-based detection using union-find algorithms to group cell boundaries
- **[`src/tables/detect_heuristic.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_heuristic.rs)** – Alignment-based heuristics for table structure inference
- **[`src/tables/grid.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/grid.rs)** – Grid normalization and cell boundary resolution
- **[`src/tables/format.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/format.rs)** – Output formatting for various target formats

### Markdown Conversion (`src/markdown/`)

The Markdown pipeline analyzes typography and document structure:

- **[`src/markdown/analysis.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/analysis.rs)** – Font statistics and heading tier detection based on size and weight
- **[`src/markdown/preprocess.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/preprocess.rs)** – Heading and list preprocessing
- **[`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs)** – Core conversion loop from positioned text to Markdown
- **[`src/markdown/classify.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/classify.rs)** – Block-level classification (paragraph, heading, list, code block)
- **[`src/markdown/postprocess.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/postprocess.rs)** – Cleanup and normalization of output

## Public API Entry Points

The library exposes two primary functions in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs):

- **`process_pdf`** – Runs the complete pipeline: detection, extraction, table detection, and Markdown conversion
- **`classify_pdf`** – Runs detection only, returning PDF type and OCR recommendations without full extraction

Both functions accept options for customizing font handling, table detection thresholds, and Markdown output formatting.

## Usage Examples by Language

### Python (PyO3 Bindings)

```python
import pdf_inspector

# Process a PDF file through the full pipeline

result = pdf_inspector.process_pdf("sample.pdf")

print("PDF type :", result.pdf_type)        # e.g., "text_based"

print("Markdown  :", result.markdown[:200]) # first 200 characters of output

```

For complete API documentation, see [[`docs/python.md`](https://github.com/firecrawl/pdf-inspector/blob/main/docs/python.md)](https://github.com/firecrawl/pdf-inspector/blob/main/docs/python.md).

### Rust

```rust
use pdf_inspector::process_pdf;

fn main() -> Result<(), pdf_inspector::Error> {
    let result = process_pdf("sample.pdf")?;
    println!("PDF type: {:?}", result.pdf_type);
    if let Some(md) = result.markdown {
        println!("{}", md);
    }
    Ok(())
}

```

Reference: [[`docs/rust-api.md`](https://github.com/firecrawl/pdf-inspector/blob/main/docs/rust-api.md)](https://github.com/firecrawl/pdf-inspector/blob/main/docs/rust-api.md)

### Node.js (N-API)

```javascript
import { readFileSync } from "fs";
import { processPdf } from "@firecrawl/pdf-inspector";

const pdfData = readFileSync("sample.pdf");
const result = processPdf(pdfData);

console.log("PDF type :", result.pdfType);   // "TextBased", "Scanned", etc.
console.log("Markdown :", result.markdown?.slice(0, 200));

```

Reference: [[`napi/README.md`](https://github.com/firecrawl/pdf-inspector/blob/main/napi/README.md)](https://github.com/firecrawl/pdf-inspector/blob/main/napi/README.md)

### Command Line

```bash

# Convert a PDF to Markdown

pdf2md sample.pdf

# Get JSON with PDF type and positioned text items

pdf2md sample.pdf --json

# Detect PDF type only (faster, no extraction)

detect_pdf sample.pdf

```

The CLI tools are implemented in [`src/bin/pdf2md.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs) and [`src/bin/detect_pdf.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/detect_pdf.rs).

## Key Implementation Files

| File | Purpose |
|------|---------|
| [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) | Public API (`process_pdf`, `classify_pdf`) and options builder |
| [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) | Fast PDF-type detection logic |
| [`src/extractor/layout.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs) | Column detection and reading-order computation |
| [`src/tables/detect_rects.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_rects.rs) | Rectangle-based table detection (union-find) |
| [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs) | Core Markdown conversion loop |
| [`src/markdown/analysis.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/analysis.rs) | Font statistics and heading tier detection |
| [`src/text_utils.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/text_utils.rs) | CJK, RTL, ligature expansion, bold/italic detection |
| [`src/bin/pdf2md.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs) | CLI binary for Markdown conversion |
| [`src/bin/detect_pdf.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/detect_pdf.rs) | CLI for PDF type detection only |

## Summary

- **Firecrawl pdf-inspector is open source under the MIT License**, permitting commercial and private use with minimal attribution requirements
- The **single-load architecture** parses PDFs once, then routes through detection or full extraction as needed
- **Multi-language bindings** support Python, Node.js, WebAssembly, and Rust-native usage
- **Modular design** separates concerns: detection ([`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs)), extraction (`src/extractor/`), table detection (`src/tables/`), and Markdown conversion (`src/markdown/`)
- **Two primary APIs** (`process_pdf` and `classify_pdf` in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)) serve different performance and accuracy tradeoffs

## Frequently Asked Questions

### Can I use Firecrawl pdf-inspector in a commercial product?

Yes. The MIT License explicitly permits commercial use, modification, and distribution. You must include the original copyright notice and license text, but no other restrictions apply. This makes pdf-inspector suitable for both proprietary software and SaaS products.

### Does the MIT License require me to open-source my modifications?

No. The MIT License is **permissive**, not copyleft. You can modify pdf-inspector and keep your changes proprietary. The only requirement is preserving the original copyright notice and license text in any distributed copies of the library itself.

### What makes pdf-inspector faster than OCR-based PDF tools?

Pdf-inspector uses **detection before extraction**: the [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) module quickly classifies PDF type by analyzing structure without rendering. For text-based PDFs, [`src/extractor/content_stream.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs) parses native text positions directly from PDF content streams, avoiding expensive image rasterization and OCR entirely. OCR is only recommended for scanned pages identified during detection.

### How do I extend pdf-inspector with custom Markdown formatting?

Fork the repository and modify the `src/markdown/` modules. The pipeline structure in [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs) processes lines through analysis, preprocessing, conversion, classification, and postprocessing stages. You can customize heading detection logic in [`src/markdown/analysis.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/analysis.rs) or add new block types in [`src/markdown/classify.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/classify.rs). Since the project uses MIT licensing, your modifications can remain private or be contributed back via pull request.