Is Firecrawl pdf-inspector Open Source? License, Architecture, and Usage Guide
Yes. Firecrawl pdf-inspector is open source under the MIT License, with all source code publicly available on GitHub for reading, modification, and redistribution.
This article explains the licensing terms, examines the Rust-based architecture, and demonstrates how to use the library from Python, Node.js, WebAssembly, and the command line. Whether you're evaluating pdf-inspector for commercial use or planning to extend its capabilities, the MIT License grants broad permissions with minimal restrictions.
MIT License Terms and Permissions
The firecrawl/pdf-inspector repository includes a LICENSE file at the repository root that explicitly states the MIT terms. According to the source code, this license allows you to:
- Use the software for any purpose, including commercial applications
- Modify the source code to suit your needs
- Distribute original or modified versions
- Sublicense the software under compatible terms
The only requirement is preserving the copyright notice and license text in any distributed copies. No copyleft provisions apply, making pdf-inspector suitable for both open-source and proprietary projects.
Core Architecture: Detection Before Extraction
Pdf-inspector implements a single-load pipeline that reads a PDF once, then shares the parsed structure between detection and extraction stages. This eliminates redundant I/O and enables fast classification without full text extraction.
PDF Type Detection (src/detector.rs)
The detector quickly categorizes PDFs into four types and identifies which pages might need OCR:
- TextBased – Native text with extractable content streams
- Scanned – Image-only pages requiring OCR
- ImageBased – PDFs containing mostly embedded images
- Mixed – Hybrid documents with varied page types
The detection logic in src/detector.rs analyzes PDF structure without rendering, making it significantly faster than full parsing.
Extraction Pipeline (src/extractor/)
Once detection completes, the extractor processes fonts and content streams:
| Module | File | Responsibility |
|---|---|---|
| Font handling | src/extractor/fonts.rs |
CFF, TrueType, Type1 font parsing |
| Content streams | src/extractor/content_stream.rs |
Page content parsing and text positioning |
| Layout analysis | src/extractor/layout.rs |
Column detection and reading-order computation |
| XObjects | src/extractor/xobjects.rs |
Form objects and external references |
| Links | src/extractor/links.rs |
Hyperlink extraction and annotation parsing |
Table Detection (src/tables/)
The table detection system uses two complementary approaches:
src/tables/detect_rects.rs– Rectangle-based detection using union-find algorithms to group cell boundariessrc/tables/detect_heuristic.rs– Alignment-based heuristics for table structure inferencesrc/tables/grid.rs– Grid normalization and cell boundary resolutionsrc/tables/format.rs– Output formatting for various target formats
Markdown Conversion (src/markdown/)
The Markdown pipeline analyzes typography and document structure:
src/markdown/analysis.rs– Font statistics and heading tier detection based on size and weightsrc/markdown/preprocess.rs– Heading and list preprocessingsrc/markdown/convert.rs– Core conversion loop from positioned text to Markdownsrc/markdown/classify.rs– Block-level classification (paragraph, heading, list, code block)src/markdown/postprocess.rs– Cleanup and normalization of output
Public API Entry Points
The library exposes two primary functions in src/lib.rs:
process_pdf– Runs the complete pipeline: detection, extraction, table detection, and Markdown conversionclassify_pdf– Runs detection only, returning PDF type and OCR recommendations without full extraction
Both functions accept options for customizing font handling, table detection thresholds, and Markdown output formatting.
Usage Examples by Language
Python (PyO3 Bindings)
import pdf_inspector
# Process a PDF file through the full pipeline
result = pdf_inspector.process_pdf("sample.pdf")
print("PDF type :", result.pdf_type) # e.g., "text_based"
print("Markdown :", result.markdown[:200]) # first 200 characters of output
For complete API documentation, see [docs/python.md](https://github.com/firecrawl/pdf-inspector/blob/main/docs/python.md).
Rust
use pdf_inspector::process_pdf;
fn main() -> Result<(), pdf_inspector::Error> {
let result = process_pdf("sample.pdf")?;
println!("PDF type: {:?}", result.pdf_type);
if let Some(md) = result.markdown {
println!("{}", md);
}
Ok(())
}
Reference: [docs/rust-api.md](https://github.com/firecrawl/pdf-inspector/blob/main/docs/rust-api.md)
Node.js (N-API)
import { readFileSync } from "fs";
import { processPdf } from "@firecrawl/pdf-inspector";
const pdfData = readFileSync("sample.pdf");
const result = processPdf(pdfData);
console.log("PDF type :", result.pdfType); // "TextBased", "Scanned", etc.
console.log("Markdown :", result.markdown?.slice(0, 200));
Reference: [napi/README.md](https://github.com/firecrawl/pdf-inspector/blob/main/napi/README.md)
Command Line
# Convert a PDF to Markdown
pdf2md sample.pdf
# Get JSON with PDF type and positioned text items
pdf2md sample.pdf --json
# Detect PDF type only (faster, no extraction)
detect_pdf sample.pdf
The CLI tools are implemented in src/bin/pdf2md.rs and src/bin/detect_pdf.rs.
Key Implementation Files
| File | Purpose |
|---|---|
src/lib.rs |
Public API (process_pdf, classify_pdf) and options builder |
src/detector.rs |
Fast PDF-type detection logic |
src/extractor/layout.rs |
Column detection and reading-order computation |
src/tables/detect_rects.rs |
Rectangle-based table detection (union-find) |
src/markdown/convert.rs |
Core Markdown conversion loop |
src/markdown/analysis.rs |
Font statistics and heading tier detection |
src/text_utils.rs |
CJK, RTL, ligature expansion, bold/italic detection |
src/bin/pdf2md.rs |
CLI binary for Markdown conversion |
src/bin/detect_pdf.rs |
CLI for PDF type detection only |
Summary
- Firecrawl pdf-inspector is open source under the MIT License, permitting commercial and private use with minimal attribution requirements
- The single-load architecture parses PDFs once, then routes through detection or full extraction as needed
- Multi-language bindings support Python, Node.js, WebAssembly, and Rust-native usage
- Modular design separates concerns: detection (
src/detector.rs), extraction (src/extractor/), table detection (src/tables/), and Markdown conversion (src/markdown/) - Two primary APIs (
process_pdfandclassify_pdfinsrc/lib.rs) serve different performance and accuracy tradeoffs
Frequently Asked Questions
Can I use Firecrawl pdf-inspector in a commercial product?
Yes. The MIT License explicitly permits commercial use, modification, and distribution. You must include the original copyright notice and license text, but no other restrictions apply. This makes pdf-inspector suitable for both proprietary software and SaaS products.
Does the MIT License require me to open-source my modifications?
No. The MIT License is permissive, not copyleft. You can modify pdf-inspector and keep your changes proprietary. The only requirement is preserving the original copyright notice and license text in any distributed copies of the library itself.
What makes pdf-inspector faster than OCR-based PDF tools?
Pdf-inspector uses detection before extraction: the src/detector.rs module quickly classifies PDF type by analyzing structure without rendering. For text-based PDFs, src/extractor/content_stream.rs parses native text positions directly from PDF content streams, avoiding expensive image rasterization and OCR entirely. OCR is only recommended for scanned pages identified during detection.
How do I extend pdf-inspector with custom Markdown formatting?
Fork the repository and modify the src/markdown/ modules. The pipeline structure in src/markdown/convert.rs processes lines through analysis, preprocessing, conversion, classification, and postprocessing stages. You can customize heading detection logic in src/markdown/analysis.rs or add new block types in src/markdown/classify.rs. Since the project uses MIT licensing, your modifications can remain private or be contributed back via pull request.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →