How to Contribute to the firecrawl pdf-inspector Project: A Complete Developer Guide

Fork the repository, create a feature branch, ensure cargo fmt, cargo clippy -- -D warnings, and cargo test all pass, then submit a PR targeting main with updated documentation.

The firecrawl pdf-inspector project is a Rust-based library and CLI tool for extracting structured markdown from PDF documents. Whether you want to fix bugs, improve table detection, or add new language bindings, understanding the project's architecture and contribution workflow is essential. This guide walks through everything you need to know to contribute effectively.

Understanding the pdf-inspector Architecture

Before making changes, familiarize yourself with how the codebase is organized:

Layer Purpose Key Implementation
Public API Exposes main entry points like process_pdf_with_options and encoding-issue detection src/lib.rs
PDF Type Detection Classifies PDFs as TextBased, Scanned, or Mixed; performs tiled-scan detection src/detector.rs
Extraction Engine Orchestrates content-stream parsing, font handling, layout analysis, and reading order src/extractor/mod.rs
Table Detection Three-stage detection: rect-based → line-based → heuristic with union-find clustering src/tables/detect_rects.rs, src/tables/detect_lines.rs, src/tables/detect_heuristic.rs
Markdown Generation Converts extracted lines to clean markdown with post-processing src/markdown/convert.rs
CLI Binaries pdf2md for PDF→Markdown conversion and detect-pdf for classification src/bin/pdf2md.rs, src/bin/detect_pdf.rs
N-API / WASM Bindings JavaScript and Python bindings for the core library napi/src/lib.rs, wasm/src/lib.rs

Contribution Workflow: Step by Step

1. Fork and Branch

Fork the repository on GitHub, then create a descriptive branch:

git clone https://github.com/YOUR_USERNAME/pdf-inspector.git
cd pdf-inspector
git checkout -b fix-table-detection-edge-case

2. Run Quality Checks Locally

All three commands must pass before committing:

cargo fmt                          # Enforce consistent formatting

cargo clippy -- -D warnings        # Linting: zero warnings allowed

cargo test                         # Run unit and integration tests

The CI pipeline enforces these strictly. Any warning will fail the build.

3. Follow Project Conventions

According to the source code conventions in AGENTS.md:

  • Use is_some_and instead of map_or(false, …) for cleaner Option handling
  • Never introduce new compiler or clippy warnings
  • Respect table rendering limits: 25 columns maximum, skip propagate_merged_cells for tables wider than 10 columns

4. Update Documentation

If you modify public APIs, update the corresponding markdown files in docs/:

Documentation File Purpose
docs/rust-api.md Rust library API reference
docs/python-api.md Python bindings documentation
docs/debugging.md Debugging techniques and logging
docs/benchmarking.md Performance measurement guide

Verify performance hasn't regressed:

cd pdf-evals
bench.py test -q   # Quick benchmark suite

6. Submit Your Pull Request

Target the main branch. The CI automatically runs formatting, linting, testing, and benchmark checks. Address any review feedback promptly.

Code Examples for Common Contributions

Using the Rust API

From src/lib.rs, the primary entry point is process_pdf_with_options:

use pdf_inspector::process_pdf_with_options;
use pdf_inspector::ProcessOptions;

fn main() -> Result<(), Box<dyn std::error::Error>> {
    let opts = ProcessOptions::default()
        .with_json(true)          // Request structured JSON output
        .with_dry_run(false);     // Actually extract the PDF

    let markdown = process_pdf_with_options("sample.pdf", opts)?;
    println!("{}", markdown);
    Ok(())
}

Adding a New Table Detection Heuristic

The table detection system in src/tables/detect_heuristic.rs accepts new strategies:

// src/tables/detect_heuristic.rs
pub fn detect_new_heuristic(page: &PdfPage) -> Option<Table> {
    // 1️⃣ Compute horizontal gaps between text items
    // 2️⃣ Group items that share a common baseline
    // 3️⃣ Validate that the group forms a rectangular grid
    //    (use the existing `Grid` helper from `tables/grid.rs`)
    // 4️⃣ Return Some(Table) if the heuristic succeeds
}

Running the CLI Tools


# Extract PDF to markdown with debug logging

RUST_LOG=pdf_inspector::extractor=debug cargo run --bin pdf2md -- path/to/document.pdf

# Classify PDF type with JSON output

cargo run --bin detect_pdf -- path/to/document.pdf --json

Key Source Files for Contributors

File Role
src/lib.rs Public library entry point and process_pdf_with_options
src/detector.rs PDF classification and tiled-scan detection
src/extractor/mod.rs Core extraction orchestration
src/tables/detect_rects.rs Rectangle-based table detection (first strategy)
src/tables/detect_lines.rs Line-based table detection fallback
src/markdown/convert.rs Markdown generation and post-processing
src/bin/pdf2md.rs PDF→Markdown CLI binary
napi/src/lib.rs Node.js N-API bindings
wasm/src/lib.rs WebAssembly bindings

Summary

  • Fork and branch from main with descriptive names
  • Quality gates: cargo fmt, cargo clippy -- -D warnings, and cargo test must all pass
  • Follow conventions: is_some_and over map_or, no warnings, respect table column limits
  • Document changes in docs/ when modifying public APIs
  • Benchmark performance-critical changes using pdf-evals/bench.py
  • Target main with pull requests and respond to review feedback

Frequently Asked Questions

Does pdf-inspector require specific Rust toolchain versions?

The project uses standard Cargo tooling with no special version requirements beyond stable Rust. The CI enforces formatting and linting with the latest stable toolchain, so ensure your local environment is up to date.

What happens if my PR introduces a clippy warning?

The CI runs cargo clippy -- -D warnings, which treats all warnings as errors. Your PR will fail automated checks. Run clippy locally before pushing to catch issues early.

Can I contribute bindings for other languages?

Yes. The project already includes N-API (Node.js) and WASM bindings in napi/src/lib.rs and wasm/src/lib.rs respectively. Follow these patterns and add corresponding documentation in docs/ for any new language bindings.

How do I test table detection improvements?

The table detection system spans src/tables/detect_rects.rs, src/tables/detect_lines.rs, and src/tables/detect_heuristic.rs. Add test cases to the integration test suite, run cargo test, and verify with pdf-evals/bench.py test -q for performance validation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →