# How to Contribute to the firecrawl pdf-inspector Project: A Complete Developer Guide

> Learn how to contribute to the firecrawl pdf-inspector project. Follow our developer guide to fork, branch, test, and submit your first pull request to enhance the PDF inspection tool.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-07

---

**Fork the repository, create a feature branch, ensure `cargo fmt`, `cargo clippy -- -D warnings`, and `cargo test` all pass, then submit a PR targeting `main` with updated documentation.**

The **firecrawl pdf-inspector** project is a Rust-based library and CLI tool for extracting structured markdown from PDF documents. Whether you want to fix bugs, improve table detection, or add new language bindings, understanding the project's architecture and contribution workflow is essential. This guide walks through everything you need to know to contribute effectively.

## Understanding the pdf-inspector Architecture

Before making changes, familiarize yourself with how the codebase is organized:

| Layer | Purpose | Key Implementation |
|-------|---------|-------------------|
| **Public API** | Exposes main entry points like `process_pdf_with_options` and encoding-issue detection | [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) |
| **PDF Type Detection** | Classifies PDFs as TextBased, Scanned, or Mixed; performs tiled-scan detection | [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) |
| **Extraction Engine** | Orchestrates content-stream parsing, font handling, layout analysis, and reading order | [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) |
| **Table Detection** | Three-stage detection: rect-based → line-based → heuristic with union-find clustering | [`src/tables/detect_rects.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_rects.rs), [`src/tables/detect_lines.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_lines.rs), [`src/tables/detect_heuristic.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_heuristic.rs) |
| **Markdown Generation** | Converts extracted lines to clean markdown with post-processing | [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs) |
| **CLI Binaries** | `pdf2md` for PDF→Markdown conversion and `detect-pdf` for classification | [`src/bin/pdf2md.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs), [`src/bin/detect_pdf.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/detect_pdf.rs) |
| **N-API / WASM Bindings** | JavaScript and Python bindings for the core library | [`napi/src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/napi/src/lib.rs), [`wasm/src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/wasm/src/lib.rs) |

## Contribution Workflow: Step by Step

### 1. Fork and Branch

Fork the repository on GitHub, then create a descriptive branch:

```bash
git clone https://github.com/YOUR_USERNAME/pdf-inspector.git
cd pdf-inspector
git checkout -b fix-table-detection-edge-case

```

### 2. Run Quality Checks Locally

All three commands must pass before committing:

```bash
cargo fmt                          # Enforce consistent formatting

cargo clippy -- -D warnings        # Linting: zero warnings allowed

cargo test                         # Run unit and integration tests

```

The CI pipeline enforces these strictly. Any warning will fail the build.

### 3. Follow Project Conventions

According to the source code conventions in [`AGENTS.md`](https://github.com/firecrawl/pdf-inspector/blob/main/AGENTS.md):

- Use `is_some_and` instead of `map_or(false, …)` for cleaner Option handling
- Never introduce new compiler or clippy warnings
- Respect table rendering limits: **25 columns maximum**, skip `propagate_merged_cells` for tables wider than **10 columns**

### 4. Update Documentation

If you modify public APIs, update the corresponding markdown files in `docs/`:

| Documentation File | Purpose |
|-------------------|---------|
| [`docs/rust-api.md`](https://github.com/firecrawl/pdf-inspector/blob/main/docs/rust-api.md) | Rust library API reference |
| [`docs/python-api.md`](https://github.com/firecrawl/pdf-inspector/blob/main/docs/python-api.md) | Python bindings documentation |
| [`docs/debugging.md`](https://github.com/firecrawl/pdf-inspector/blob/main/docs/debugging.md) | Debugging techniques and logging |
| [`docs/benchmarking.md`](https://github.com/firecrawl/pdf-inspector/blob/main/docs/benchmarking.md) | Performance measurement guide |

### 5. Run Benchmarks (Recommended)

Verify performance hasn't regressed:

```bash
cd pdf-evals
bench.py test -q   # Quick benchmark suite

```

### 6. Submit Your Pull Request

Target the `main` branch. The CI automatically runs formatting, linting, testing, and benchmark checks. Address any review feedback promptly.

## Code Examples for Common Contributions

### Using the Rust API

From [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs), the primary entry point is `process_pdf_with_options`:

```rust
use pdf_inspector::process_pdf_with_options;
use pdf_inspector::ProcessOptions;

fn main() -> Result<(), Box<dyn std::error::Error>> {
    let opts = ProcessOptions::default()
        .with_json(true)          // Request structured JSON output
        .with_dry_run(false);     // Actually extract the PDF

    let markdown = process_pdf_with_options("sample.pdf", opts)?;
    println!("{}", markdown);
    Ok(())
}

```

### Adding a New Table Detection Heuristic

The table detection system in [`src/tables/detect_heuristic.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_heuristic.rs) accepts new strategies:

```rust
// src/tables/detect_heuristic.rs
pub fn detect_new_heuristic(page: &PdfPage) -> Option<Table> {
    // 1️⃣ Compute horizontal gaps between text items
    // 2️⃣ Group items that share a common baseline
    // 3️⃣ Validate that the group forms a rectangular grid
    //    (use the existing `Grid` helper from `tables/grid.rs`)
    // 4️⃣ Return Some(Table) if the heuristic succeeds
}

```

### Running the CLI Tools

```bash

# Extract PDF to markdown with debug logging

RUST_LOG=pdf_inspector::extractor=debug cargo run --bin pdf2md -- path/to/document.pdf

# Classify PDF type with JSON output

cargo run --bin detect_pdf -- path/to/document.pdf --json

```

## Key Source Files for Contributors

| File | Role |
|------|------|
| [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) | Public library entry point and `process_pdf_with_options` |
| [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) | PDF classification and tiled-scan detection |
| [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) | Core extraction orchestration |
| [`src/tables/detect_rects.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_rects.rs) | Rectangle-based table detection (first strategy) |
| [`src/tables/detect_lines.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_lines.rs) | Line-based table detection fallback |
| [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs) | Markdown generation and post-processing |
| [`src/bin/pdf2md.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs) | PDF→Markdown CLI binary |
| [`napi/src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/napi/src/lib.rs) | Node.js N-API bindings |
| [`wasm/src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/wasm/src/lib.rs) | WebAssembly bindings |

## Summary

- **Fork and branch** from `main` with descriptive names
- **Quality gates**: `cargo fmt`, `cargo clippy -- -D warnings`, and `cargo test` must all pass
- **Follow conventions**: `is_some_and` over `map_or`, no warnings, respect table column limits
- **Document changes** in `docs/` when modifying public APIs
- **Benchmark** performance-critical changes using [`pdf-evals/bench.py`](https://github.com/firecrawl/pdf-inspector/blob/main/pdf-evals/bench.py)
- **Target `main`** with pull requests and respond to review feedback

## Frequently Asked Questions

### Does pdf-inspector require specific Rust toolchain versions?

The project uses standard Cargo tooling with no special version requirements beyond stable Rust. The CI enforces formatting and linting with the latest stable toolchain, so ensure your local environment is up to date.

### What happens if my PR introduces a clippy warning?

The CI runs `cargo clippy -- -D warnings`, which treats all warnings as errors. Your PR will fail automated checks. Run clippy locally before pushing to catch issues early.

### Can I contribute bindings for other languages?

Yes. The project already includes N-API (Node.js) and WASM bindings in [`napi/src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/napi/src/lib.rs) and [`wasm/src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/wasm/src/lib.rs) respectively. Follow these patterns and add corresponding documentation in `docs/` for any new language bindings.

### How do I test table detection improvements?

The table detection system spans [`src/tables/detect_rects.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_rects.rs), [`src/tables/detect_lines.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_lines.rs), and [`src/tables/detect_heuristic.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_heuristic.rs). Add test cases to the integration test suite, run `cargo test`, and verify with `pdf-evals/bench.py test -q` for performance validation.