# How to Contribute to the pdf-inspector Project: A Complete Guide for Rust Developers

> Learn how to contribute to the pdf-inspector Rust project. This guide covers selecting a task, running tests, and submitting PRs via GitHub for successful contributions.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-04

---

**Contributing to pdf-inspector involves selecting a subsystem to improve, running the test suite, following the PR workflow, and submitting changes through GitHub.**

The **pdf-inspector** project is a fast Rust library maintained by Firecrawl that performs PDF type detection, text extraction, and Markdown conversion. Its modular architecture deliberately separates concerns into independent stages, making it straightforward for contributors to work on specific components without understanding the entire pipeline. This guide walks you through how to contribute to pdf-inspector effectively, based on the actual source code structure.

## Understanding the pdf-inspector Architecture

Before contributing, you need to understand how the codebase is organized. The project is split into distinct stages, each with clear responsibilities and dedicated source files.

### Core Subsystems and File Locations

- **Detection** — Quickly classifies PDFs as `TextBased`, `Scanned`, `ImageBased`, or `Mixed` using only the X-ref table and lightweight content stream scanning. Core logic resides in [[`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs).

- **Extraction** — Walks PDF operators (`Tj`, `TJ`, `Td`, `Tm`, etc.) to produce positioned `TextItem`s and `PdfRect`s, handling fonts, ToUnicode CMaps, and XObjects. Implemented in [[`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) and [[`src/extractor/content_stream.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs).

- **Layout** — Detects multi-column reading order and prepares data for table detection. Found in [[`src/extractor/layout.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs).

- **Table Detection** — Uses a three-stage strategy (rectangle-based → line-based → heuristic) to build `Table` structures and format them as Markdown. Spans [[`src/tables/detect_rects.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_rects.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_rects.rs), [[`src/tables/detect_lines.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_lines.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_lines.rs), [[`src/tables/detect_heuristic.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_heuristic.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_heuristic.rs), [[`src/tables/grid.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/grid.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/grid.rs), and [[`src/tables/format.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/format.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/format.rs).

- **Markdown Generation** — Analyzes font statistics to construct headings, lists, code blocks, and tables with post-processing. Located in [[`src/markdown/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/mod.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/mod.rs) and [[`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs).

- **Public API** — Provides convenience functions like `process_pdf`, `detect_pdf`, and `classify_pdf_mem` plus the `PdfOptions` builder. Defined in [[`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs).

- **Language Bindings** — Python (`pdf_inspector`), Node.js (`@firecrawl/pdf-inspector`), and WebAssembly. See [[`src/python.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs), [[`napi/README.md`](https://github.com/firecrawl/pdf-inspector/blob/main/napi/README.md)](https://github.com/firecrawl/pdf-inspector/blob/main/napi/README.md), and [[`wasm/README.md`](https://github.com/firecrawl/pdf-inspector/blob/main/wasm/README.md)](https://github.com/firecrawl/pdf-inspector/blob/main/wasm/README.md).

- **CLI Binaries** — `pdf2md` for full extraction and `detect-pdf` for metadata-only operations. In [[`src/bin/pdf2md.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs) and [[`src/bin/detect_pdf.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/detect_pdf.rs)](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/detect_pdf.rs).

## How to Start Contributing to pdf-inspector

### 1. Pick Your Subsystem

**Table detection** (`src/tables/*`) is currently the most active area for contributions, with opportunities to improve heuristic detection and handling of sparse layouts. The **detection logic** ([`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs)) is stable but welcomes performance optimizations. Language bindings and CLI usability improvements are also valuable entry points.

### 2. Set Up Your Development Environment

```bash

# Fork and clone the repository

git clone https://github.com/<your-username>/pdf-inspector.git
cd pdf-inspector

# Build the project to verify your toolchain

cargo build --release

```

### 3. Follow the Contribution Workflow

The project enforces strict quality standards through CI. Your PR must pass all checks defined in [`.github/workflows/ci.yml`](https://github.com/firecrawl/pdf-inspector/blob/main/.github/workflows/ci.yml).

```bash

# Create a feature branch

git checkout -b improve-table-heuristics

# Make your changes, then run formatting

cargo fmt

# Run linting (warnings are treated as errors)

cargo clippy -- -D warnings

# Execute the full test suite

cargo test

```

The repository contains **267+ unit tests** and **73+ integration tests**. Every contribution requires corresponding test additions or modifications.

### 4. Submit Your Pull Request

Push your branch and open a PR on GitHub. Include:
- Clear description of the problem and solution
- Benchmarks or examples if applicable
- Updated documentation in `docs/*.md` for public API changes

## Working with the Public API

Understanding the API helps you write better tests and contributions. Here are practical examples for each supported interface:

### Rust

```rust
use pdf_inspector::{process_pdf, PdfOptions, ProcessMode};

let opts = PdfOptions::new()
    .mode(ProcessMode::Full)
    .pages([1, 3, 5]);

let result = process_pdf_with_options("report.pdf", opts).unwrap();
println!("PDF type: {:?}", result.pdf_type);
if let Some(md) = result.markdown {
    println!("Markdown output:\n{md}");
}

```

### Python

```python
import pdf_inspector

# Fast detection only

info = pdf_inspector.detect_pdf("invoice.pdf")
print(info.pdf_type)          # "text_based", "scanned", etc.

# Full extraction

res = pdf_inspector.process_pdf("contract.pdf")
print(res.markdown)

```

### Node.js

```javascript
import { readFileSync } from 'fs';
import { processPdf } from '@firecrawl/pdf-inspector';

const buffer = readFileSync('manual.pdf');
const result = processPdf(buffer);
console.log(result.pdfType);
console.log(result.markdown);

```

### CLI

```bash

# Detection only

detect-pdf financials.pdf --json

# Full Markdown conversion

pdf2md research-paper.pdf > paper.md

```

## Advanced Contribution: Hybrid OCR Pipelines

For contributors working on OCR integration, the per-page extraction API is particularly relevant:

```rust
let pages = pdf_inspector::extract_pages_markdown("mixed.pdf", Some(&[0, 2]))?;
for page in pages.pages {
    if page.needs_ocr {
        // Send to OCR service
    } else {
        println!("Page {} markdown:\n{}", page.page + 1, page.markdown);
    }
}

```

## Key Files Every Contributor Should Know

| File | Purpose |
|------|---------|
| [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) | Public API and `PdfOptions` builder |
| [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) | Fast PDF type detection |
| [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) | Text extraction coordination |
| [`src/extractor/content_stream.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs) | PDF operator state machine |
| [`src/tables/detect_heuristic.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_heuristic.rs) | Alignment-based table detection |
| [`src/tables/grid.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/grid.rs) | Cell assignment and boundary logic |
| [`src/markdown/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/mod.rs) | Markdown generation entry point |
| [`src/python.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs) | PyO3 Python bindings |
| [`src/bin/pdf2md.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs) | CLI conversion tool |

## Summary

- **pdf-inspector** is organized into independent stages (detection, extraction, layout, tables, Markdown generation) that you can contribute to separately.
- **Table detection** (`src/tables/*`) offers the most active contribution opportunities, while **detection performance** welcomes optimizations.
- **Quality gates**: `cargo fmt`, `cargo clippy -- -D warnings`, and `cargo test` must all pass—267+ unit and 73+ integration tests enforce correctness.
- **Documentation updates** in `docs/*.md` are required for public API changes.
- **Multiple interfaces** (Rust, Python, Node.js, WebAssembly, CLI) mean contributions can improve accessibility across ecosystems.

## Frequently Asked Questions

### What Rust knowledge is required to contribute to pdf-inspector?

You need intermediate Rust skills including ownership, lifetimes, and working with `cargo`. Familiarity with PDF specifications (ISO 32000) or parsing techniques helps for extraction and detection work, but you can start with simpler improvements to CLI tools or bindings.

### How do I add a new test for my contribution?

Add unit tests in `src/` files alongside your changes (using `#[cfg(test)]` modules) or integration tests in the `tests/` directory. Run `cargo test` to verify. The CI enforces that all tests pass and code coverage remains adequate.

### Can I contribute without modifying Rust code?

Yes. Documentation improvements, bug reports with reproducible test cases, and enhancements to Python or Node.js bindings are valuable. The WASM and CLI interfaces also accept contributions focused on usability and cross-platform packaging.

### What is the review process for pdf-inspector PRs?

Maintainers review for correctness, test coverage, performance impact, and adherence to the existing code style. The CI pipeline automatically runs formatting, linting, and tests. Most substantive changes receive feedback within a few days, with faster turnaround for documentation fixes and minor bug corrections.