How to Contribute to the pdf-inspector Project: A Complete Guide for Rust Developers
Contributing to pdf-inspector involves selecting a subsystem to improve, running the test suite, following the PR workflow, and submitting changes through GitHub.
The pdf-inspector project is a fast Rust library maintained by Firecrawl that performs PDF type detection, text extraction, and Markdown conversion. Its modular architecture deliberately separates concerns into independent stages, making it straightforward for contributors to work on specific components without understanding the entire pipeline. This guide walks you through how to contribute to pdf-inspector effectively, based on the actual source code structure.
Understanding the pdf-inspector Architecture
Before contributing, you need to understand how the codebase is organized. The project is split into distinct stages, each with clear responsibilities and dedicated source files.
Core Subsystems and File Locations
-
Detection — Quickly classifies PDFs as
TextBased,Scanned,ImageBased, orMixedusing only the X-ref table and lightweight content stream scanning. Core logic resides in [src/detector.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs). -
Extraction — Walks PDF operators (
Tj,TJ,Td,Tm, etc.) to produce positionedTextItems andPdfRects, handling fonts, ToUnicode CMaps, and XObjects. Implemented in [src/extractor/mod.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) and [src/extractor/content_stream.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/content_stream.rs). -
Layout — Detects multi-column reading order and prepares data for table detection. Found in [
src/extractor/layout.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/layout.rs). -
Table Detection — Uses a three-stage strategy (rectangle-based → line-based → heuristic) to build
Tablestructures and format them as Markdown. Spans [src/tables/detect_rects.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_rects.rs), [src/tables/detect_lines.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_lines.rs), [src/tables/detect_heuristic.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_heuristic.rs), [src/tables/grid.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/grid.rs), and [src/tables/format.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/format.rs). -
Markdown Generation — Analyzes font statistics to construct headings, lists, code blocks, and tables with post-processing. Located in [
src/markdown/mod.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/mod.rs) and [src/markdown/convert.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs). -
Public API — Provides convenience functions like
process_pdf,detect_pdf, andclassify_pdf_memplus thePdfOptionsbuilder. Defined in [src/lib.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs). -
Language Bindings — Python (
pdf_inspector), Node.js (@firecrawl/pdf-inspector), and WebAssembly. See [src/python.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/python.rs), [napi/README.md](https://github.com/firecrawl/pdf-inspector/blob/main/napi/README.md), and [wasm/README.md](https://github.com/firecrawl/pdf-inspector/blob/main/wasm/README.md). -
CLI Binaries —
pdf2mdfor full extraction anddetect-pdffor metadata-only operations. In [src/bin/pdf2md.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs) and [src/bin/detect_pdf.rs](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/detect_pdf.rs).
How to Start Contributing to pdf-inspector
1. Pick Your Subsystem
Table detection (src/tables/*) is currently the most active area for contributions, with opportunities to improve heuristic detection and handling of sparse layouts. The detection logic (src/detector.rs) is stable but welcomes performance optimizations. Language bindings and CLI usability improvements are also valuable entry points.
2. Set Up Your Development Environment
# Fork and clone the repository
git clone https://github.com/<your-username>/pdf-inspector.git
cd pdf-inspector
# Build the project to verify your toolchain
cargo build --release
3. Follow the Contribution Workflow
The project enforces strict quality standards through CI. Your PR must pass all checks defined in .github/workflows/ci.yml.
# Create a feature branch
git checkout -b improve-table-heuristics
# Make your changes, then run formatting
cargo fmt
# Run linting (warnings are treated as errors)
cargo clippy -- -D warnings
# Execute the full test suite
cargo test
The repository contains 267+ unit tests and 73+ integration tests. Every contribution requires corresponding test additions or modifications.
4. Submit Your Pull Request
Push your branch and open a PR on GitHub. Include:
- Clear description of the problem and solution
- Benchmarks or examples if applicable
- Updated documentation in
docs/*.mdfor public API changes
Working with the Public API
Understanding the API helps you write better tests and contributions. Here are practical examples for each supported interface:
Rust
use pdf_inspector::{process_pdf, PdfOptions, ProcessMode};
let opts = PdfOptions::new()
.mode(ProcessMode::Full)
.pages([1, 3, 5]);
let result = process_pdf_with_options("report.pdf", opts).unwrap();
println!("PDF type: {:?}", result.pdf_type);
if let Some(md) = result.markdown {
println!("Markdown output:\n{md}");
}
Python
import pdf_inspector
# Fast detection only
info = pdf_inspector.detect_pdf("invoice.pdf")
print(info.pdf_type) # "text_based", "scanned", etc.
# Full extraction
res = pdf_inspector.process_pdf("contract.pdf")
print(res.markdown)
Node.js
import { readFileSync } from 'fs';
import { processPdf } from '@firecrawl/pdf-inspector';
const buffer = readFileSync('manual.pdf');
const result = processPdf(buffer);
console.log(result.pdfType);
console.log(result.markdown);
CLI
# Detection only
detect-pdf financials.pdf --json
# Full Markdown conversion
pdf2md research-paper.pdf > paper.md
Advanced Contribution: Hybrid OCR Pipelines
For contributors working on OCR integration, the per-page extraction API is particularly relevant:
let pages = pdf_inspector::extract_pages_markdown("mixed.pdf", Some(&[0, 2]))?;
for page in pages.pages {
if page.needs_ocr {
// Send to OCR service
} else {
println!("Page {} markdown:\n{}", page.page + 1, page.markdown);
}
}
Key Files Every Contributor Should Know
| File | Purpose |
|---|---|
src/lib.rs |
Public API and PdfOptions builder |
src/detector.rs |
Fast PDF type detection |
src/extractor/mod.rs |
Text extraction coordination |
src/extractor/content_stream.rs |
PDF operator state machine |
src/tables/detect_heuristic.rs |
Alignment-based table detection |
src/tables/grid.rs |
Cell assignment and boundary logic |
src/markdown/mod.rs |
Markdown generation entry point |
src/python.rs |
PyO3 Python bindings |
src/bin/pdf2md.rs |
CLI conversion tool |
Summary
- pdf-inspector is organized into independent stages (detection, extraction, layout, tables, Markdown generation) that you can contribute to separately.
- Table detection (
src/tables/*) offers the most active contribution opportunities, while detection performance welcomes optimizations. - Quality gates:
cargo fmt,cargo clippy -- -D warnings, andcargo testmust all pass—267+ unit and 73+ integration tests enforce correctness. - Documentation updates in
docs/*.mdare required for public API changes. - Multiple interfaces (Rust, Python, Node.js, WebAssembly, CLI) mean contributions can improve accessibility across ecosystems.
Frequently Asked Questions
What Rust knowledge is required to contribute to pdf-inspector?
You need intermediate Rust skills including ownership, lifetimes, and working with cargo. Familiarity with PDF specifications (ISO 32000) or parsing techniques helps for extraction and detection work, but you can start with simpler improvements to CLI tools or bindings.
How do I add a new test for my contribution?
Add unit tests in src/ files alongside your changes (using #[cfg(test)] modules) or integration tests in the tests/ directory. Run cargo test to verify. The CI enforces that all tests pass and code coverage remains adequate.
Can I contribute without modifying Rust code?
Yes. Documentation improvements, bug reports with reproducible test cases, and enhancements to Python or Node.js bindings are valuable. The WASM and CLI interfaces also accept contributions focused on usability and cross-platform packaging.
What is the review process for pdf-inspector PRs?
Maintainers review for correctness, test coverage, performance impact, and adherence to the existing code style. The CI pipeline automatically runs formatting, linting, and tests. Most substantive changes receive feedback within a few days, with faster turnaround for documentation fixes and minor bug corrections.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →