Where to Find Documentation for firecrawl/pdf‑inspector: A Complete Guide to the Official Docs

The official documentation for firecrawl pdf‑inspector lives entirely inside the repository at README.md and the docs/ folder, with Python API docs at docs/python.md, Rust API docs at docs/rust-api.md, and specialized guides for debugging, benchmarking, and publishing.

If you're building with firecrawl pdf‑inspector, you'll want to know exactly where to find authoritative reference material. Unlike projects that scatter docs across external wikis, this Rust‑based PDF extraction engine keeps everything version‑controlled and browsable in the source tree. Whether you're calling binaries from the shell, importing the Python wrapper, or integrating the Rust crate directly, the documentation structure maps cleanly to each use case.

Primary Documentation Locations

The project maintains seven core documentation files that cover the full surface area of the tool:

Document Purpose Location
README.md Quick‑start, CLI usage, build instructions Repository root
docs/python.md Python wrapper installation and API reference docs/python.md
docs/rust-api.md Rust library public API and type definitions docs/rust-api.md
docs/debugging.md Verbose logging configuration for components docs/debugging.md
docs/benchmarking.md Performance testing and result interpretation docs/benchmarking.md
docs/publishing.md Crate release procedures and CI checks docs/publishing.md
site/index.html Interactive HTML rendering of all markdown docs site/index.html

All paths are relative to https://github.com/firecrawl/pdf-inspector/blob/main/.

CLI Documentation: README.md

The README.md serves as the entry point for command‑line users. It documents the two primary binaries installed with the crate:

  • pdf2md — Converts PDF files to Markdown or structured JSON
  • detect-pdf — Classifies PDF types and detects tiled‑scan patterns

Basic CLI Usage


# Extract PDF to Markdown

pdf2md my-document.pdf

# Output structured JSON

pdf2md my-document.pdf --json > output.json

# Detect PDF type with analysis

detect-pdf my-document.pdf
detect-pdf my-document.pdf --json --analyze

The README also covers building from source with cargo build --release and running the test suite.

Python API Documentation

For Python developers, docs/python.md provides the complete reference for the pdf_inspector package distributed via PyPI.

Installing and Using the Python Wrapper

pip install pdf-inspector
from pdf_inspector import PDFInspector

# Initialize with default options

inspector = PDFInspector()

# Convert to Markdown

markdown = inspector.pdf_to_md("my-document.pdf")
print(markdown)

# Get structured JSON representation

json_data = inspector.pdf_to_json("my-document.pdf")
print(json_data)

The Python wrapper delegates to the underlying Rust implementation, so performance characteristics match the native crate.

Rust API Documentation

Systems programmers working directly with the crate should consult docs/rust-api.md. This document details the primary entry point process_pdf_with_options defined in src/lib.rs, along with configuration through the ProcessOptions struct.

Rust Library Integration

use pdf_inspector::process_pdf_with_options;
use pdf_inspector::ProcessOptions;

fn main() -> Result<(), Box<dyn std::error::Error>> {
    let opts = ProcessOptions::default();
    let result = process_pdf_with_options("my-document.pdf", opts)?;
    println!("{}", result.markdown);
    Ok(())
}

The Rust docs also explain how to access individual pipeline stages—layout analysis in src/extractor/layout.rs, table detection in src/tables/, and Markdown formatting in src/markdown/—for custom extraction workflows.

Specialized Development Guides

Beyond API references, three documents support ** advanced usage and maintenance**:

Debugging Guide (docs/debugging.md)

Configure component‑level logging with RUST_LOG environment variables:


# Debug layout engine decisions

RUST_LOG=pdf_inspector::extractor::layout=debug cargo run --bin pdf2md -- my-document.pdf

# Trace table detection heuristics

RUST_LOG=pdf_inspector::tables=debug cargo run --bin pdf2md -- my-document.pdf

Other log targets include pdf_inspector::extractor for the main pipeline and pdf_inspector::detector for PDF classification.

Benchmarking Guide (docs/benchmarking.md)

Describes how to run the built‑in benchmark suite and interpret throughput metrics for large‑document processing.

Publishing Guide (docs/publishing.md)

Maintainer documentation covering version bumping in Cargo.toml, changelog updates, and CI‑verified release tags.

Browsing Documentation Offline

The site/ directory contains a pre‑built HTML site generated from the markdown sources. Open site/index.html in a browser for offline navigation with search and syntax highlighting—useful when working without network access or reviewing documentation for a specific git checkout.

Key Source Files Referenced in Documentation

The documentation consistently points to these implementation files for readers who need to understand internal behavior:

File Role
src/lib.rs Public API: process_pdf_with_options, encoding detection
src/detector.rs PDF classification and page sampling
src/types.rs Core structures: TextItem, TextLine, PdfRect
src/extractor/layout.rs Column detection and layout classification
src/tables/*.rs Table detection and formatting
src/markdown/*.rs Markdown generation and post‑processing

These paths appear throughout the documentation when explaining edge cases or extension points.

Summary

  • Entry point: Start with README.md for CLI usage and project overview
  • Python developers: Reference docs/python.md for wrapper installation and API
  • Rust developers: Use docs/rust-api.md for crate integration and ProcessOptions
  • Troubleshooting: Enable verbose logging via docs/debugging.md
  • Offline reading: Open site/index.html for rendered HTML documentation
  • All docs are version‑controlled: Check the corresponding git tag for your installed version

Frequently Asked Questions

Does firecrawl pdf‑inspector have external documentation or a docs website?

No. All official documentation is maintained inside the repository under README.md and the docs/ directory. The site/index.html file provides a browsable HTML rendering, but there is no separate documentation site or wiki. This ensures documentation stays synchronized with code changes.

How do I access documentation for an older version of pdf‑inspector?

Check out the git tag corresponding to your installed version. The docs/ folder at that commit reflects the accurate API and behavior for that release. The repository uses semantic versioning, so tags like v0.5.0 map directly to documented feature sets.

What documentation exists for extending the PDF extraction pipeline?

The Rust API documentation in docs/rust-api.md describes how to use lower‑level components. For deep customization, read the source files referenced throughout the docs: src/extractor/layout.rs for layout algorithms, src/tables/*.rs for table detection, and src/markdown/*.rs for output formatting. The debugging guide helps trace execution through these stages.

Is there API reference documentation generated from source code comments?

The project uses hand‑written markdown documentation rather than auto‑generated API docs. The docs/rust-api.md and docs/python.md files provide curated explanations with examples. For function‑level details, consult the source files directly—the public API surface in src/lib.rs is intentionally small and well‑documented inline.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →