Where to Find Documentation for firecrawl/pdf‑inspector: A Complete Guide to the Official Docs
The official documentation for firecrawl pdf‑inspector lives entirely inside the repository at README.md and the docs/ folder, with Python API docs at docs/python.md, Rust API docs at docs/rust-api.md, and specialized guides for debugging, benchmarking, and publishing.
If you're building with firecrawl pdf‑inspector, you'll want to know exactly where to find authoritative reference material. Unlike projects that scatter docs across external wikis, this Rust‑based PDF extraction engine keeps everything version‑controlled and browsable in the source tree. Whether you're calling binaries from the shell, importing the Python wrapper, or integrating the Rust crate directly, the documentation structure maps cleanly to each use case.
Primary Documentation Locations
The project maintains seven core documentation files that cover the full surface area of the tool:
| Document | Purpose | Location |
|---|---|---|
| README.md | Quick‑start, CLI usage, build instructions | Repository root |
| docs/python.md | Python wrapper installation and API reference | docs/python.md |
| docs/rust-api.md | Rust library public API and type definitions | docs/rust-api.md |
| docs/debugging.md | Verbose logging configuration for components | docs/debugging.md |
| docs/benchmarking.md | Performance testing and result interpretation | docs/benchmarking.md |
| docs/publishing.md | Crate release procedures and CI checks | docs/publishing.md |
| site/index.html | Interactive HTML rendering of all markdown docs | site/index.html |
All paths are relative to https://github.com/firecrawl/pdf-inspector/blob/main/.
CLI Documentation: README.md
The README.md serves as the entry point for command‑line users. It documents the two primary binaries installed with the crate:
pdf2md— Converts PDF files to Markdown or structured JSONdetect-pdf— Classifies PDF types and detects tiled‑scan patterns
Basic CLI Usage
# Extract PDF to Markdown
pdf2md my-document.pdf
# Output structured JSON
pdf2md my-document.pdf --json > output.json
# Detect PDF type with analysis
detect-pdf my-document.pdf
detect-pdf my-document.pdf --json --analyze
The README also covers building from source with cargo build --release and running the test suite.
Python API Documentation
For Python developers, docs/python.md provides the complete reference for the pdf_inspector package distributed via PyPI.
Installing and Using the Python Wrapper
pip install pdf-inspector
from pdf_inspector import PDFInspector
# Initialize with default options
inspector = PDFInspector()
# Convert to Markdown
markdown = inspector.pdf_to_md("my-document.pdf")
print(markdown)
# Get structured JSON representation
json_data = inspector.pdf_to_json("my-document.pdf")
print(json_data)
The Python wrapper delegates to the underlying Rust implementation, so performance characteristics match the native crate.
Rust API Documentation
Systems programmers working directly with the crate should consult docs/rust-api.md. This document details the primary entry point process_pdf_with_options defined in src/lib.rs, along with configuration through the ProcessOptions struct.
Rust Library Integration
use pdf_inspector::process_pdf_with_options;
use pdf_inspector::ProcessOptions;
fn main() -> Result<(), Box<dyn std::error::Error>> {
let opts = ProcessOptions::default();
let result = process_pdf_with_options("my-document.pdf", opts)?;
println!("{}", result.markdown);
Ok(())
}
The Rust docs also explain how to access individual pipeline stages—layout analysis in src/extractor/layout.rs, table detection in src/tables/, and Markdown formatting in src/markdown/—for custom extraction workflows.
Specialized Development Guides
Beyond API references, three documents support ** advanced usage and maintenance**:
Debugging Guide (docs/debugging.md)
Configure component‑level logging with RUST_LOG environment variables:
# Debug layout engine decisions
RUST_LOG=pdf_inspector::extractor::layout=debug cargo run --bin pdf2md -- my-document.pdf
# Trace table detection heuristics
RUST_LOG=pdf_inspector::tables=debug cargo run --bin pdf2md -- my-document.pdf
Other log targets include pdf_inspector::extractor for the main pipeline and pdf_inspector::detector for PDF classification.
Benchmarking Guide (docs/benchmarking.md)
Describes how to run the built‑in benchmark suite and interpret throughput metrics for large‑document processing.
Publishing Guide (docs/publishing.md)
Maintainer documentation covering version bumping in Cargo.toml, changelog updates, and CI‑verified release tags.
Browsing Documentation Offline
The site/ directory contains a pre‑built HTML site generated from the markdown sources. Open site/index.html in a browser for offline navigation with search and syntax highlighting—useful when working without network access or reviewing documentation for a specific git checkout.
Key Source Files Referenced in Documentation
The documentation consistently points to these implementation files for readers who need to understand internal behavior:
| File | Role |
|---|---|
src/lib.rs |
Public API: process_pdf_with_options, encoding detection |
src/detector.rs |
PDF classification and page sampling |
src/types.rs |
Core structures: TextItem, TextLine, PdfRect |
src/extractor/layout.rs |
Column detection and layout classification |
src/tables/*.rs |
Table detection and formatting |
src/markdown/*.rs |
Markdown generation and post‑processing |
These paths appear throughout the documentation when explaining edge cases or extension points.
Summary
- Entry point: Start with
README.mdfor CLI usage and project overview - Python developers: Reference
docs/python.mdfor wrapper installation and API - Rust developers: Use
docs/rust-api.mdfor crate integration andProcessOptions - Troubleshooting: Enable verbose logging via
docs/debugging.md - Offline reading: Open
site/index.htmlfor rendered HTML documentation - All docs are version‑controlled: Check the corresponding git tag for your installed version
Frequently Asked Questions
Does firecrawl pdf‑inspector have external documentation or a docs website?
No. All official documentation is maintained inside the repository under README.md and the docs/ directory. The site/index.html file provides a browsable HTML rendering, but there is no separate documentation site or wiki. This ensures documentation stays synchronized with code changes.
How do I access documentation for an older version of pdf‑inspector?
Check out the git tag corresponding to your installed version. The docs/ folder at that commit reflects the accurate API and behavior for that release. The repository uses semantic versioning, so tags like v0.5.0 map directly to documented feature sets.
What documentation exists for extending the PDF extraction pipeline?
The Rust API documentation in docs/rust-api.md describes how to use lower‑level components. For deep customization, read the source files referenced throughout the docs: src/extractor/layout.rs for layout algorithms, src/tables/*.rs for table detection, and src/markdown/*.rs for output formatting. The debugging guide helps trace execution through these stages.
Is there API reference documentation generated from source code comments?
The project uses hand‑written markdown documentation rather than auto‑generated API docs. The docs/rust-api.md and docs/python.md files provide curated explanations with examples. For function‑level details, consult the source files directly—the public API surface in src/lib.rs is intentionally small and well‑documented inline.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →