# How to Build the `pdf2md` and `detect-pdf` CLI Tools from Source: A Complete Guide

> Learn to build pdf2md and detect-pdf CLI tools from source. This guide covers compiling the firecrawl pdf-inspector Rust project with cargo build --release for optimized binaries.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-06

---

**Build the `pdf2md` and `detect-pdf` CLI tools by compiling the Rust source in `firecrawl/pdf-inspector` using `cargo build --release`, which produces optimized binaries in `target/release/`.**

The **pdf-inspector** repository from Firecrawl provides fast, accurate PDF-to-Markdown conversion and PDF type detection. If you want the latest features or need to modify the behavior, building these CLI tools from source is straightforward. This guide walks through the complete build process for the `pdf2md` and `detect-pdf` binaries.

## Prerequisites: Install Rust

Before building, you need a working Rust toolchain. The recommended approach is [rustup](https://rustup.rs):

```bash
curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh
source $HOME/.cargo/env
rustc --version   # verify installation

```

Any recent stable Rust version will work.

## Clone the pdf-inspector Repository

```bash
git clone https://github.com/firecrawl/pdf-inspector.git
cd pdf-inspector

```

The source structure places CLI entry points in `src/bin/` and shared logic in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs).

## Build the Binaries with Cargo

### Release Build (Recommended)

For production use, compile with optimizations enabled:

```bash
cargo build --release

```

This creates:
- `target/release/pdf2md`
- `target/release/detect-pdf`

These are fully optimized binaries suitable for processing large PDF collections.

### Debug Build (For Development)

For faster compilation during development or debugging:

```bash
cargo build

```

Binaries appear in `target/debug/` with full debug symbols but slower execution.

## Run the Built Binaries

Execute directly without system-wide installation:

```bash

# PDF to Markdown conversion

./target/release/pdf2md sample.pdf --compact > output.md

# PDF type detection with JSON output

./target/release/detect-pdf sample.pdf --json

```

Both binaries support the same core flags:
- `--select-pages 1,3,5-10` — process specific page ranges
- `--password <pw>` — handle encrypted PDFs
- `--json` — emit structured JSON instead of human-readable output

## Install to PATH (Optional)

To make the tools available anywhere in your shell:

```bash
cargo install --path .

```

This installs both `pdf2md` and `detect-pdf` to `~/.cargo/bin/`. Ensure this directory is on your `$PATH`.

## Source Code Architecture

Understanding the key files helps when customizing or debugging:

| Path | Purpose |
|------|---------|
| [`src/bin/pdf2md.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs) | CLI entry point for Markdown conversion |
| [`src/bin/detect_pdf.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/detect_pdf.rs) | CLI entry point for PDF type detection and analysis |
| [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) | Public API (`process_pdf_with_options`, `PdfOptions`) shared by both tools |
| [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) | Fast PDF classification (text-based vs. scanned) |
| `src/extractor/` | Core text extraction and layout pipeline |
| `src/markdown/` | Markdown generation, heading detection, list handling |
| [`Cargo.toml`](https://github.com/firecrawl/pdf-inspector/blob/main/Cargo.toml) | Crate metadata with `[[bin]]` entries defining both executables |

The binaries in `src/bin/` are thin wrappers. All heavy lifting happens in the library modules, making the tools easy to test and extend.

## Usage Examples

```bash

# Compact Markdown output (no metadata headers)

./target/release/pdf2md document.pdf --compact > clean.md

# Full JSON result with layout analysis

./target/release/pdf2md report.pdf --json | jq '.pages[].text'

# Detect PDF type with full structural analysis

./target/release/detect-pdf scanned.pdf --analyze --json

# Process password-protected document

./target/release/pdf2md encrypted.pdf --password "userpass" > output.md

```

## Summary

- **Install Rust** via rustup to get the `cargo` build tool
- **Clone** `firecrawl/pdf-inspector` from GitHub
- **Build** with `cargo build --release` for optimized binaries in `target/release/`
- **Key files**: [`src/bin/pdf2md.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/pdf2md.rs), [`src/bin/detect_pdf.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/bin/detect_pdf.rs), and [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) define the tool functionality
- **Install** permanently with `cargo install --path .` to add to your PATH

## Frequently Asked Questions

### Why does the release build take longer than the debug build?

Release builds enable **LLVM optimizations** that significantly improve runtime performance. According to the [`Cargo.toml`](https://github.com/firecrawl/pdf-inspector/blob/main/Cargo.toml) configuration, these optimizations include link-time optimization and strip symbols, making the binary faster and smaller. For one-off conversions, debug builds work fine; for batch processing, always use `--release`.

### Can I build only one of the two binaries?

Yes. The [`Cargo.toml`](https://github.com/firecrawl/pdf-inspector/blob/main/Cargo.toml) defines both binaries explicitly with `[[bin]]` entries. To build just `pdf2md`, use `cargo build --release --bin pdf2md`. Similarly, `cargo build --release --bin detect-pdf` builds only the detector tool.

### Where are the core algorithms implemented?

The extraction and layout logic lives outside the CLI binaries. Check `src/extractor/` for text extraction, `src/markdown/` for Markdown generation, and [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) for fast PDF classification. Both `pdf2md` and `detect-pdf` call into [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) for these capabilities.