# Pre‑Built Docker Images for pdf‑inspector: What You Need to Know

> Discover if pre-built Docker images exist for firecrawl/pdf-inspector. Learn why you must build your own container from source for this tool.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: getting-started
- Published: 2026-08-04

---

**No, the firecrawl/pdf‑inspector repository does not provide pre‑built Docker images—users must build from source and create their own container.**

The pdf‑inspector project is a Rust‑based toolkit for PDF analysis and Markdown conversion. Despite its utility for extracting structured text and detecting encoding issues, the maintainers have not published official Docker images to Docker Hub or GitHub Container Registry. This article explains why and shows you exactly how to containerize the binaries yourself.

## Why pdf‑inspector Ships Without Pre‑Built Images

A thorough search of the source tree reveals no `Dockerfile`, [`docker-compose.yml`](https://github.com/firecrawl/pdf-inspector/blob/main/docker-compose.yml), or image‑build scripts. The repository focuses on native binary distribution through Cargo.

The only Docker reference appears in [`.github/workflows/publish.yml`](https://github.com/firecrawl/pdf-inspector/blob/main/.github/workflows/publish.yml), where the CI pipeline spins up a temporary container solely for testing. This ephemeral usage does not produce or publish a reusable image—it validates the build in an isolated environment and then discards it.

Consequently, containerized deployments require a manual build step. The good news: pdf‑inspector compiles to static binaries with minimal system dependencies, making custom images straightforward to create.

## Building pdf‑inspector Binaries Locally

Before containerizing, compile the Rust project on your host machine or in a multi‑stage build.

```bash

# Clone the repository

git clone https://github.com/firecrawl/pdf-inspector.git
cd pdf-inspector

# Build optimized release binaries

cargo build --release

```

This produces two key executables:

- `target/release/pdf2md` — Converts PDFs to Markdown with layout preservation
- `target/release/detect-pdf` — Classifies PDFs as text‑based, scanned, or mixed

These binaries link against `libgcc` and `libstdc++` but otherwise have no heavy runtime dependencies.

## Creating a Minimal Custom Docker Image

Package the compiled binaries into a lightweight base image. Alpine Linux keeps the final size small while providing the required C++ standard libraries.

```dockerfile

# Dockerfile – custom container for pdf‑inspector

FROM alpine:3.20 AS runtime
RUN apk add --no-cache libgcc libstdc++

# Copy pre‑built binaries from host (or use multi‑stage build)

COPY target/release/pdf2md /usr/local/bin/pdf2md
COPY target/release/detect-pdf /usr/local/bin/

# Default to the Markdown converter

ENTRYPOINT ["pdf2md"]

```

Build and tag your image:

```bash
docker build -t firecrawl/pdf-inspector:latest .

```

## Running pdf‑inspector Inside Docker

Mount PDFs as volumes and invoke either binary directly.

```bash

# Convert a PDF to Markdown (JSON output)

docker run --rm \
  -v "$(pwd)":/data \
  firecrawl/pdf-inspector:latest \
  pdf2md /data/example.pdf --json > output.json

# Detect PDF type

docker run --rm \
  -v "$(pwd)":/data \
  firecrawl/pdf-inspector:latest \
  detect-pdf /data/example.pdf

```

For a simplified command, override the entrypoint:

```bash
docker run --rm --entrypoint detect-pdf \
  -v "$(pwd)":/data \
  firecrawl/pdf-inspector:latest \
  /data/example.pdf

```

## Multi‑Stage Build Alternative (No Host Compilation)

Eliminate host‑side Rust toolchains with a two‑stage Dockerfile that builds inside the container:

```dockerfile

# Stage 1: Build

FROM rust:1.82-alpine AS builder
RUN apk add --no-cache build-base libgcc libstdc++
WORKDIR /src
COPY . .
RUN cargo build --release

# Stage 2: Runtime

FROM alpine:3.20
RUN apk add --no-cache libgcc libstdc++
COPY --from=builder /src/target/release/pdf2md /usr/local/bin/
COPY --from=builder /src/target/release/detect-pdf /usr/local/bin/
ENTRYPOINT ["pdf2md"]

```

Build with:

```bash
docker build -t firecrawl/pdf-inspector:latest .

```

This approach yields a sub‑50 MB image containing only the compiled binaries and their shared libraries.

## Core Components You Will Containerize

Understanding the source structure helps when customizing your build:

| File | Purpose |
|------|---------|
| [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) | Public API with `process_pdf_with_options` and encoding‑issue detection |
| [`src/detector.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/detector.rs) | PDF classification (text‑based, scanned, mixed) |
| [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) | Orchestrates content‑stream text extraction |
| [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs) | Layout‑aware Markdown generation |
| [`src/tables/detect_rects.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_rects.rs) | Rectangle‑based table detection |
| [`src/tables/detect_lines.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_lines.rs) | Line‑based table fallback |
| [`src/tables/detect_heuristic.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/detect_heuristic.rs) | Heuristic table detection |
| [`examples/basic_usage.py`](https://github.com/firecrawl/pdf-inspector/blob/main/examples/basic_usage.py) | Python binding examples |

These modules compile into the `pdf2md` and `detect-pdf` binaries you will copy into your image.

## Summary

- **No pre‑built Docker images exist** for firecrawl/pdf‑inspector—verify by checking for missing Dockerfile and CI publish steps in [`.github/workflows/publish.yml`](https://github.com/firecrawl/pdf-inspector/blob/main/.github/workflows/publish.yml)
- **Build from source** using `cargo build --release` to obtain `pdf2md` and `detect-pdf`
- **Create minimal images** using Alpine or `rust:slim` bases with `libgcc` and `libstdc++`
- **Use multi‑stage builds** to avoid installing Rust toolchains on deployment hosts
- **Mount volumes** for PDF input/output when running containers

## Frequently Asked Questions

### Does pdf‑inspector publish images to Docker Hub?

No. The repository contains no image‑push workflows and no Dockerfile. The only Docker usage in [`.github/workflows/publish.yml`](https://github.com/firecrawl/pdf-inspector/blob/main/.github/workflows/publish.yml) runs temporary test containers that are destroyed after CI completion.

### What base image works best for pdf‑inspector?

**Alpine 3.20** or **Debian Slim** both work. Alpine produces smaller images (~15 MB plus binaries) but requires installing `libgcc` and `libstdc++`. Debian requires no extra packages but yields larger images.

### Can I use the Python bindings inside Docker?

Yes. Install the compiled Rust library and Python wrapper in the same image, or use the provided [`examples/basic_usage.py`](https://github.com/firecrawl/pdf-inspector/blob/main/examples/basic_usage.py) as a template for your own Dockerfile that includes Python runtime dependencies.

### Why not add an official Dockerfile to the repository?

That decision rests with the maintainers. The current design prioritizes Cargo‑based distribution and keeps the repository focused on source code. Community pull requests adding optional container support may be welcomed—check the project's contribution guidelines.