# How to Customize the Markdown Output Format in pdf-inspector

> Customize Markdown output in pdf-inspector using ProcessOptions, CLI JSON output, or by forking the repository. Learn how to tailor your PDF to Markdown conversions for specific needs.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-08

---

**You can customize the Markdown output format in pdf-inspector by passing configuration options via the `ProcessOptions` struct, consuming the `--json` CLI output to re-render content externally, or forking the repository to modify the conversion logic in [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs) and related modules.**

The pdf-inspector library from Firecrawl converts PDF documents to structured Markdown through a modular Rust pipeline. Whether you need to adjust heading prefixes, change list bullet styles, or implement custom table formatting, you can customize the Markdown output format at several strategic extension points without rewriting the entire extraction engine.

## Understanding the Conversion Pipeline

The transformation from PDF to Markdown follows a five-stage pipeline defined in the source code. Understanding these stages helps identify where to inject custom behavior.

1. **Extraction** – The core extractor in [`src/extractor/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/extractor/mod.rs) walks PDF content streams, resolves fonts, and builds a list of `TextLine` objects.
2. **Classification** – [`src/markdown/classify.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/classify.rs) examines each line (analyzing fonts, indentation, and punctuation) and assigns a role such as **Header**, **List**, **Code**, or **Paragraph**.
3. **Pre-processing** – [`src/markdown/preprocess.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/preprocess.rs) performs layout-aware clean-ups including drop-cap merging, heading line merging, and table-spanning line masking.
4. **Conversion** – [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs) iterates over classified lines, emitting Markdown tokens according to detected roles. This module applies default formatting such as `#` for headings and `-` for lists.
5. **Post-processing** – [`src/markdown/postprocess.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/postprocess.rs) applies final polish including dot-leader removal, hyphenation fixes, page-number stripping, and URL re-formatting.

## Customization Approaches

You have three primary strategies to alter the Markdown output, ranging from configuration-based to source-level modifications.

### Using the Library API with ProcessOptions

The most maintainable method involves using the public API exposed in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs). The `process_pdf_with_options` function accepts a `ProcessOptions` struct that controls pipeline behavior without requiring source changes.

```rust
use pdf_inspector::lib::{process_pdf_with_options, ProcessOptions};

let opts = ProcessOptions {
    // Use asterisks instead of dashes for list items
    list_marker: Some("*".into()),
    // Disable drop-cap merging for cleaner poetry extraction
    enable_drop_cap_merge: false,
    ..Default::default()
};

let markdown = process_pdf_with_options("sample.pdf", opts).unwrap();
println!("{}", markdown);

```

This approach allows you to toggle features like `preserve_linebreaks`, `custom_heading_prefix`, or `list_marker` through strongly-typed configuration rather than code modifications.

### Consuming JSON Output for External Rendering

The CLI binary `pdf2md` exposes a `--json` flag that outputs a structured representation instead of raw Markdown. By consuming this JSON, you can implement completely custom rendering logic in any programming language.

First, extract the structured data:

```bash
pdf2md --json report.pdf > report.json

```

Then process the JSON to emit Markdown with your own rules:

```rust
use serde_json::Value;
use std::fs;

let data: Value = serde_json::from_str(&fs::read_to_string("report.json").unwrap()).unwrap();

for page in data["pages"].as_array().unwrap() {
    for line in page["lines"].as_array().unwrap() {
        // Custom heading prefix using Setext-style underline
        if line["role"] == "Header" {
            println!("{}\n===", line["text"]);
        } else {
            println!("{}", line["text"]);
        }
    }
}

```

This method completely bypasses the internal Markdown renderer while preserving all layout and classification metadata.

### Forking and Modifying Source Code

For deep customization of rendering logic, fork the repository and modify the conversion modules directly. The most common target is [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs), specifically the `convert_line` function where role-to-Markdown mapping occurs.

To change heading styles from ATX (`# Heading`) to Setext (`Heading\n===`):

1. Edit [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs) to replace the `#` generation logic with underline syntax for level-1 headings.
2. Rebuild the library: `cargo build --release`
3. Run the modified binary: `cargo run --bin pdf2md sample.pdf`

This approach gives you direct control over every aspect of Markdown generation but requires maintaining a fork.

## Key Extension Points

When modifying source code or designing configuration options, target these specific modules and functions:

### Header Styles

Modify [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs) in the `convert_line` function to change how heading levels map to Markdown syntax. You can replace `#` prefixes with Setext-style underlines or adjust the heading level offset (e.g., mapping PDF "Title" to `##` instead of `#`).

### List Markers

List rendering logic resides in [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs). You can switch from `-` to `*` or `+` for unordered lists, or enforce numbered lists by detecting list roles and prepending sequence numbers instead of bullets.

### Table Formatting

For table-specific adjustments, edit [`src/tables/format.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/format.rs), specifically the `format_table` function. This controls column alignment indicators, pipe spacing, and Markdown-compatible table borders. You can tweak the output to support GitHub-flavored Markdown tables or alternative alignment syntax.

### Pre- and Post-Processing Tweaks

Add custom cleaning logic in [`src/markdown/preprocess.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/preprocess.rs) to handle document-specific layouts before conversion, or in [`src/markdown/postprocess.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/postprocess.rs) to refine output after Markdown generation. Common modifications include customizing hyphenation rules, filtering specific page number patterns, or preserving specific typographic elements.

## Summary

- **Use `ProcessOptions`** via `process_pdf_with_options` in [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs) for configuration-driven customization of list markers, drop-cap handling, and linebreak preservation.
- **Consume `--json` output** from the `pdf2md` CLI to re-render Markdown externally with complete control over syntax and style.
- **Modify [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs)** to change core rendering logic for headings, lists, and code blocks when you need behavior changes beyond configuration options.
- **Adjust [`src/tables/format.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/tables/format.rs)** to customize table alignment and border formatting.
- **Extend pre/post-processors** in [`src/markdown/preprocess.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/preprocess.rs) and [`src/markdown/postprocess.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/postprocess.rs) for document-specific cleaning and formatting polish.

## Frequently Asked Questions

### How do I change the bullet style from dashes to asterisks?

Use the `ProcessOptions` struct when calling `process_pdf_with_options` from [`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs). Set the `list_marker` field to `Some("*".into())` to replace the default `-` with `*` for all unordered lists. Alternatively, modify the list rendering section in [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs) if you need to change the hardcoded default.

### Can I convert PDFs to Markdown without installing Rust?

Yes. Use the `pdf2md` CLI tool with the `--json` flag to extract structured JSON output, then process that JSON using Python, Node.js, or any language that supports JSON parsing. This allows you to generate custom Markdown without compiling or modifying the Rust source code.

### Where is the heading level determined during conversion?

Heading levels are determined in [`src/markdown/classify.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/classify.rs) based on font size and styling, but the actual Markdown syntax (how many `#` characters to prepend) is applied in [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs) within the `convert_line` function. To change heading level mapping or switch to Setext-style headings, modify the logic in [`convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/convert.rs).

### How can I preserve specific page layout elements like drop caps?

Control drop-cap handling through the `enable_drop_cap_merge` option in `ProcessOptions` (set to `false` to preserve drop caps as separate lines). For more granular control, edit [`src/markdown/preprocess.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/preprocess.rs) where the `merge_drop_caps` logic identifies and merges drop-cap characters with subsequent text lines.