# How pdf‑inspector's Markdown Converter Handles Bullet, Numbered, and Lettered Lists

> Discover how pdf-inspector's Markdown converter classifies and normalizes bullet, numbered, and lettered lists from PDFs. Learn about its pattern matching in Rust.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-06

---

**The pdf‑inspector Markdown converter detects list items using pattern matching in [`src/markdown/classify.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/classify.rs), then normalizes bullet markers to `-` while preserving numbered and lettered prefixes exactly as they appear in the PDF.**

The `pdf-inspector` library transforms raw PDF text into clean Markdown through a two-phase process: classification followed by formatting. Understanding how this **Markdown converter** processes different list types helps you predict output structure when extracting documents with mixed formatting.

## List Detection in pdf‑inspector

The classification logic resides in [`src/markdown/classify.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/classify.rs). The `is_list_item(text: &str) -> bool` function examines each line to determine if it belongs to a list.

### Bullet List Detection

**Bulleted items** match when a line starts with common bullet symbols followed by a space:

```rust
// Recognized bullet characters: •, -, *, ○, ●, ◦

```

These symbols trigger list classification regardless of surrounding styling markup.

### Numbered List Detection

**Numbered items** require a decimal number followed by `.` or `)`:

```rust
// Examples: 1., 2), 10.
// Implementation examines first 5 characters for digit + delimiter pattern

```

The validation ensures the prefix contains only digits before the delimiter. Source lines 84-106 implement this strict numeric check.

### Lettered List Detection

**Lettered items** match single ASCII letters with the same delimiters:

```rust
// Examples: a., b), (a)
// Source lines 110-118 handle parenthesized and suffixed variants

```

This covers both `a.` style and `(a)` style ordering common in legal documents and academic papers.

## Normalizing Markdown Syntax

Once classified, `format_list_item(text: &str) -> String` in [`src/markdown/classify.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/classify.rs) (source lines 124-152) transforms markers to standard Markdown:

| Original Marker | Output |
|---------------|--------|
| `•`, `○`, `●`, `◦` | `-` (canonical bullet) |
| `-` or `*` | unchanged (preserves author intent) |
| `1.`, `2)`, `a.`, `b)` | preserved exactly |

The function also handles styled bullets like `**● Item**` or `<u>● Item</u>` by stripping the bullet character while preserving surrounding tags (source lines 33-45).

## Integration in the Conversion Pipeline

The [`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs) file orchestrates list processing through this flow:

1. Trim whitespace from each extracted line
2. Call `is_list_item` when list detection is enabled
3. Pass matches to `format_list_item`
4. Track `in_list` state to group subsequent indented lines

This state machine ensures multi-line list items remain properly associated with their markers.

## Practical Example

Consider a PDF with heterogeneous list formatting:

```bash
pdf2md examples/mixed-list.pdf > mixed-list.md

```

The resulting [`mixed-list.md`](https://github.com/firecrawl/pdf-inspector/blob/main/mixed-list.md):

```markdown
- First bullet point
- Second bullet point
1. First numbered item
2) Second numbered item
a. First lettered item
b) Second lettered item

```

All bullet variants collapse to `-`. Numbered and lettered sequences pass through unchanged, relying on Markdown's native ordered list rendering.

## Key Implementation Files

- **[`src/markdown/classify.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/classify.rs)** — Pattern matching for `is_list_item` and `format_list_item`
- **[`src/markdown/convert.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/convert.rs)** — Conversion loop applying classification results
- **[`src/markdown/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/mod.rs)** — Public API `process_pdf_with_options` coordinating extraction

## Summary

- **Bullet lists**: Normalized to `-` regardless of original symbol (`•`, `●`, `○`, `◦`, `*`)
- **Numbered lists**: Preserved with original digits and delimiters (`.` or `)`)
- **Lettered lists**: Maintained as extracted, supporting both suffix and parenthesis styles
- **Styling preservation**: Formatting tags around bullets survive the conversion process

## Frequently Asked Questions

### How does pdf‑inspector distinguish between a dash used as a bullet versus a dash in regular text?

The converter requires a space after the dash and validates against the `is_list_item` pattern set. A standalone `-` followed by space at line start classifies as a bullet; dashes within sentences or without trailing space remain plain text.

### Can pdf‑inspector handle nested lists with different marker styles?

The current implementation tracks a boolean `in_list` flag for basic continuity. Deep nesting with alternating bullet and numbered styles processes sequentially, though indentation preservation depends on the PDF's whitespace structure rather than explicit nesting depth detection.

### Why are numbered list markers not normalized to `1.` format?

Markdown natively supports both `.` and `)` delimiters for ordered lists. Preserving the original marker maintains author intent and handles cases where `)` specifically indicates legal or formal document conventions that `.` would alter semantically.