How pdf‑inspector's Markdown Converter Handles Bullet, Numbered, and Lettered Lists

The pdf‑inspector Markdown converter detects list items using pattern matching in src/markdown/classify.rs, then normalizes bullet markers to - while preserving numbered and lettered prefixes exactly as they appear in the PDF.

The pdf-inspector library transforms raw PDF text into clean Markdown through a two-phase process: classification followed by formatting. Understanding how this Markdown converter processes different list types helps you predict output structure when extracting documents with mixed formatting.

List Detection in pdf‑inspector

The classification logic resides in src/markdown/classify.rs. The is_list_item(text: &str) -> bool function examines each line to determine if it belongs to a list.

Bullet List Detection

Bulleted items match when a line starts with common bullet symbols followed by a space:

// Recognized bullet characters: •, -, *, ○, ●, ◦

These symbols trigger list classification regardless of surrounding styling markup.

Numbered List Detection

Numbered items require a decimal number followed by . or ):

// Examples: 1., 2), 10.
// Implementation examines first 5 characters for digit + delimiter pattern

The validation ensures the prefix contains only digits before the delimiter. Source lines 84-106 implement this strict numeric check.

Lettered List Detection

Lettered items match single ASCII letters with the same delimiters:

// Examples: a., b), (a)
// Source lines 110-118 handle parenthesized and suffixed variants

This covers both a. style and (a) style ordering common in legal documents and academic papers.

Normalizing Markdown Syntax

Once classified, format_list_item(text: &str) -> String in src/markdown/classify.rs (source lines 124-152) transforms markers to standard Markdown:

Original Marker Output
•, ○, ●, ◦ - (canonical bullet)
- or * unchanged (preserves author intent)
1., 2), a., b) preserved exactly

The function also handles styled bullets like **● Item** or <u>● Item</u> by stripping the bullet character while preserving surrounding tags (source lines 33-45).

Integration in the Conversion Pipeline

The src/markdown/convert.rs file orchestrates list processing through this flow:

  1. Trim whitespace from each extracted line
  2. Call is_list_item when list detection is enabled
  3. Pass matches to format_list_item
  4. Track in_list state to group subsequent indented lines

This state machine ensures multi-line list items remain properly associated with their markers.

Practical Example

Consider a PDF with heterogeneous list formatting:

pdf2md examples/mixed-list.pdf > mixed-list.md

The resulting mixed-list.md:

- First bullet point
- Second bullet point
1. First numbered item
2) Second numbered item
a. First lettered item
b) Second lettered item

All bullet variants collapse to -. Numbered and lettered sequences pass through unchanged, relying on Markdown's native ordered list rendering.

Key Implementation Files

Summary

  • Bullet lists: Normalized to - regardless of original symbol (•, ●, ○, ◦, *)
  • Numbered lists: Preserved with original digits and delimiters (. or ))
  • Lettered lists: Maintained as extracted, supporting both suffix and parenthesis styles
  • Styling preservation: Formatting tags around bullets survive the conversion process

Frequently Asked Questions

How does pdf‑inspector distinguish between a dash used as a bullet versus a dash in regular text?

The converter requires a space after the dash and validates against the is_list_item pattern set. A standalone - followed by space at line start classifies as a bullet; dashes within sentences or without trailing space remain plain text.

Can pdf‑inspector handle nested lists with different marker styles?

The current implementation tracks a boolean in_list flag for basic continuity. Deep nesting with alternating bullet and numbered styles processes sequentially, though indentation preservation depends on the PDF's whitespace structure rather than explicit nesting depth detection.

Why are numbered list markers not normalized to 1. format?

Markdown natively supports both . and ) delimiters for ordered lists. Preserving the original marker maintains author intent and handles cases where ) specifically indicates legal or formal document conventions that . would alter semantically.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →