How pdf‑inspector's Markdown Converter Handles Bullet, Numbered, and Lettered Lists
The pdf‑inspector Markdown converter detects list items using pattern matching in src/markdown/classify.rs, then normalizes bullet markers to - while preserving numbered and lettered prefixes exactly as they appear in the PDF.
The pdf-inspector library transforms raw PDF text into clean Markdown through a two-phase process: classification followed by formatting. Understanding how this Markdown converter processes different list types helps you predict output structure when extracting documents with mixed formatting.
List Detection in pdf‑inspector
The classification logic resides in src/markdown/classify.rs. The is_list_item(text: &str) -> bool function examines each line to determine if it belongs to a list.
Bullet List Detection
Bulleted items match when a line starts with common bullet symbols followed by a space:
// Recognized bullet characters: •, -, *, ○, ●, ◦
These symbols trigger list classification regardless of surrounding styling markup.
Numbered List Detection
Numbered items require a decimal number followed by . or ):
// Examples: 1., 2), 10.
// Implementation examines first 5 characters for digit + delimiter pattern
The validation ensures the prefix contains only digits before the delimiter. Source lines 84-106 implement this strict numeric check.
Lettered List Detection
Lettered items match single ASCII letters with the same delimiters:
// Examples: a., b), (a)
// Source lines 110-118 handle parenthesized and suffixed variants
This covers both a. style and (a) style ordering common in legal documents and academic papers.
Normalizing Markdown Syntax
Once classified, format_list_item(text: &str) -> String in src/markdown/classify.rs (source lines 124-152) transforms markers to standard Markdown:
| Original Marker | Output |
|---|---|
•, ○, ●, ◦ |
- (canonical bullet) |
- or * |
unchanged (preserves author intent) |
1., 2), a., b) |
preserved exactly |
The function also handles styled bullets like **● Item** or <u>● Item</u> by stripping the bullet character while preserving surrounding tags (source lines 33-45).
Integration in the Conversion Pipeline
The src/markdown/convert.rs file orchestrates list processing through this flow:
- Trim whitespace from each extracted line
- Call
is_list_itemwhen list detection is enabled - Pass matches to
format_list_item - Track
in_liststate to group subsequent indented lines
This state machine ensures multi-line list items remain properly associated with their markers.
Practical Example
Consider a PDF with heterogeneous list formatting:
pdf2md examples/mixed-list.pdf > mixed-list.md
The resulting mixed-list.md:
- First bullet point
- Second bullet point
1. First numbered item
2) Second numbered item
a. First lettered item
b) Second lettered item
All bullet variants collapse to -. Numbered and lettered sequences pass through unchanged, relying on Markdown's native ordered list rendering.
Key Implementation Files
src/markdown/classify.rs— Pattern matching foris_list_itemandformat_list_itemsrc/markdown/convert.rs— Conversion loop applying classification resultssrc/markdown/mod.rs— Public APIprocess_pdf_with_optionscoordinating extraction
Summary
- Bullet lists: Normalized to
-regardless of original symbol (•,●,○,◦,*) - Numbered lists: Preserved with original digits and delimiters (
.or)) - Lettered lists: Maintained as extracted, supporting both suffix and parenthesis styles
- Styling preservation: Formatting tags around bullets survive the conversion process
Frequently Asked Questions
How does pdf‑inspector distinguish between a dash used as a bullet versus a dash in regular text?
The converter requires a space after the dash and validates against the is_list_item pattern set. A standalone - followed by space at line start classifies as a bullet; dashes within sentences or without trailing space remain plain text.
Can pdf‑inspector handle nested lists with different marker styles?
The current implementation tracks a boolean in_list flag for basic continuity. Deep nesting with alternating bullet and numbered styles processes sequentially, though indentation preservation depends on the PDF's whitespace structure rather than explicit nesting depth detection.
Why are numbered list markers not normalized to 1. format?
Markdown natively supports both . and ) delimiters for ordered lists. Preserving the original marker maintains author intent and handles cases where ) specifically indicates legal or formal document conventions that . would alter semantically.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →