How pdf-inspector Merges Fragmented Text Items: A Deep Dive into the Heuristic Table Detector
pdf-inspector merges fragmented text items by grouping glyphs by Y-coordinate, sorting them left-to-right, and iteratively combining adjacent fragments that share similar font sizes, decoration states, and fall within a half-font-size horizontal gap threshold.
PDF rendering engines often output individual glyphs as separate objects, creating hundreds of one-character TextItem instances that break downstream processing. The pdf-inspector Rust crate solves this through a purpose-built merging algorithm in its heuristic table detection pipeline. This article examines exactly how the codebase transforms scattered glyphs into coherent textual runs for accurate column detection and table extraction.
Where the Merging Logic Lives
The core implementation resides in src/tables/detect_heuristic.rs. The public entry point merge_adjacent_items forwards to merge_adjacent_items_preserving at lines 21-23:
// src/tables/detect_heuristic.rs#L21-L23
pub fn merge_adjacent_items(items: &[TextItem]) -> (Vec<TextItem>, Vec<Vec<usize>>) {
merge_adjacent_items_preserving(items, &[])
}
This function returns two structures: merged_items containing consolidated text runs, and index_map preserving traceability to original fragment indices.
The Four-Stage Merging Algorithm
Stage 1: Line Grouping by Y-Coordinate
The algorithm first buckets items by vertical position using a 5 pt tolerance to capture glyphs on the same visual line. This grouping loop appears at lines 33-46:
// src/tables/detect_heuristic.rs#L33-L46
let mut line_groups: Vec<Vec<&TextItem>> = Vec::new();
for item in items {
let y = item.y;
// Find existing group within 5.0 pt tolerance
let found = line_groups.iter_mut().find(|g| {
(g[0].y - y).abs() < 5.0
});
match found {
Some(g) => g.push(item),
None => line_groups.push(vec![item]),
}
}
Stage 2: Horizontal Sorting
Each line group is sorted by X-coordinate to establish left-to-right reading order (lines 48-55):
// src/tables/detect_heuristic.rs#L48-L55
for group in &mut line_groups {
group.sort_by(|a, b| {
a.x.partial_cmp(&b.x).unwrap_or(std::cmp::Ordering::Equal)
});
}
Stage 3: Iterative Merging With Multi-Factor Validation
For each sorted group, the code walks items with nested while loops. Four criteria determine whether adjacent fragments merge:
Font-size similarity — The next item must be within 20% of the first item's size (lines 76-80):
// src/tables/detect_heuristic.rs#L76-L80
let size_ratio = candidate.font_size / first.font_size;
if size_ratio < 0.8 || size_ratio > 1.2 {
// Sizes differ by more than 20% — do not merge
break;
}
Decoration preservation — If underline/strike-out flags differ and either item is marked preserved, the merge halts (lines 82-97). This protects redline edits and intentional formatting boundaries.
Gap threshold — The horizontal gap must not exceed ½ × font-size. Larger gaps indicate column boundaries or intentional spaces (lines 99-104):
// src/tables/detect_heuristic.rs#L99-L104
let gap = candidate.x - (current.x + current.width);
if gap > current.font_size * 0.5 {
// Gap too large — treat as separate token
break;
}
Word-boundary spacing — When gaps exceed 0.08 × font-size, a space character is inserted before appending (lines 11-13 of inner loop), distinguishing intra-word spacing from inter-word separation.
Stage 4: Result Construction
A new TextItem is emitted with concatenated text, combined width, and attributes inherited from the first fragment (lines 20-35). Original indices populate index_map at line 37-38 for downstream traceability.
Practical Usage Example
use pdf_inspector::types::TextItem;
use pdf_inspector::tables::detect_heuristic::merge_adjacent_items;
// Raw glyph extraction typically yields fragmented items
let raw_items: Vec<TextItem> = vec![
TextItem { text: "H".into(), x: 10.0, y: 720.0, width: 5.0, height: 10.0,
font_size: 12.0, ..Default::default() },
TextItem { text: "e".into(), x: 15.0, y: 720.0, width: 5.0, height: 10.0,
font_size: 12.0, ..Default::default() },
TextItem { text: "llo".into(), x: 20.0, y: 720.0, width: 15.0, height: 10.0,
font_size: 12.0, ..Default::default() },
];
// Merge into coherent words
let (merged, index_map) = merge_adjacent_items(&raw_items);
println!("Merged: '{}' (from fragments {:?})", merged[0].text, index_map[0]);
// Output: Merged: 'Hello' (from fragments [0, 1, 2])
Integration With Table Detection
The merging step is mandatory for src/tables/mod.rs, which orchestrates the full pipeline. Without consolidation, column-boundary histograms would operate on isolated glyphs rather than realistic word tokens, destroying detection accuracy. The index_map enables later stages to reference original PDF positions for highlighting or verification.
Key Files and Their Roles
| File | Purpose |
|---|---|
src/types.rs |
Defines TextItem struct with x, y, width, font_size, and decoration flags |
src/tables/detect_heuristic.rs |
Full merging implementation including merge_adjacent_items and merge_adjacent_items_preserving |
src/tables/mod.rs |
Pipeline orchestration that invokes merging before column detection |
Summary
- pdf-inspector performs text item merging in
src/tables/detect_heuristic.rsas a prerequisite for table detection - The algorithm uses 5 pt Y-tolerance for line grouping, 20% font-size variance, ½×font-size gap threshold, and 0.08×font-size word-boundary detection
- Decoration flags and preservation markers prevent improper merging of formatted content
- The function returns both consolidated items and an index map for traceability
- This merging enables accurate column detection and markdown conversion downstream
Frequently Asked Questions
How does pdf-inspector determine which text fragments belong on the same line?
pdf-inspector uses a 5 pt vertical tolerance to group items by Y-coordinate, as implemented in src/tables/detect_heuristic.rs lines 33-46. Fragments whose Y positions differ by less than 5 points are considered part of the same visual line.
What happens if two text fragments have different font sizes?
The merging algorithm permits ±20% size variance between adjacent fragments. If the ratio falls outside 0.8–1.2, the merge breaks and a new text run begins. This prevents inappropriately combining headers with body text or footnotes.
Why does pdf-inspector need to merge text items before table detection?
PDF engines often emit individual glyphs as separate objects. Without merging, column-boundary histograms would analyze hundreds of single-character items rather than meaningful word tokens, making reliable table structure detection impossible according to the src/tables/mod.rs pipeline design.
Can the merging algorithm preserve track changes or redline formatting?
Yes. The merge_adjacent_items_preserving variant accepts a list of indices to preserve. When underline or strike-out flags differ and either item is marked preserved, the merge halts to maintain edit boundaries—see lines 82-97 of src/tables/detect_heuristic.rs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →