How to Use `detect_vector_grid_in_region_mem` for TSR-Compatible Table Structure Detection in pdf-inspector
Use detect_vector_grid_in_region_mem in pdf-inspector to extract line-based vector grids from PDF regions, producing TSR-compatible table structures for downstream semantic labeling and markdown conversion.
The pdf-inspector Rust library implements three table detection strategies—rect-based, line-based, and heuristic—with the TSR (Table Structure Recognition)-compatible workflow built on the line-based stage. The function detect_vector_grid_in_region_mem in src/tables/grid.rs performs the core grid extraction, converting vector lines into a structured Grid that matches TSR expectations. This guide shows how to invoke it directly for precise, memory-efficient table detection.
What detect_vector_grid_in_region_mem Does
This function implements a five-step pipeline for TSR-ready table extraction:
- Region selection – Accepts a
RegionMem(memory-mapped PDF slice) and isolates the target page area - Vector extraction – Scans for horizontal and vertical line objects, converting them to resolution-independent vectors
- Grid construction – Merges intersecting vectors and computes column/row boundaries stored in a
Gridstruct - Validation – Enforces minimum density constraints (≥2 intersecting lines per row/column)
- Output – Returns
DetectedTablecontaining theGridplus cell-spanning metadata
The *_mem suffix indicates memory-efficient operation on region-level slices, making it suitable for multi-column documents or sub-page extraction.
Prerequisites: Obtaining a RegionMem
Before calling the detector, extract a memory region from your PDF using src/extractor/content_stream.rs:
use pdf_inspector::extractor::content_stream::extract_region;
use pdf_inspector::types::PdfRect;
// Define region in PDF points (1/72 inch)
let region = PdfRect {
x: 50.0,
y: 100.0,
width: 500.0,
height: 300.0,
};
let region_mem = extract_region(&page, region).unwrap();
Alternative: use extract_all_regions(&page) to obtain candidate regions across an entire page.
Direct Grid Extraction from a Region
Call detect_vector_grid_in_region_mem with a prepared RegionMem:
use pdf_inspector::tables::detect_vector_grid_in_region_mem;
if let Some(detected) = detect_vector_grid_in_region_mem(®ion_mem) {
println!("Columns: {}, Rows: {}",
detected.grid.columns.len(),
detected.grid.rows.len());
// Access TSR-compatible grid structure
let grid = &detected.grid;
// grid.columns: Vec<f64> (x-positions)
// grid.rows: Vec<f64> (y-positions)
} else {
eprintln!("No valid table grid detected");
}
The DetectedTable struct provides:
grid: Core column/row boundary vectorsspanning_cells: Merged cell metadata for complex tables
Full-Document TSR Pipeline
Iterate all pages and regions for complete table extraction:
use pdf_inspector::{
extractor::content_stream::extract_all_regions,
tables::{detect_vector_grid_in_region_mem, DetectedTable},
};
fn extract_tsr_tables(pdf_path: &str) -> Vec<DetectedTable> {
let doc = pdf_inspector::load_pdf(pdf_path).expect("cannot open PDF");
let mut tables = Vec::new();
for page in doc.pages() {
for region_mem in extract_all_regions(&page) {
if let Some(tbl) = detect_vector_grid_in_region_mem(®ion_mem) {
tables.push(tbl);
}
}
}
tables
}
Key integration points:
- Combine with
src/extractor/layout.rsto enrich cells with OCR-derived content - Feed to
src/markdown/format.rsfor TSR-compatible markdown output - Use
--jsonflag inpdf2mdCLI for downstream model consumption
Exporting Grids for External TSR Models
Serialize the vector grid structure for frameworks like Table-Structure-Recognition:
use serde_json::json;
if let Some(tbl) = detect_vector_grid_in_region_mem(®ion_mem) {
let payload = json!({
"columns": tbl.grid.columns,
"rows": tbl.grid.rows,
"spans": tbl.spanning_cells,
});
println!("{}", serde_json::to_string_pretty(&payload).unwrap());
}
This format separates structure (grid geometry) from content (cell text), aligning with two-stage TSR architectures.
Architecture: Where the Function Fits
The detection hierarchy in src/tables/mod.rs orchestrates strategies:
| Priority | Strategy | File | When Used |
|---|---|---|---|
| 1 | Rect-based | src/tables/detect_rects.rs |
Clear rectangular text blocks |
| 2 | Line-based | src/tables/detect_lines.rs → calls detect_vector_grid_in_region_mem |
TSR-compatible vector grids |
| 3 | Heuristic | src/tables/detect_heuristic.rs |
Fallback to font/spacing analysis |
By invoking detect_vector_grid_in_region_mem directly, you bypass earlier stages and force line-based TSR-compatible detection.
Performance Characteristics
- Memory: Region-mapped; O(1) additional allocation regardless of PDF size
- Speed: Linear in line object count; single-pass vector extraction
- Precision: Sub-point accuracy for column/row boundaries
- Robustness: Fails gracefully to
Nonefor invalid grids rather than emitting spurious structure
Summary
detect_vector_grid_in_region_meminsrc/tables/grid.rsis the core TSR-compatible grid extractor inpdf-inspector- Requires a
RegionMemfromextract_regionorextract_all_regionsinsrc/extractor/content_stream.rs - Returns
DetectedTablewith resolution-independentGridstructure suitable for downstream TSR models - Skip rect/heuristic detection by calling directly for forced line-based extraction
- Export via JSON for external TSR pipelines, or convert to markdown via
src/markdown/format.rs
Frequently Asked Questions
What does "TSR-compatible" mean in this context?
TSR (Table Structure Recognition) refers to extracting the logical row/column geometry independently of cell content. detect_vector_grid_in_region_mem produces a Grid struct containing precise boundary vectors that match TSR framework expectations—structural metadata without OCR text—enabling clean separation between layout analysis and semantic labeling stages.
How is detect_vector_grid_in_region_mem different from the main process_pdf_with_options API?
The public API in src/lib.rs runs all three detection strategies automatically. Calling detect_vector_grid_in_region_mem directly bypasses rect-based and heuristic detection, forcing pure line-based extraction. This is essential when you need guaranteed TSR-compatible output or when integrating into custom pipelines that handle their own region selection.
What input formats trigger successful grid detection?
The function requires PDFs with explicit vector line objects drawing table borders. Scanned images without embedded line vectors will return None—use OCR-based preprocessing or fall back to heuristic detection in src/tables/detect_heuristic.rs for such cases.
Can I use this with multi-page or multi-column documents?
Yes. The *_mem design processes arbitrary RegionMem slices, making it ideal for complex layouts. Extract regions per-column using extract_all_regions, then invoke detect_vector_grid_in_region_mem on each candidate independently.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →