How to Use `detect_vector_grid_in_region_mem` for TSR-Compatible Table Structure Detection in pdf-inspector

Use detect_vector_grid_in_region_mem in pdf-inspector to extract line-based vector grids from PDF regions, producing TSR-compatible table structures for downstream semantic labeling and markdown conversion.

The pdf-inspector Rust library implements three table detection strategies—rect-based, line-based, and heuristic—with the TSR (Table Structure Recognition)-compatible workflow built on the line-based stage. The function detect_vector_grid_in_region_mem in src/tables/grid.rs performs the core grid extraction, converting vector lines into a structured Grid that matches TSR expectations. This guide shows how to invoke it directly for precise, memory-efficient table detection.

What detect_vector_grid_in_region_mem Does

This function implements a five-step pipeline for TSR-ready table extraction:

  1. Region selection – Accepts a RegionMem (memory-mapped PDF slice) and isolates the target page area
  2. Vector extraction – Scans for horizontal and vertical line objects, converting them to resolution-independent vectors
  3. Grid construction – Merges intersecting vectors and computes column/row boundaries stored in a Grid struct
  4. Validation – Enforces minimum density constraints (≥2 intersecting lines per row/column)
  5. Output – Returns DetectedTable containing the Grid plus cell-spanning metadata

The *_mem suffix indicates memory-efficient operation on region-level slices, making it suitable for multi-column documents or sub-page extraction.

Prerequisites: Obtaining a RegionMem

Before calling the detector, extract a memory region from your PDF using src/extractor/content_stream.rs:

use pdf_inspector::extractor::content_stream::extract_region;
use pdf_inspector::types::PdfRect;

// Define region in PDF points (1/72 inch)
let region = PdfRect {
    x: 50.0,
    y: 100.0,
    width: 500.0,
    height: 300.0,
};

let region_mem = extract_region(&page, region).unwrap();

Alternative: use extract_all_regions(&page) to obtain candidate regions across an entire page.

Direct Grid Extraction from a Region

Call detect_vector_grid_in_region_mem with a prepared RegionMem:

use pdf_inspector::tables::detect_vector_grid_in_region_mem;

if let Some(detected) = detect_vector_grid_in_region_mem(&region_mem) {
    println!("Columns: {}, Rows: {}",
             detected.grid.columns.len(),
             detected.grid.rows.len());
    
    // Access TSR-compatible grid structure
    let grid = &detected.grid;
    // grid.columns: Vec<f64> (x-positions)
    // grid.rows: Vec<f64> (y-positions)
} else {
    eprintln!("No valid table grid detected");
}

The DetectedTable struct provides:

  • grid: Core column/row boundary vectors
  • spanning_cells: Merged cell metadata for complex tables

Full-Document TSR Pipeline

Iterate all pages and regions for complete table extraction:

use pdf_inspector::{
    extractor::content_stream::extract_all_regions,
    tables::{detect_vector_grid_in_region_mem, DetectedTable},
};

fn extract_tsr_tables(pdf_path: &str) -> Vec<DetectedTable> {
    let doc = pdf_inspector::load_pdf(pdf_path).expect("cannot open PDF");
    let mut tables = Vec::new();

    for page in doc.pages() {
        for region_mem in extract_all_regions(&page) {
            if let Some(tbl) = detect_vector_grid_in_region_mem(&region_mem) {
                tables.push(tbl);
            }
        }
    }
    tables
}

Key integration points:

Exporting Grids for External TSR Models

Serialize the vector grid structure for frameworks like Table-Structure-Recognition:

use serde_json::json;

if let Some(tbl) = detect_vector_grid_in_region_mem(&region_mem) {
    let payload = json!({
        "columns": tbl.grid.columns,
        "rows": tbl.grid.rows,
        "spans": tbl.spanning_cells,
    });
    println!("{}", serde_json::to_string_pretty(&payload).unwrap());
}

This format separates structure (grid geometry) from content (cell text), aligning with two-stage TSR architectures.

Architecture: Where the Function Fits

The detection hierarchy in src/tables/mod.rs orchestrates strategies:

Priority Strategy File When Used
1 Rect-based src/tables/detect_rects.rs Clear rectangular text blocks
2 Line-based src/tables/detect_lines.rs → calls detect_vector_grid_in_region_mem TSR-compatible vector grids
3 Heuristic src/tables/detect_heuristic.rs Fallback to font/spacing analysis

By invoking detect_vector_grid_in_region_mem directly, you bypass earlier stages and force line-based TSR-compatible detection.

Performance Characteristics

  • Memory: Region-mapped; O(1) additional allocation regardless of PDF size
  • Speed: Linear in line object count; single-pass vector extraction
  • Precision: Sub-point accuracy for column/row boundaries
  • Robustness: Fails gracefully to None for invalid grids rather than emitting spurious structure

Summary

  • detect_vector_grid_in_region_mem in src/tables/grid.rs is the core TSR-compatible grid extractor in pdf-inspector
  • Requires a RegionMem from extract_region or extract_all_regions in src/extractor/content_stream.rs
  • Returns DetectedTable with resolution-independent Grid structure suitable for downstream TSR models
  • Skip rect/heuristic detection by calling directly for forced line-based extraction
  • Export via JSON for external TSR pipelines, or convert to markdown via src/markdown/format.rs

Frequently Asked Questions

What does "TSR-compatible" mean in this context?

TSR (Table Structure Recognition) refers to extracting the logical row/column geometry independently of cell content. detect_vector_grid_in_region_mem produces a Grid struct containing precise boundary vectors that match TSR framework expectations—structural metadata without OCR text—enabling clean separation between layout analysis and semantic labeling stages.

How is detect_vector_grid_in_region_mem different from the main process_pdf_with_options API?

The public API in src/lib.rs runs all three detection strategies automatically. Calling detect_vector_grid_in_region_mem directly bypasses rect-based and heuristic detection, forcing pure line-based extraction. This is essential when you need guaranteed TSR-compatible output or when integrating into custom pipelines that handle their own region selection.

What input formats trigger successful grid detection?

The function requires PDFs with explicit vector line objects drawing table borders. Scanned images without embedded line vectors will return None—use OCR-based preprocessing or fall back to heuristic detection in src/tables/detect_heuristic.rs for such cases.

Can I use this with multi-page or multi-column documents?

Yes. The *_mem design processes arbitrary RegionMem slices, making it ideal for complex layouts. Extract regions per-column using extract_all_regions, then invoke detect_vector_grid_in_region_mem on each candidate independently.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →