# How to Get Layout Complexity Metadata (Tables & Columns per Page) from pdf‑inspector

> Easily get layout complexity metadata, including tables and columns per page, from pdf-inspector. Learn how to use the LayoutComplexity struct with process_pdf() and other functions.

- Repository: [Firecrawl/pdf-inspector](https://github.com/firecrawl/pdf-inspector)
- Tags: how-to-guide
- Published: 2026-08-10

---

**pdf‑inspector computes layout complexity metadata—including table locations and multi‑column page detection—through the `LayoutComplexity` struct returned by `process_pdf()`, `detect_pdf()`, or `extract_pages_markdown()`.**

The firecrawl/pdf‑inspector Rust library analyzes PDF structure to identify complex layouts without relying on external OCR services. This guide explains how to retrieve **tables per page** and **columns per page** metadata using the library's public API.

## Understanding Layout Complexity Results

pdf‑inspector exposes layout analysis through two result types:

- **`PdfProcessResult`** – returned by `process_pdf()` and `detect_pdf()`; contains a `layout: LayoutComplexity` field with global document analysis
- **`PagesExtractionResult`** – returned by `extract_pages_markdown()` and `extract_pages_markdown_mem()`; exposes the same data as `pages_with_tables` and `pages_with_columns` fields alongside per‑page markdown

Both provide identical underlying data: 1‑indexed page numbers where tables or columns appear, plus an `is_complex` boolean flag.

## The LayoutComplexity Data Structure

The metadata definition lives in **[`src/types.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/types.rs)**:

```rust
// src/types.rs lines 62-74
pub struct LayoutComplexity {
    pub is_complex: bool,
    pub pages_with_tables: Vec<u32>,
    pub pages_with_columns: Vec<u32>,
}

```

This struct stores three critical pieces of information:
- `is_complex` – true if any page contains tables or multiple columns
- `pages_with_tables` – vector of page numbers (1‑indexed) with detected tables
- `pages_with_columns` – vector of page numbers (1‑indexed) with multi‑column layout

## How Layout Complexity Is Computed

The core computation happens in **`compute_layout_complexity_with_chart_regions`** in **[`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)** (lines 58‑112). This function:

1. Iterates every distinct page in the document
2. Groups PDF content items per page
3. Runs **three table detectors** per page: rectangle‑based, line‑based, and heuristic detection
4. Applies column histogram analysis via `markdown::split_side_by_side` to detect multi‑column layouts

For tables, the algorithm aggregates results from all three detectors in **`src/tables/`**:
- Rectangle detector – finds table‑like bounding boxes
- Line detector – identifies ruling lines forming cell structures
- Heuristic detector – catches remaining table patterns by content alignment

For columns, the system analyzes text flow patterns and records pages where more than one text column is detected.

## Method 1: Full Processing with process_pdf()

Use **`process_pdf()`** for detection plus full markdown extraction. The function resides at **[`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)** lines 65‑71:

```rust
use pdf_inspector::{process_pdf, LayoutComplexity};

fn main() -> Result<(), pdf_inspector::PdfError> {
    let result = process_pdf("sample.pdf")?;

    // Destructure the layout field
    let LayoutComplexity {
        is_complex,
        pages_with_tables,
        pages_with_columns,
    } = result.layout;

    println!("Complex document: {}", is_complex);
    println!("Tables on pages: {:?}", pages_with_tables);
    println!("Multi-column pages: {:?}", pages_with_columns);
    
    Ok(())
}

```

This runs the complete pipeline: PDF loading, table detection, column analysis, and markdown generation.

## Method 2: Fast Metadata-Only Detection with detect_pdf()

When you need layout complexity metadata **without** markdown extraction, use **`detect_pdf()`** at **[`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)** lines 73‑78:

```rust
use pdf_inspector::detect_pdf;

fn main() -> Result<(), pdf_inspector::PdfError> {
    let info = detect_pdf("sample.pdf")?;
    
    println!("Tables found on: {:?}", info.layout.pages_with_tables);
    println!("Columns found on: {:?}", info.layout.pages_with_columns);
    
    Ok(())
}

```

This is a thin wrapper around `process_pdf_with_options()` using `DetectOnly` mode—significantly faster for large documents when you only need structural metadata.

## Method 3: Per-Page Extraction with extract_pages_markdown()

For page‑level markdown with embedded layout data, use **`extract_pages_markdown()`** at **[`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)** lines 46‑56:

```rust
use pdf_inspector::extract_pages_markdown;

fn main() -> Result<(), pdf_inspector::PdfError> {
    let pages = extract_pages_markdown("sample.pdf", None)?;
    
    // Direct access to layout fields on the result
    println!("Table pages: {:?}", pages.pages_with_tables);
    println!("Column pages: {:?}", pages.pages_with_columns);
    
    // Also available via the layout field
    assert_eq!(pages.pages_with_tables, pages.layout.pages_with_tables);
    
    Ok(())
}

```

The `PagesExtractionResult` struct mirrors `LayoutComplexity` fields directly for convenience.

## CLI Usage for Quick Inspection

pdf‑inspector's CLI tool **`pdf2md`** provides layout complexity metadata in JSON output:

```bash

# Fast detection only - prints layout metadata

pdf2md --detect-only sample.pdf

# Full extraction with JSON output containing layout block

pdf2md --json sample.pdf

```

The `--json` flag serializes the `LayoutComplexity` struct as a nested `layout` object.

## Key Source Files Reference

| File | Purpose |
|------|---------|
| **[`src/types.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/types.rs)** (lines 62‑74) | `LayoutComplexity` struct definition |
| **[`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)** (lines 58‑112) | Core `compute_layout_complexity_with_chart_regions` function; public API entry points |
| **[`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)** (lines 65‑71) | `process_pdf()` implementation |
| **[`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)** (lines 73‑78) | `detect_pdf()` implementation |
| **`src/tables/`** | Three table detector implementations (rect, line, heuristic) |
| **[`src/markdown/mod.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/markdown/mod.rs)** | `split_side_by_side` column detection logic |

## Practical Applications

Layout complexity metadata enables conditional processing pipelines:

- **Selective OCR** – run expensive OCR only on `pages_with_tables` or complex layouts
- **Layout-specific rendering** – apply different CSS or formatting for multi-column pages
- **Content classification** – flag documents requiring special handling based on `is_complex`
- **Performance optimization** – skip table detection heuristics on known simple documents

## Summary

- **`process_pdf()`** returns `LayoutComplexity` in `result.layout` for full processing
- **`detect_pdf()`** provides the same metadata faster by skipping markdown generation
- **`extract_pages_markdown()`** exposes `pages_with_tables` and `pages_with_columns` directly on the result
- All methods return **1‑indexed page numbers** where tables or columns are detected
- The underlying detection runs **three table detectors per page** plus column histogram analysis

## Frequently Asked Questions

### What page numbering does pdf‑inspector use for layout metadata?

pdf‑inspector uses **1‑indexed page numbering** for all `pages_with_tables` and `pages_with_columns` vectors. Page 1 in the output corresponds to the first page of the PDF document, matching human‑readable page numbers rather than zero‑based array indices.

### Can I get layout complexity without extracting markdown text?

Yes. Use **`detect_pdf("file.pdf")`** instead of `process_pdf()`. This runs the same table and column detection but skips markdown generation, returning only the `PdfProcessResult` with populated `layout` field. According to the source code, this is implemented as a convenience wrapper at **[`src/lib.rs`](https://github.com/firecrawl/pdf-inspector/blob/main/src/lib.rs)** lines 73‑78.

### How accurate is the table detection in pdf‑inspector?

The library runs **three complementary detectors** per page: rectangle‑based (bounding boxes), line‑based (ruling lines), and heuristic (content alignment). These are implemented in **`src/tables/`** and aggregated in `compute_layout_complexity_with_chart_regions`. No single method catches all table types, so the combined approach reduces false negatives at the cost of occasional false positives.

### Does layout complexity detection work on scanned PDFs?

pdf‑inspector operates on **PDF content streams**—text elements, vector graphics, and metadata—not pixel data. Scanned documents that are image‑only PDFs without embedded text will show minimal or no detected structure. For such documents, OCR preprocessing (outside pdf‑inspector) is required before meaningful layout analysis can occur.