OpenDataLoader PDF `table_method` Default vs Cluster: Implementation Guide

Use table_method='default' for border‑only detection on digitally generated PDFs with complete grid lines, and table_method='cluster' for border‑plus‑text clustering when handling scanned documents, missing borders, or complex multi‑section layouts.

The table_method parameter in the opendataloader‑project/opendataloader‑pdf repository controls how the engine identifies tabular structures during PDF conversion. Understanding the differences between the default and cluster strategies is essential for optimizing both accuracy and performance when extracting tables from different document types.

What Is the table_method Parameter?

The table_method option selects the table‑detection algorithm applied during the conversion process. It is exposed as:

  • CLI flag: --table-method
  • Java API: Config#setTableMethod(String)
  • Node.js API: ConvertOptions.tableMethod

The allowed values are defined as constants in java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/api/Config.java at lines 87‑90:

public static final String TABLE_METHOD_DEFAULT = "default";
public static final String TABLE_METHOD_CLUSTER = "cluster";

These values are registered as configuration options at lines 110‑111 in the same file.

The Default Method: Border‑Based Detection

How It Works

The default method performs border‑only detection. It scans the PDF’s vector graphics stream for continuous lines, rectangles, and ruling objects that form closed rectangular grids. Cells are grouped when they share common geometric boundaries, and the engine builds row/column structures directly from this geometry.

Source Code Implementation

According to the TypeScript definitions in node/opendataloader-pdf/src/convert-options.generated.ts (lines 24‑28), the default method is documented as "border‑based". The core implementation relies entirely on the vector graphics parser without invoking text clustering algorithms.

When to Use Default

  • Digitally generated PDFs exported from spreadsheets or reporting tools with complete grid lines.
  • Performance‑critical workflows where minimal CPU and memory overhead is required.
  • Documents with clear, unbroken borders and no scanned content.

The Cluster Method: Border Plus Text Clustering

How It Works

The cluster method implements border + text clustering detection. It first executes the same border analysis as the default method, then applies a spatial‑clustering algorithm to the extracted text boxes. This secondary pass infers row and column structures even when borders are missing, broken, or faint. The final result merges the border‑based grid with the clustering‑inferred grid.

Source Code Implementation

The cluster method is defined alongside default in Config.java and exposed through the same API surface. The clustering logic runs entirely on the Java side, analyzing text box proximity to group words into horizontal lines (rows) and vertical alignments (columns).

When to Use Cluster

  • Scanned PDFs with OCR text layers where original borders may be degraded.
  • Documents with missing, weak, or partial borders.
  • Complex multi‑section tables or mixed‑layout documents where geometric analysis alone fails.
  • Cases requiring higher accuracy: the README notes a TEDS score improvement from approximately 0.49 → 0.93 for hard cases when using advanced detection modes.

Code Examples

CLI Usage


# Fastest option: border-only detection

opendataloader-pdf mydoc.pdf --table-method default

# Better accuracy for scanned or borderless tables

opendataloader-pdf mydoc.pdf --table-method cluster

Java API

import org.opendataloader.pdf.api.Config;

// Default border-based detection
Config defaultConfig = new Config();  // Implicitly uses "default"

// Cluster-based detection for complex layouts
Config clusterConfig = new Config();
clusterConfig.setTableMethod("cluster");

Node.js API

import { convert } from "opendataloader-pdf";

// Border-only (default behavior)
await convert("mydoc.pdf");

// Border + clustering for improved structure detection
await convert("mydoc.pdf", { tableMethod: "cluster" });

Performance and Accuracy Comparison

Characteristic default Method cluster Method
Detection Strategy Border geometry only Border geometry + spatial text clustering
CPU Usage Minimal Moderate (additional clustering pass)
Memory Usage Low (vector graphics only) Higher (text box analysis)
Accuracy on Digital PDFs High High (equivalent)
Accuracy on Scanned PDFs Low High (recovers structure from text)
TEDS Score (hard cases) ~0.49 ~0.93

Summary

  • The default table_method provides the fastest table detection by analyzing only vector borders, making it ideal for digitally generated PDFs with complete grid lines.
  • The cluster table_method adds spatial text clustering to border analysis, significantly improving accuracy on scanned documents, borderless tables, and complex layouts at the cost of slightly higher resource consumption.
  • Both methods are defined in Config.java (lines 87‑90 and 110‑111) and exposed through CLI, Java, and Node.js APIs.

Frequently Asked Questions

What is the default value for table_method in OpenDataLoader PDF?

The default value is "default", which performs border‑only table detection. When you instantiate a new Config object in Java or call the conversion API without specifying the parameter, the engine automatically uses the border‑based method defined in Config.java at line 87.

Can I switch between table methods without reinstalling the library?

Yes. The table_method is a runtime configuration parameter passed via the --table-method CLI flag, the setTableMethod() Java method, or the tableMethod Node.js option. No recompilation or reinstallation is required; simply change the parameter value for each conversion operation.

Does the cluster method work with scanned PDFs?

Yes. The cluster method is specifically designed for scanned PDFs and documents with degraded borders. It applies spatial clustering to OCR text boxes to infer table structure even when original ruling lines are missing or faint, delivering significantly higher accuracy than the border‑only approach on scanned content.

Which method consumes more memory?

The cluster method consumes more memory because it loads and analyzes text box coordinates in addition to vector graphics. While the default method processes only border geometry with minimal overhead, the clustering algorithm requires extra heap space to perform spatial analysis on the extracted text, with memory usage scaling with document complexity.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →