# OpenDataLoader PDF `table_method` Default vs Cluster: Implementation Guide

> Explore OpenDataLoader PDF's table_method default vs cluster. Learn when to use default for border-only detection and cluster for scanned docs or complex layouts. An essential implementation guide.

- Repository: [opendataloader-project/opendataloader-pdf](https://github.com/opendataloader-project/opendataloader-pdf)
- Tags: implementation-guide
- Published: 2026-03-20

---

**Use `table_method='default'` for border‑only detection on digitally generated PDFs with complete grid lines, and `table_method='cluster'` for border‑plus‑text clustering when handling scanned documents, missing borders, or complex multi‑section layouts.**

The `table_method` parameter in the **opendataloader‑project/opendataloader‑pdf** repository controls how the engine identifies tabular structures during PDF conversion. Understanding the differences between the `default` and `cluster` strategies is essential for optimizing both accuracy and performance when extracting tables from different document types.

## What Is the `table_method` Parameter?

The `table_method` option selects the table‑detection algorithm applied during the conversion process. It is exposed as:

- **CLI flag:** `--table-method`
- **Java API:** `Config#setTableMethod(String)`
- **Node.js API:** `ConvertOptions.tableMethod`

The allowed values are defined as constants in [`java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/api/Config.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/api/Config.java) at lines 87‑90:

```java
public static final String TABLE_METHOD_DEFAULT = "default";
public static final String TABLE_METHOD_CLUSTER = "cluster";

```

These values are registered as configuration options at lines 110‑111 in the same file.

## The Default Method: Border‑Based Detection

### How It Works

The `default` method performs **border‑only detection**. It scans the PDF’s vector graphics stream for continuous lines, rectangles, and ruling objects that form closed rectangular grids. Cells are grouped when they share common geometric boundaries, and the engine builds row/column structures directly from this geometry.

### Source Code Implementation

According to the TypeScript definitions in [`node/opendataloader-pdf/src/convert-options.generated.ts`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/node/opendataloader-pdf/src/convert-options.generated.ts) (lines 24‑28), the `default` method is documented as "border‑based". The core implementation relies entirely on the vector graphics parser without invoking text clustering algorithms.

### When to Use Default

- **Digitally generated PDFs** exported from spreadsheets or reporting tools with complete grid lines.
- **Performance‑critical** workflows where minimal CPU and memory overhead is required.
- Documents with **clear, unbroken borders** and no scanned content.

## The Cluster Method: Border Plus Text Clustering

### How It Works

The `cluster` method implements **border + text clustering** detection. It first executes the same border analysis as the `default` method, then applies a spatial‑clustering algorithm to the extracted text boxes. This secondary pass infers row and column structures even when borders are missing, broken, or faint. The final result merges the border‑based grid with the clustering‑inferred grid.

### Source Code Implementation

The `cluster` method is defined alongside `default` in [`Config.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/Config.java) and exposed through the same API surface. The clustering logic runs entirely on the Java side, analyzing text box proximity to group words into horizontal lines (rows) and vertical alignments (columns).

### When to Use Cluster

- **Scanned PDFs** with OCR text layers where original borders may be degraded.
- Documents with **missing, weak, or partial borders**.
- **Complex multi‑section tables** or mixed‑layout documents where geometric analysis alone fails.
- Cases requiring higher accuracy: the README notes a TEDS score improvement from approximately **0.49 → 0.93** for hard cases when using advanced detection modes.

## Code Examples

### CLI Usage

```bash

# Fastest option: border-only detection

opendataloader-pdf mydoc.pdf --table-method default

# Better accuracy for scanned or borderless tables

opendataloader-pdf mydoc.pdf --table-method cluster

```

### Java API

```java
import org.opendataloader.pdf.api.Config;

// Default border-based detection
Config defaultConfig = new Config();  // Implicitly uses "default"

// Cluster-based detection for complex layouts
Config clusterConfig = new Config();
clusterConfig.setTableMethod("cluster");

```

### Node.js API

```javascript
import { convert } from "opendataloader-pdf";

// Border-only (default behavior)
await convert("mydoc.pdf");

// Border + clustering for improved structure detection
await convert("mydoc.pdf", { tableMethod: "cluster" });

```

## Performance and Accuracy Comparison

| Characteristic | `default` Method | `cluster` Method |
|----------------|------------------|------------------|
| **Detection Strategy** | Border geometry only | Border geometry + spatial text clustering |
| **CPU Usage** | Minimal | Moderate (additional clustering pass) |
| **Memory Usage** | Low (vector graphics only) | Higher (text box analysis) |
| **Accuracy on Digital PDFs** | High | High (equivalent) |
| **Accuracy on Scanned PDFs** | Low | High (recovers structure from text) |
| **TEDS Score (hard cases)** | ~0.49 | ~0.93 |

## Summary

- The **`default`** `table_method` provides the fastest table detection by analyzing only vector borders, making it ideal for digitally generated PDFs with complete grid lines.
- The **`cluster`** `table_method` adds spatial text clustering to border analysis, significantly improving accuracy on scanned documents, borderless tables, and complex layouts at the cost of slightly higher resource consumption.
- Both methods are defined in [`Config.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/Config.java) (lines 87‑90 and 110‑111) and exposed through CLI, Java, and Node.js APIs.

## Frequently Asked Questions

### What is the default value for `table_method` in OpenDataLoader PDF?

The default value is `"default"`, which performs border‑only table detection. When you instantiate a new `Config` object in Java or call the conversion API without specifying the parameter, the engine automatically uses the border‑based method defined in [`Config.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/Config.java) at line 87.

### Can I switch between table methods without reinstalling the library?

Yes. The `table_method` is a runtime configuration parameter passed via the `--table-method` CLI flag, the `setTableMethod()` Java method, or the `tableMethod` Node.js option. No recompilation or reinstallation is required; simply change the parameter value for each conversion operation.

### Does the `cluster` method work with scanned PDFs?

Yes. The `cluster` method is specifically designed for scanned PDFs and documents with degraded borders. It applies spatial clustering to OCR text boxes to infer table structure even when original ruling lines are missing or faint, delivering significantly higher accuracy than the border‑only approach on scanned content.

### Which method consumes more memory?

The `cluster` method consumes more memory because it loads and analyzes text box coordinates in addition to vector graphics. While the `default` method processes only border geometry with minimal overhead, the clustering algorithm requires extra heap space to perform spatial analysis on the extracted text, with memory usage scaling with document complexity.