# How the detect_strikethrough Feature Works in OpenDataLoader PDF: Algorithm and Output Format

> Learn how OpenDataLoader PDF's detect_strikethrough feature finds and marks crossed-out text. Explore the algorithm and its clear Markdown output format.

- Repository: [opendataloader-project/opendataloader-pdf](https://github.com/opendataloader-project/opendataloader-pdf)
- Tags: deep-dive
- Published: 2026-03-20

---

**The `detect_strikethrough` feature scans PDF pages for horizontal lines that intersect text chunks and wraps matching content with Markdown strikethrough syntax (`~~text~~`).**

The `opendataloader-project/opendataloader-pdf` library provides optional strikethrough detection that identifies visually crossed-out text during PDF extraction. When enabled, the processor mutates the internal `TextChunk` values to include standard Markdown delimiters, ensuring downstream serializers preserve the semantic markup. This article explains the detection algorithm implemented in [`StrikethroughProcessor.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/StrikethroughProcessor.java) and the exact output format produced.

## Enabling the Strikethrough Detection Feature

The feature is disabled by default and controlled by the **`detectStrikethrough`** boolean flag in [`Config.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/Config.java) (lines 85-86). You can activate it through either the Java API or the command-line interface.

**Java API activation:**

```java
import org.opendataloader.pdf.api.Config;

Config config = new Config();
config.setDetectStrikethrough(true);  // Enable detection

```

**CLI activation:**

```bash
opendataloader-pdf-cli \
    --input document.pdf \
    --detect-strikethrough \
    --output-format markdown

```

When the flag is `true`, the main pipelines conditionally invoke the processor. In [`DocumentProcessor.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/DocumentProcessor.java) (lines 155-156) and [`HybridDocumentProcessor.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/HybridDocumentProcessor.java) (lines 282-283), the code checks `config.isDetectStrikethrough()` before calling `StrikethroughProcessor.processStrikethroughs(pageContents)`.

## The Three-Stage Detection Algorithm

The `StrikethroughProcessor` class implements a geometric analysis pipeline that runs in three distinct stages to minimize false positives while capturing legitimate strikethroughs.

### Candidate Collection

The processor first splits page objects into horizontal `LineChunk` and `TextChunk` instances. Only horizontal lines—verified via `line.isHorizontalLine()`—are retained for further analysis (lines 61-74 in [`StrikethroughProcessor.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/StrikethroughProcessor.java)). This initial filter eliminates vertical rules and diagonal markings that cannot represent standard strikethrough formatting.

### False Positive Filtering

Before geometric matching, the algorithm applies two heuristic filters to discard decorative lines:

- **Table border exclusion**: The method `isTableBorderLine` (lines 11-18) detects and ignores lines that form table grid boundaries.
- **Stroke thickness validation**: Lines exceeding the `MAX_STROKE_TO_TEXT_HEIGHT_RATIO` threshold are rejected (lines 33-37), preventing thick separators or graphic elements from being mistaken for text strikethroughs.

### Geometric Validation

The core detection logic resides in `isStrikethroughLine` (lines 123-173), which performs four geometric checks against each candidate line and text pair:

1. **Vertical center alignment**: The line's Y-coordinate must fall within `VERTICAL_CENTER_TOLERANCE` (20% of the text height) of the `TextChunk`'s vertical center.
2. **Horizontal overlap**: The line must overlap the text horizontally by at least `MIN_HORIZONTAL_OVERLAP_RATIO` (0.8 or 80%).
3. **Width proportionality**: The line width cannot exceed `MAX_LINE_TO_TEXT_WIDTH_RATIO` (1.5x) of the text width, preventing long underlines from triggering false matches.

When all constraints pass, the `TextChunk` is flagged as struck-through.

## Output Format and Markdown Wrapping

Upon detection, the processor mutates the `TextChunk` value directly by wrapping the raw text with double tilde delimiters. As implemented in lines 98-99 of [`StrikethroughProcessor.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/StrikethroughProcessor.java):

```java
String value = chunk.getValue();
if (!value.startsWith("~~")) {
    chunk.setValue("~~" + value + "~~");
}

```

**The output format is Markdown strikethrough syntax** embedded within the text content itself. This design choice ensures that all downstream serializers—whether emitting JSON, plain text, or HTML—preserve the semantic markup:

- **Markdown files**: `~~deleted text~~` renders visually as struck-through text.
- **HTML conversion**: Processors typically convert the tildes to `<del>deleted text</del>` tags.
- **JSON output**: The string field contains the literal value `"~~deleted text~~"`.

## Integration with Document Processors

The strikethrough detection integrates seamlessly into both standard and hybrid processing workflows. The `DocumentProcessor` class invokes the feature at line 155-156, while the `HybridDocumentProcessor` includes identical conditional logic at lines 282-283. Both check the configuration flag before processing, ensuring zero overhead when the feature is disabled.

**Example extraction result:**

If a PDF contains the sentence *"This feature is ~~deprecated~~ and will be removed"*, the extracted output becomes:

```markdown
This feature is ~~deprecated~~ and will be removed

```

The delimiters persist through the entire pipeline, allowing rendering engines to display the visual strikethrough while maintaining machine-readable text content.

## Summary

- The `detect_strikethrough` feature is controlled by a boolean flag in [`Config.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/Config.java) and defaults to `false`.
- Detection requires three validation stages: candidate collection, false-positive filtering (table borders and thick strokes), and geometric alignment checks.
- The geometric test requires 80% horizontal overlap, vertical center alignment within 20% tolerance, and width proportionality under 1.5x.
- Valid matches are wrapped with `~~` delimiters directly in the `TextChunk` value (lines 98-99 of [`StrikethroughProcessor.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/StrikethroughProcessor.java)).
- Output propagates through all serializers as Markdown strikethrough syntax, compatible with HTML `<del>` tags and JSON string fields.

## Frequently Asked Questions

### What is the exact output format of the detect_strikethrough feature?

The feature outputs standard Markdown strikethrough syntax. When the processor identifies a struck-through text chunk, it prefixes and suffixes the content with double tilde characters (`~~`), resulting in strings like `~~crossed out~~`. This format is preserved in JSON exports and renders visually in Markdown-compatible viewers.

### How do I enable strikethrough detection in my Java application?

Import `org.opendataloader.pdf.api.Config`, instantiate a `Config` object, and call `setDetectStrikethrough(true)` before passing the configuration to your document processor. The processor will automatically invoke `StrikethroughProcessor.processStrikethroughs()` during page analysis.

### What geometric constraints prevent false positives?

The algorithm requires the horizontal line to overlap at least 80% of the text width (`MIN_HORIZONTAL_OVERLAP_RATIO`), stay within 20% of the text's vertical center (`VERTICAL_CENTER_TOLERANCE`), and not exceed 1.5 times the text width (`MAX_LINE_TO_TEXT_WIDTH_RATIO`). Additionally, table borders and strokes thicker than the maximum height ratio are filtered out before geometric testing.

### Does the feature detect strikethroughs in tables?

No. The `isTableBorderLine` method (lines 11-18) explicitly filters out lines identified as table grid borders to prevent table formatting from being mistaken for text strikethroughs. Only non-border horizontal lines intersecting text chunks trigger the detection.