How the detect_strikethrough Feature Works in OpenDataLoader PDF: Algorithm and Output Format

The detect_strikethrough feature scans PDF pages for horizontal lines that intersect text chunks and wraps matching content with Markdown strikethrough syntax (~~text~~).

The opendataloader-project/opendataloader-pdf library provides optional strikethrough detection that identifies visually crossed-out text during PDF extraction. When enabled, the processor mutates the internal TextChunk values to include standard Markdown delimiters, ensuring downstream serializers preserve the semantic markup. This article explains the detection algorithm implemented in StrikethroughProcessor.java and the exact output format produced.

Enabling the Strikethrough Detection Feature

The feature is disabled by default and controlled by the detectStrikethrough boolean flag in Config.java (lines 85-86). You can activate it through either the Java API or the command-line interface.

Java API activation:

import org.opendataloader.pdf.api.Config;

Config config = new Config();
config.setDetectStrikethrough(true);  // Enable detection

CLI activation:

opendataloader-pdf-cli \
    --input document.pdf \
    --detect-strikethrough \
    --output-format markdown

When the flag is true, the main pipelines conditionally invoke the processor. In DocumentProcessor.java (lines 155-156) and HybridDocumentProcessor.java (lines 282-283), the code checks config.isDetectStrikethrough() before calling StrikethroughProcessor.processStrikethroughs(pageContents).

The Three-Stage Detection Algorithm

The StrikethroughProcessor class implements a geometric analysis pipeline that runs in three distinct stages to minimize false positives while capturing legitimate strikethroughs.

Candidate Collection

The processor first splits page objects into horizontal LineChunk and TextChunk instances. Only horizontal lines—verified via line.isHorizontalLine()—are retained for further analysis (lines 61-74 in StrikethroughProcessor.java). This initial filter eliminates vertical rules and diagonal markings that cannot represent standard strikethrough formatting.

False Positive Filtering

Before geometric matching, the algorithm applies two heuristic filters to discard decorative lines:

  • Table border exclusion: The method isTableBorderLine (lines 11-18) detects and ignores lines that form table grid boundaries.
  • Stroke thickness validation: Lines exceeding the MAX_STROKE_TO_TEXT_HEIGHT_RATIO threshold are rejected (lines 33-37), preventing thick separators or graphic elements from being mistaken for text strikethroughs.

Geometric Validation

The core detection logic resides in isStrikethroughLine (lines 123-173), which performs four geometric checks against each candidate line and text pair:

  1. Vertical center alignment: The line's Y-coordinate must fall within VERTICAL_CENTER_TOLERANCE (20% of the text height) of the TextChunk's vertical center.
  2. Horizontal overlap: The line must overlap the text horizontally by at least MIN_HORIZONTAL_OVERLAP_RATIO (0.8 or 80%).
  3. Width proportionality: The line width cannot exceed MAX_LINE_TO_TEXT_WIDTH_RATIO (1.5x) of the text width, preventing long underlines from triggering false matches.

When all constraints pass, the TextChunk is flagged as struck-through.

Output Format and Markdown Wrapping

Upon detection, the processor mutates the TextChunk value directly by wrapping the raw text with double tilde delimiters. As implemented in lines 98-99 of StrikethroughProcessor.java:

String value = chunk.getValue();
if (!value.startsWith("~~")) {
    chunk.setValue("~~" + value + "~~");
}

The output format is Markdown strikethrough syntax embedded within the text content itself. This design choice ensures that all downstream serializers—whether emitting JSON, plain text, or HTML—preserve the semantic markup:

  • Markdown files: ~~deleted text~~ renders visually as struck-through text.
  • HTML conversion: Processors typically convert the tildes to <del>deleted text</del> tags.
  • JSON output: The string field contains the literal value "~~deleted text~~".

Integration with Document Processors

The strikethrough detection integrates seamlessly into both standard and hybrid processing workflows. The DocumentProcessor class invokes the feature at line 155-156, while the HybridDocumentProcessor includes identical conditional logic at lines 282-283. Both check the configuration flag before processing, ensuring zero overhead when the feature is disabled.

Example extraction result:

If a PDF contains the sentence "This feature is deprecated and will be removed", the extracted output becomes:

This feature is ~~deprecated~~ and will be removed

The delimiters persist through the entire pipeline, allowing rendering engines to display the visual strikethrough while maintaining machine-readable text content.

Summary

  • The detect_strikethrough feature is controlled by a boolean flag in Config.java and defaults to false.
  • Detection requires three validation stages: candidate collection, false-positive filtering (table borders and thick strokes), and geometric alignment checks.
  • The geometric test requires 80% horizontal overlap, vertical center alignment within 20% tolerance, and width proportionality under 1.5x.
  • Valid matches are wrapped with ~~ delimiters directly in the TextChunk value (lines 98-99 of StrikethroughProcessor.java).
  • Output propagates through all serializers as Markdown strikethrough syntax, compatible with HTML <del> tags and JSON string fields.

Frequently Asked Questions

What is the exact output format of the detect_strikethrough feature?

The feature outputs standard Markdown strikethrough syntax. When the processor identifies a struck-through text chunk, it prefixes and suffixes the content with double tilde characters (~~), resulting in strings like ~~crossed out~~. This format is preserved in JSON exports and renders visually in Markdown-compatible viewers.

How do I enable strikethrough detection in my Java application?

Import org.opendataloader.pdf.api.Config, instantiate a Config object, and call setDetectStrikethrough(true) before passing the configuration to your document processor. The processor will automatically invoke StrikethroughProcessor.processStrikethroughs() during page analysis.

What geometric constraints prevent false positives?

The algorithm requires the horizontal line to overlap at least 80% of the text width (MIN_HORIZONTAL_OVERLAP_RATIO), stay within 20% of the text's vertical center (VERTICAL_CENTER_TOLERANCE), and not exceed 1.5 times the text width (MAX_LINE_TO_TEXT_WIDTH_RATIO). Additionally, table borders and strokes thicker than the maximum height ratio are filtered out before geometric testing.

Does the feature detect strikethroughs in tables?

No. The isTableBorderLine method (lines 11-18) explicitly filters out lines identified as table grid borders to prevent table formatting from being mistaken for text strikethroughs. Only non-border horizontal lines intersecting text chunks trigger the detection.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →