How AI Safety Filters Like 'hidden-text' and 'off-page' Are Implemented in opendataloader-pdf

AI safety filters in opendataloader-pdf automatically remove invisible or out-of-view content by analyzing contrast ratios for hidden text and bounding box geometry for off-page objects before the content reaches downstream LLMs.

The opendataloader-pdf repository implements these AI safety filters to protect large language model pipelines from ingesting misleading or malicious content embedded in PDFs. By default, the system scrubs transparent glyphs, off-canvas objects, and other invisible elements that human readers never see.

Architecture of AI Safety Filters

The filter system consists of three coordinated layers that handle configuration, processing, and CLI integration.

Configuration Layer

The FilterConfig class in java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/api/FilterConfig.java holds the default enablement state for each AI safety filter. By default, filterHiddenText, filterOutOfPage, filterTinyText, and filterHiddenOCG are all set to true.

Processing Engine

The ContentFilterProcessor in java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/processors/ContentFilterProcessor.java executes the geometric and visual analysis required by the off-page and hidden-text filters.

CLI Bridge

The CLIOptions class in java/opendataloader-pdf-cli/src/main/java/org/opendataloader/pdf/cli/CLIOptions.java parses the --content-safety-off argument and toggles the corresponding booleans in FilterConfig.

Hidden-Text Filter Implementation

The hidden-text filter identifies glyphs that are invisible or nearly invisible to human readers by measuring color contrast.

Contrast Detection Logic

Each TextChunk extracted by the VeraPDF WCAG engine is passed to a ContrastRatioConsumer. If the computed contrast ratio falls below 1.2 (defined as MIN_CONTRAST_RATIO in the source), the chunk is flagged as hidden text.

Filtering Action

When filterHiddenText is enabled, the processor drops the chunk entirely. When disabled, the chunk is retained but marked with textChunk.setHiddenText(true) so downstream components can identify it.

The implementation resides in java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/processors/HiddenTextProcessor.java:

private static final double MIN_CONTRAST_RATIO = 1.2d;

for (IObject content : contents) {
    if (content instanceof TextChunk) {
        TextChunk textChunk = (TextChunk) content;
        contrastRatioConsumer.calculateContrastRatio(textChunk);
        if (textChunk.getContrastRatio() < MIN_CONTRAST_RATIO) {
            if (!isFilterHiddenText) {
                textChunk.setHiddenText(true);   // keep but flag
            } else {
                continue;                       // drop from output
            }
        }
    }
    result.add(content);
}

Off-Page Filter Implementation

The off-page filter removes objects that lie outside the visible page area, such as content hidden in the margins or placed on non-visible layers.

Bounding Box Calculation

The filter retrieves the page bounding box via DocumentProcessor.getPageBoundingBox(pageNumber), which returns the MediaBox or CropBox rectangle. The coordinates are normalized by translating the lower-left corner to the origin using pageBoundingBox.move(-leftX, -bottomY).

Geometric Overlap Test

Every IObject on the page is tested against the normalized page box. If pageBoundingBox.notOverlaps(object.getBoundingBox()) returns true, the object is replaced with null. After processing, DocumentProcessor.removeNullObjectsFromList scrubs the placeholders.

The logic is implemented in java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/processors/ContentFilterProcessor.java:

private static void filterOutOfPageContents(int pageNumber, List<IObject> contents) {
    BoundingBox pageBoundingBox = DocumentProcessor.getPageBoundingBox(pageNumber);
    if (pageBoundingBox == null) {
        return;
    }
    pageBoundingBox.move(-pageBoundingBox.getLeftX(), -pageBoundingBox.getBottomY());
    for (int index = 0; index < contents.size(); index++) {
        IObject object = contents.get(index);
        if (object != null && pageBoundingBox.notOverlaps(object.getBoundingBox())) {
            contents.set(index, null);           // mark for removal
        }
    }
}

End-to-End Processing Pipeline

The ContentFilterProcessor.getFilteredContents method orchestrates the entire AI safety pipeline. It conditionally invokes the off-page filter, then delegates to HiddenTextProcessor.findHiddenText for contrast analysis, and finally returns the sanitized content list.

public static List<IObject> getFilteredContents(..., Config config) {
    // ... other clean-up steps ...
    if (config.getFilterConfig().isFilterOutOfPage()) {
        filterOutOfPageContents(pageNumber, pageContents);
    }
    // ... more processing ...
    pageContents = HiddenTextProcessor.findHiddenText(
        inputPdfName, pageContents,
        config.getFilterConfig().isFilterHiddenText(),
        config.getPassword());
    // ... final sanitisation steps ...
    return pageContents;
}

Disabling AI Safety Filters via CLI

Users can selectively disable filters using the --content-safety-off argument. The CLIOptions.applyContentSafetyOption method parses comma-separated values (hidden-text, off-page, tiny, hidden-ocg, all) and sets the corresponding FilterConfig booleans to false.

case "hidden-text":
    config.getFilterConfig().setFilterHiddenText(false);
    break;
case "off-page":
    config.getFilterConfig().setFilterOutOfPage(false);
    break;

For example, to process a PDF while keeping hidden text but removing off-page objects:

opendataloader-pdf mydoc.pdf --content-safety-off hidden-text

Summary

  • AI safety filters in opendataloader-pdf automatically strip invisible content before it reaches LLM pipelines.
  • Hidden-text filter detects glyphs with contrast ratios below 1.2 using HiddenTextProcessor and removes them when enabled.
  • Off-page filter uses geometric bounding box tests in ContentFilterProcessor to discard objects outside the visible page area.
  • Configuration is managed by FilterConfig with safe defaults (all filters enabled) and exposed via the --content-safety-off CLI flag.
  • Integration happens in ContentFilterProcessor.getFilteredContents, which orchestrates the entire sanitization pipeline.

Frequently Asked Questions

What triggers the hidden-text filter in opendataloader-pdf?

The hidden-text filter triggers when a text chunk's contrast ratio falls below 1.2, as defined by the MIN_CONTRAST_RATIO constant in HiddenTextProcessor. The processor uses ContrastRatioConsumer to calculate the ratio between glyph color and background; values below the threshold indicate transparent or invisible text.

How does the off-page filter handle coordinate transformations?

The off-page filter normalizes page coordinates by translating the bounding box so the lower-left corner becomes the origin. It calls pageBoundingBox.move(-leftX, -bottomY) before testing object overlap. This ensures consistent geometric comparison regardless of where the PDF places its coordinate system origin.

Can I disable AI safety filters programmatically?

Yes. Instantiate a Config object, retrieve its FilterConfig via getFilterConfig(), and call setters like setFilterHiddenText(false) or setFilterOutOfPage(false). Pass the modified config to ContentFilterProcessor.getFilteredContents() to run the pipeline with selective filtering disabled.

Where are the default filter settings defined?

Default settings are defined in java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/api/FilterConfig.java. The class initializes filterHiddenText, filterOutOfPage, filterTinyText, and filterHiddenOCG to true, ensuring all AI safety filters run unless explicitly disabled via CLI or API.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →