# How AI Safety Filters Like 'hidden-text' and 'off-page' Are Implemented in opendataloader-pdf

> Discover how opendataloader-pdf implements AI safety filters like hidden text and off-page removal. Learn to auto-remove invisible content before it reaches LLMs.

- Repository: [opendataloader-project/opendataloader-pdf](https://github.com/opendataloader-project/opendataloader-pdf)
- Tags: internals
- Published: 2026-03-20

---

**AI safety filters in opendataloader-pdf automatically remove invisible or out-of-view content by analyzing contrast ratios for hidden text and bounding box geometry for off-page objects before the content reaches downstream LLMs.**

The opendataloader-pdf repository implements these AI safety filters to protect large language model pipelines from ingesting misleading or malicious content embedded in PDFs. By default, the system scrubs transparent glyphs, off-canvas objects, and other invisible elements that human readers never see.

## Architecture of AI Safety Filters

The filter system consists of three coordinated layers that handle configuration, processing, and CLI integration.

### Configuration Layer

The `FilterConfig` class in [`java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/api/FilterConfig.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/api/FilterConfig.java) holds the default enablement state for each AI safety filter. By default, `filterHiddenText`, `filterOutOfPage`, `filterTinyText`, and `filterHiddenOCG` are all set to **true**.

### Processing Engine

The `ContentFilterProcessor` in [`java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/processors/ContentFilterProcessor.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/processors/ContentFilterProcessor.java) executes the geometric and visual analysis required by the off-page and hidden-text filters.

### CLI Bridge

The `CLIOptions` class in [`java/opendataloader-pdf-cli/src/main/java/org/opendataloader/pdf/cli/CLIOptions.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/java/opendataloader-pdf-cli/src/main/java/org/opendataloader/pdf/cli/CLIOptions.java) parses the `--content-safety-off` argument and toggles the corresponding booleans in `FilterConfig`.

## Hidden-Text Filter Implementation

The hidden-text filter identifies glyphs that are invisible or nearly invisible to human readers by measuring color contrast.

### Contrast Detection Logic

Each `TextChunk` extracted by the VeraPDF WCAG engine is passed to a `ContrastRatioConsumer`. If the computed contrast ratio falls below **1.2** (defined as `MIN_CONTRAST_RATIO` in the source), the chunk is flagged as hidden text.

### Filtering Action

When `filterHiddenText` is enabled, the processor drops the chunk entirely. When disabled, the chunk is retained but marked with `textChunk.setHiddenText(true)` so downstream components can identify it.

The implementation resides in [`java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/processors/HiddenTextProcessor.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/processors/HiddenTextProcessor.java):

```java
private static final double MIN_CONTRAST_RATIO = 1.2d;

for (IObject content : contents) {
    if (content instanceof TextChunk) {
        TextChunk textChunk = (TextChunk) content;
        contrastRatioConsumer.calculateContrastRatio(textChunk);
        if (textChunk.getContrastRatio() < MIN_CONTRAST_RATIO) {
            if (!isFilterHiddenText) {
                textChunk.setHiddenText(true);   // keep but flag
            } else {
                continue;                       // drop from output
            }
        }
    }
    result.add(content);
}

```

## Off-Page Filter Implementation

The off-page filter removes objects that lie outside the visible page area, such as content hidden in the margins or placed on non-visible layers.

### Bounding Box Calculation

The filter retrieves the page bounding box via `DocumentProcessor.getPageBoundingBox(pageNumber)`, which returns the `MediaBox` or `CropBox` rectangle. The coordinates are normalized by translating the lower-left corner to the origin using `pageBoundingBox.move(-leftX, -bottomY)`.

### Geometric Overlap Test

Every `IObject` on the page is tested against the normalized page box. If `pageBoundingBox.notOverlaps(object.getBoundingBox())` returns true, the object is replaced with `null`. After processing, `DocumentProcessor.removeNullObjectsFromList` scrubs the placeholders.

The logic is implemented in [`java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/processors/ContentFilterProcessor.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/processors/ContentFilterProcessor.java):

```java
private static void filterOutOfPageContents(int pageNumber, List<IObject> contents) {
    BoundingBox pageBoundingBox = DocumentProcessor.getPageBoundingBox(pageNumber);
    if (pageBoundingBox == null) {
        return;
    }
    pageBoundingBox.move(-pageBoundingBox.getLeftX(), -pageBoundingBox.getBottomY());
    for (int index = 0; index < contents.size(); index++) {
        IObject object = contents.get(index);
        if (object != null && pageBoundingBox.notOverlaps(object.getBoundingBox())) {
            contents.set(index, null);           // mark for removal
        }
    }
}

```

## End-to-End Processing Pipeline

The `ContentFilterProcessor.getFilteredContents` method orchestrates the entire AI safety pipeline. It conditionally invokes the off-page filter, then delegates to `HiddenTextProcessor.findHiddenText` for contrast analysis, and finally returns the sanitized content list.

```java
public static List<IObject> getFilteredContents(..., Config config) {
    // ... other clean-up steps ...
    if (config.getFilterConfig().isFilterOutOfPage()) {
        filterOutOfPageContents(pageNumber, pageContents);
    }
    // ... more processing ...
    pageContents = HiddenTextProcessor.findHiddenText(
        inputPdfName, pageContents,
        config.getFilterConfig().isFilterHiddenText(),
        config.getPassword());
    // ... final sanitisation steps ...
    return pageContents;
}

```

## Disabling AI Safety Filters via CLI

Users can selectively disable filters using the `--content-safety-off` argument. The `CLIOptions.applyContentSafetyOption` method parses comma-separated values (`hidden-text`, `off-page`, `tiny`, `hidden-ocg`, `all`) and sets the corresponding `FilterConfig` booleans to false.

```java
case "hidden-text":
    config.getFilterConfig().setFilterHiddenText(false);
    break;
case "off-page":
    config.getFilterConfig().setFilterOutOfPage(false);
    break;

```

For example, to process a PDF while keeping hidden text but removing off-page objects:

```bash
opendataloader-pdf mydoc.pdf --content-safety-off hidden-text

```

## Summary

- **AI safety filters** in opendataloader-pdf automatically strip invisible content before it reaches LLM pipelines.
- **Hidden-text filter** detects glyphs with contrast ratios below 1.2 using `HiddenTextProcessor` and removes them when enabled.
- **Off-page filter** uses geometric bounding box tests in `ContentFilterProcessor` to discard objects outside the visible page area.
- **Configuration** is managed by `FilterConfig` with safe defaults (all filters enabled) and exposed via the `--content-safety-off` CLI flag.
- **Integration** happens in `ContentFilterProcessor.getFilteredContents`, which orchestrates the entire sanitization pipeline.

## Frequently Asked Questions

### What triggers the hidden-text filter in opendataloader-pdf?

The hidden-text filter triggers when a text chunk's contrast ratio falls below **1.2**, as defined by the `MIN_CONTRAST_RATIO` constant in `HiddenTextProcessor`. The processor uses `ContrastRatioConsumer` to calculate the ratio between glyph color and background; values below the threshold indicate transparent or invisible text.

### How does the off-page filter handle coordinate transformations?

The off-page filter normalizes page coordinates by translating the bounding box so the lower-left corner becomes the origin. It calls `pageBoundingBox.move(-leftX, -bottomY)` before testing object overlap. This ensures consistent geometric comparison regardless of where the PDF places its coordinate system origin.

### Can I disable AI safety filters programmatically?

Yes. Instantiate a `Config` object, retrieve its `FilterConfig` via `getFilterConfig()`, and call setters like `setFilterHiddenText(false)` or `setFilterOutOfPage(false)`. Pass the modified config to `ContentFilterProcessor.getFilteredContents()` to run the pipeline with selective filtering disabled.

### Where are the default filter settings defined?

Default settings are defined in [`java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/api/FilterConfig.java`](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/api/FilterConfig.java). The class initializes `filterHiddenText`, `filterOutOfPage`, `filterTinyText`, and `filterHiddenOCG` to `true`, ensuring all AI safety filters run unless explicitly disabled via CLI or API.