What Types of Sensitive Data Does the --sanitize Option Redact by Default in OpenDataLoader PDF
The --sanitize option in OpenDataLoader PDF automatically redacts 10 categories of sensitive data—including email addresses, international phone numbers, credit card numbers, IP addresses, MAC addresses, and government IDs—replacing them with standardized placeholders to prevent PII leakage in converted documents.
The --sanitize flag is a critical privacy feature in the opendataloader-project/opendataloader-pdf repository that enables automatic detection and redaction of personally identifiable information (PII) during PDF conversion. When activated, the library leverages predefined regex patterns defined in FilterConfig.java to identify sensitive data types and replaces them with harmless placeholders via the ContentSanitizer utility before generating the final output.
How the --sanitize Option Works
The sanitization pipeline activates when the --sanitize CLI flag or the filterSensitiveData API parameter is set to true.
In FilterConfig.java, located at java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/api/FilterConfig.java, the library compiles regex patterns into SanitizationRule objects during initialization. These rules define the matching patterns for each sensitive data type and their corresponding replacement placeholders. The ContentSanitizer.java class then traverses the PDF content tree and applies these compiled rules to every text element, ensuring that all matching patterns are replaced with placeholders before the output is serialized.
Default Sensitive Data Types Redacted
The --sanitize option targets 10 distinct categories of sensitive information. Each category is defined by a specific regex pattern in FilterConfig.java and replaced with a standardized placeholder.
Email Addresses
Email addresses matching standard RFC patterns are replaced with email@example.com. The regex pattern is defined in [FilterConfig.java lines 38-41](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/api/FilterConfig.java#L38-L41).
International Phone Numbers
Phone numbers beginning with the + prefix, indicating international format, are replaced with +00-0000-0000. This pattern is configured in [FilterConfig.java lines 42-45](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/api/FilterConfig.java#L42-L45).
Alphanumeric IDs
Identification codes consisting of 1-2 letters followed by 6-9 digits—common in government ID systems and membership programs—are replaced with AA0000000. The regex is defined in [FilterConfig.java lines 46-49](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/api/FilterConfig.java#L46-L49).
Credit Card Numbers
Payment card numbers formatted as four groups of four digits, with optional hyphens, are replaced with 0000-0000-0000-0000. This pattern is specified in [FilterConfig.java lines 50-53](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/api/FilterConfig.java#L50-L53).
Long Numeric Strings
Generic sensitive numeric sequences of 10-18 consecutive digits are replaced with 0000000000000000. This pattern catches miscellaneous account numbers or identifiers not matching specific formats like credit cards. See [FilterConfig.java lines 54-57](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/api/FilterConfig.java#L54-L57).
IPv4 Addresses
Standard IPv4 addresses in dot-decimal notation are replaced with 0.0.0.0. The regex is defined in [FilterConfig.java lines 58-61](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/api/FilterConfig.java#L58-L61).
IPv6 Addresses
IPv6 addresses in standard hexadecimal colon notation are replaced with 0.0.0.0::1. This pattern is found in [FilterConfig.java lines 62-65](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/api/FilterConfig.java#L62-L65).
MAC Addresses
Media Access Control addresses in standard colon-separated format are replaced with 00:00:00:00:00:00. The regex is defined in [FilterConfig.java lines 66-68](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/api/FilterConfig.java#L66-L68).
15-Digit Numeric IDs
Specific 15-digit numeric identifiers, such as Korean resident registration numbers, are replaced with 000000000000000. This pattern is configured in [FilterConfig.java lines 70-73](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/api/FilterConfig.java#L70-L73).
URLs
HTTP and HTTPS URLs are replaced with https://example.com. This pattern is defined in [FilterConfig.java lines 74-76](https://github.com/opendataloader-project/opendataloader-pdf/blob/main/java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/api/FilterConfig.java#L74-L76).
Usage Examples
You can activate the --sanitize option through the command line interface or programmatically via the official Python and Node.js wrappers.
Command Line Interface
# Redact all default sensitive data types
opendataloader-pdf document.pdf --sanitize -f json
Python API
from opendataloader_pdf import convert
# Enable sanitization via the sanitize parameter
result = convert("document.pdf", {"sanitize": True, "format": ["json"]})
print(result) # Emails, URLs, IPs, etc. are replaced with placeholders
Node.js API
import { convert } from "opendataloader-pdf";
await convert("document.pdf", { sanitize: true, format: ["json"] });
Implementation Details
The sanitization system is implemented across three core components in the opendataloader-pdf repository.
FilterConfig.java defines the default sanitization rules. Located at java/opendataloader-pdf-core/src/main/java/org/opendataloader/pdf/api/FilterConfig.java, this class compiles regex patterns for the 10 sensitive data types into SanitizationRule objects during construction.
ContentSanitizer.java handles the execution. This utility class traverses the PDF content tree and applies the compiled rules to every text element, ensuring matching patterns are replaced before output generation.
CLIOptions.java exposes the feature to the command line. Located in the CLI module at java/opendataloader-pdf-cli/src/main/java/org/opendataloader/pdf/cli/CLIOptions.java, this class maps the --sanitize flag to the filterSensitiveData configuration parameter.
Summary
- The
--sanitizeoption in opendataloader-pdf automatically redacts 10 categories of sensitive data using predefined regex patterns. - Default redactions include emails, phone numbers, credit cards, IP addresses (IPv4 and IPv6), MAC addresses, alphanumeric IDs, 15-digit government IDs, URLs, and long numeric strings.
- Configuration is centralized in
FilterConfig.java, with execution handled byContentSanitizer.javaand CLI exposure viaCLIOptions.java. - The feature is accessible via CLI (
--sanitizeflag), Python (sanitize: True), and Node.js (sanitize: true) APIs.
Frequently Asked Questions
Does the --sanitize option modify the original PDF file?
No, the --sanitize option only affects the conversion output. The original PDF file remains completely unmodified. Sanitization is applied during the content extraction phase by the ContentSanitizer class, ensuring that sensitive data is replaced with placeholders only in the generated output before it is written to disk or returned via the API.
Can I customize which sensitive data types are redacted?
The default implementation in FilterConfig.java treats sanitization as an all-or-nothing feature. All 10 regex patterns are compiled into SanitizationRule objects during initialization, and the CLI --sanitize flag enables the entire set. Selective configuration of individual patterns would require modifying the FilterConfig class or extending the SanitizationRule configuration, as the current API does not expose granular toggles for specific data types.
Is there a performance penalty when using --sanitize?
Yes, enabling --sanitize introduces processing overhead because the ContentSanitizer must traverse the entire PDF content tree and apply regex matching against all patterns for every text element. However, the impact is generally minimal for standard documents because the regex patterns are compiled once during FilterConfig initialization, optimizing the matching performance during the sanitization pass. The privacy compliance benefits typically outweigh the modest performance cost.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →