# How the `verify_pdf.py` Tool Validates ATS Parseability in PDF Resumes

> Discover how verify_pdf.py validates ATS parseability in resumes by checking text layers, whitespace, page count, and content. Ensure your resume is processed correctly.

- Repository: [Mads Lorentzen/ai-job-search](https://github.com/MadsLorentzen/ai-job-search)
- Tags: how-to-guide
- Published: 2026-09-03

---

**The [`verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/verify_pdf.py) tool ensures ATS parseability by extracting the PDF's text layer through a priority fallback chain (pypdf → pdftotext), normalizing whitespace, and enforcing configurable checks for page count, minimum character density, and mandatory content strings.**

The `ai-job-search` repository by MadsLorentzen includes a specialized verification utility to ensure generated resumes remain machine-readable. Located at [`tools/verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/verify_pdf.py), this tool validates that a PDF contains a searchable text layer rather than just scanned images, which is critical for Applicant Tracking System (ATS) compatibility. Understanding how this verification works helps developers debug resume generation pipelines and prevent submission failures.

## Dual-Path Text Extraction Strategy

The verification process begins with robust text extraction designed to handle various PDF generation methods. According to the source code in [`tools/verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/verify_pdf.py), the tool implements a priority-based fallback mechanism to ensure text recovery even when pure-Python libraries fail.

### Primary Extraction via pypdf

First, the tool attempts extraction using the pure-Python library **pypdf** through the `_extract_pypdf` function (lines 55-69). This method reads the PDF content programmatically without requiring external system dependencies, making it ideal for portable environments.

### Fallback to pdftotext

If **pypdf** is unavailable, raises an error, or returns empty content after whitespace normalization, the tool automatically falls back to the Poppler utility **pdftotext** via `_extract_pdftotext` (lines 72-77). This external binary handles complex PDF structures—such as those with embedded fonts or unusual compression—that might trip up pure-Python parsers, ensuring maximum compatibility across different resume generators.

## Normalization and Validation Checks

Once extracted, the text undergoes normalization and a series of validation gates that ensure ATS readability.

### Whitespace Normalization

The `normalize_text` helper (lines 50-52) collapses any whitespace sequence into a single space. This creates a consistent baseline for character counting and string matching, eliminating formatting discrepancies that could trigger false negatives during verification.

### Page Count Verification

When the `--pages` argument is supplied, the tool compares the extracted page count against the expected value using `parse_page_count` (lines 11-13). A mismatch immediately raises a `VerificationError` (implemented at lines 111-115), preventing multi-page overflows or missing content from reaching recruiters.

### Minimum Character Threshold

The `--min-chars` parameter (default: 1) ensures the text layer contains sufficient machine-readable content. The implementation at lines 16-21 checks if the normalized text length falls below the threshold, raising an error if the PDF lacks a substantive text layer—a strong indicator of image-only scanned documents that ATS cannot parse.

### Required Content Validation

Using the `--contains` flag (or `required_text` argument), the tool validates mandatory strings such as candidate names or section headers. The loop at lines 23-28 normalizes each required string and searches for it within the normalized PDF text. Missing strings trigger a `VerificationError`, ensuring critical information is actually present in the extractable layer.

### Debug Text Extraction

For troubleshooting generation issues, the tool supports raw text inspection. When `--dump-text` is provided, the extracted layer is written to a specified file before any verification checks execute (lines 97-107). This preserves debugging information even when validation fails, allowing developers to inspect exactly what content is visible to ATS software.

## Command-Line and Programmatic Usage

### Command-Line Verification

```bash
python -m tools.verify_pdf my_resume.pdf --pages 2 --min-chars 100 --contains "John Doe" --dump-text ./tmp/extracted.txt

```

This command validates that `my_resume.pdf` contains exactly two pages, at least 100 normalized characters, the phrase "John Doe", and saves the extracted text to [`./tmp/extracted.txt`](https://github.com/MadsLorentzen/ai-job-search/blob/main/./tmp/extracted.txt) for inspection.

### Programmatic Integration

```python
from tools.verify_pdf import verify_pdf, VerificationError

try:
    extractor, text, pages = verify_pdf(
        "candidate.pdf",
        expected_pages=1,
        min_chars=50,
        required_text=["Experience", "Education"],
    )
    print(f"✅ PDF is ATS-parseable (extracted by {extractor})")
except VerificationError as err:
    print(f"❌ PDF verification failed: {err}")

```

This pattern allows CI/CD pipelines to automatically reject resumes that fail ATS compatibility standards, returning the name of the successful extractor (`pypdf` or `pdftotext`) along with the extracted content and page count upon success.

## Summary

- **Dual extraction strategy**: The tool prioritizes pypdf (lines 55-69) and falls back to pdftotext (lines 72-77) to maximize PDF compatibility across different generation methods.
- **Whitespace normalization**: The `normalize_text` function (lines 50-52) ensures consistent text processing by collapsing whitespace sequences before validation.
- **Configurable validation**: Parameters for page count, minimum characters, and required strings provide granular control over ATS requirements.
- **Debug support**: The `--dump-text` option (lines 97-107) extracts raw text before validation for troubleshooting extraction failures.
- **Clear error reporting**: All failures raise descriptive `VerificationError` exceptions indicating exactly which check failed, enabling rapid pipeline adjustments.

## Frequently Asked Questions

### Why does [`verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/verify_pdf.py) use two different extraction methods?

The tool uses **pypdf** as the primary extractor because it is a pure-Python dependency that is easy to install in most environments. However, some PDF generation methods create complex structures that pypdf cannot parse. The fallback to **pdftotext** (lines 72-77) ensures that even problematic PDFs can be validated for ATS parseability without requiring manual intervention or additional Python packages.

### What happens if my PDF contains scanned images without text?

If the PDF lacks a searchable text layer, both extraction methods will return empty or near-empty strings. When this occurs, the `normalize_text` function (lines 50-52) will produce minimal output, triggering the `--min-chars` validation check (lines 16-21) and raising a `VerificationError`. This prevents image-based resumes from passing verification and failing silently in ATS systems.

### Can I verify multiple required phrases at once?

Yes. The `--contains` flag accepts multiple values, and the implementation at lines 23-28 iterates through each required string in the `required_text` list. Each phrase is normalized and searched independently within the normalized PDF content. If any required phrase is missing, the tool immediately raises a `VerificationError` specifying which content was not found.

### How does the tool handle PDFs with unusual whitespace formatting?

The `normalize_text` helper collapses any sequence of whitespace characters—including tabs, newlines, and multiple consecutive spaces—into a single space. This normalization occurs before character counting and string matching, ensuring that formatting differences between PDF generators do not cause false verification failures while preserving the actual text content for ATS parsing.