# ATS Text-Layer Verification Process: How `tools/verify_pdf.py` Ensures PDF Compatibility

> Learn the ATS text-layer verification process using tools/verify_pdf.py. Ensure your PDFs are compatible with applicant tracking systems by checking for extractable text and correct page counts.

- Repository: [Mads Lorentzen/ai-job-search](https://github.com/MadsLorentzen/ai-job-search)
- Tags: how-to-guide
- Published: 2026-08-31

---

**The [`tools/verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/verify_pdf.py) script in the MadsLorentzen/ai-job-search repository validates that PDFs contain an extractable text layer and correct page counts, ensuring compatibility with applicant tracking systems (ATS) that cannot index image-based documents.**

The MadsLorentzen/ai-job-search repository includes a specialized utility for validating PDF output before submission to candidate tracking platforms. The ATS text-layer verification process implemented in [`tools/verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/verify_pdf.py) performs automated checks to confirm that generated resumes contain machine-readable text rather than rasterized images, while also verifying pagination requirements.

## What is ATS Text-Layer Verification?

Applicant tracking systems require two fundamental properties from uploaded PDFs: the correct number of pages and an embedded text layer that allows indexing and keyword searching. The [`verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/verify_pdf.py) tool validates both properties programmatically, preventing rejected applications caused by unscannable image-based PDFs.

## Architecture of the Verification Pipeline

### Command-Line Interface with `build_parser()`

The script begins by collecting validation constraints through `build_parser()`, which constructs an `argparse.ArgumentParser` instance. This function defines arguments for PDF path, expected page count, minimum character thresholds, required text snippets, and optional text dump locations.

### Dual-Strategy Text Extraction

The verification engine implements a resilient fallback mechanism for text extraction. The `_extract_pypdf()` function attempts to extract text using the pure-Python `pypdf` library first, returning `None` on any failure. When this approach fails or returns empty content, the system automatically falls back to `_extract_pdftotext()`, which invokes the Poppler `pdftotext` command-line utility. This design prioritizes speed and portability while ensuring compatibility across diverse CI environments.

### Page Count Retrieval via `pdfinfo`

When the Poppler-based extraction path activates, the script retrieves page counts by executing `pdfinfo` and parsing the output through `parse_page_count()`. This function specifically scans for the `Pages:` line in the utility's output, converting the value to an integer for validation against user-specified requirements.

### Validation Logic in `verify_pdf()`

The core `verify_pdf()` function orchestrates all verification checks, comparing extracted text against minimum character counts, searching for required string snippets, and validating page counts. Any constraint violation raises a custom `VerificationError` exception containing specific failure details. The function optionally writes extracted text to a dump file for debugging purposes.

### Robust Error Handling with `run_tool()`

All subprocess invocations route through `run_tool()`, which wraps `subprocess.run` calls to normalize exceptions and non-zero exit codes into descriptive `VerificationError` instances. This abstraction ensures consistent error reporting whether tools are missing or commands fail during execution.

## Command-Line Usage

Verify that a resume PDF contains exactly two pages and includes specific professional keywords:

```bash
python -m tools.verify_pdf resume.pdf --pages 2 --contains "Software Engineer"

```

The script exits with status code zero on success or prints an error message and returns non-zero on validation failure.

## Programmatic Integration

Embed verification directly into Python workflows:

```python
from pathlib import Path
from tools.verify_pdf import verify_pdf, VerificationError

pdf_path = Path("resume.pdf")
try:
    extractor, text, pages = verify_pdf(
        pdf_path,
        expected_pages=2,
        min_chars=10,
        required_text=("Software Engineer", "Python"),
        dump_text=Path("resume_text.txt"),
    )
    print(f"✅ Verified with {extractor}: {pages} pages")
except VerificationError as err:
    print(f"❌ Verification failed: {err}")

```

This pattern captures `VerificationError` to handle validation failures gracefully while optionally persisting extracted text for manual inspection.

## Test Coverage

The verification logic includes comprehensive unit tests in [`tests/test_verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tests/test_verify_pdf.py), which exercise the extractor fallback mechanism, page-count validation edge cases, and required-text matching functionality. These tests ensure the dual-extraction strategy and error handling paths function correctly across different environments.

## Summary

- The [`tools/verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/verify_pdf.py) script validates PDFs for ATS compatibility by checking page counts and text extractability.
- The implementation uses a primary `pypdf` extraction method with an automatic fallback to Poppler's `pdftotext` utility.
- The `verify_pdf()` function enforces constraints through the `VerificationError` exception hierarchy.
- Command-line and programmatic interfaces support both CI/CD pipelines and custom Python workflows.
- Test coverage in [`tests/test_verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tests/test_verify_pdf.py) validates the dual-strategy extraction and error handling logic.

## Frequently Asked Questions

### Why does the script use both `pypdf` and `pdftotext`?

The dual-strategy approach prioritizes the pure-Python `pypdf` library for speed and reduced dependencies, but automatically falls back to the Poppler `pdftotext` tool when the primary method fails. This ensures the ATS text-layer verification process works reliably across different operating systems and CI environments without requiring specific binary installations.

### What happens if a PDF contains scanned images instead of text?

If the PDF contains no extractable text layer, both `_extract_pypdf()` and `_extract_pdftotext()` will return empty or minimal content. The `verify_pdf()` function then compares this result against the `min_chars` parameter (defaulting to the user-specified threshold) and raises a `VerificationError` indicating the document lacks the required machine-readable content.

### How does the script validate page counts?

When using the Poppler fallback path, the script executes `pdfinfo` via the `parse_page_count()` function to parse the `Pages:` field from the metadata. This integer is compared against the `expected_pages` argument passed to `verify_pdf()`, triggering a `VerificationError` if the actual count differs from the requirement.

### Can I use this verification in automated testing pipelines?

Yes. The script is designed for CI/CD integration through its command-line interface and non-zero exit codes on failure. Alternatively, import `verify_pdf` and `VerificationError` directly into Python test suites to validate PDF generation as part of automated resume-building workflows.