# How the AI Job Search Framework Verifies ATS Parseability of Generated PDFs

> Learn how the AI Job Search Framework verifies ATS parseability of generated PDFs using a multi-layer system. Ensure your CVs are searchable with text extraction and content validation.

- Repository: [Mads Lorentzen/ai-job-search](https://github.com/MadsLorentzen/ai-job-search)
- Tags: how-to-guide
- Published: 2026-09-01

---

**The AI Job Search Framework validates ATS parseability through a multi-layer verification system in [`tools/verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/verify_pdf.py) that extracts text using pypdf or Poppler's pdftotext, then validates page count, character minimums, and required content strings to ensure every generated CV contains a searchable text layer.**

The MadsLorentzen/ai-job-search project includes a robust verification pipeline to ensure that AI-generated resumes remain machine-readable by applicant tracking systems. Since image-only PDFs fail ATS parsing and render candidates invisible to recruiters, the framework's `verify_pdf` utility automatically validates that each document contains extractable text content before final delivery.

## The Text Extraction Pipeline

The verification process begins in [`tools/verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/verify_pdf.py) with a dual-strategy extraction approach designed to maximize compatibility across different PDF generation methods.

### Primary and Fallback Extraction Methods

The framework attempts extraction through two distinct engines. First, it invokes `_extract_pypdf` at lines 54-69 to parse the PDF using the pure-Python pypdf library. If pypdf is unavailable, raises an exception, or returns empty text after normalization, the system automatically falls back to `_extract_pdftotext` at lines 72-78, which shells out to Poppler's `pdftotext` command-line utility.

This redundancy ensures that verification succeeds regardless of the Python environment's specific dependencies. The extraction function returns a tuple containing the extractor name ("pypdf" or "pdftotext"), the raw text content, and the total page count (lines 80-88), enabling downstream validators to audit which engine successfully parsed the document.

## Validation Criteria for ATS Compatibility

Beyond mere text presence, the framework enforces three specific constraints at lines 11-27 that predict ATS parsing success.

### Page Count Verification

When generating standardized CVs with strict length requirements, the tool accepts an `expected_pages` parameter (lines 11-15). The actual page count returned by the extractor is compared against this target, and a `VerificationError` is raised immediately if the counts diverge. This prevents formatting issues that might cause a two-page CV to render as three pages in certain PDF readers.

### Minimum Character Threshold

The framework validates that the extracted text contains meaningful content through the `min_chars` parameter, which defaults to 1 but can be set higher (lines 16-21). The extracted text undergoes whitespace normalization before length calculation, ensuring that invisible characters or padding whitespace cannot satisfy the requirement. Documents falling below the threshold fail verification with a descriptive error.

### Required Content Strings

Using the `--contains` CLI flag or `required_text` argument, users can specify mandatory phrases that must exist within the PDF's text layer (lines 23-27). Common requirements include section headers like "Professional Experience" or "Education." The tool normalizes both the extracted text and search strings before performing substring matching, ensuring that case or spacing variations do not cause false negatives.

## Diagnostic and Debug Capabilities

The verification utility includes a `--dump-text` option that writes the extracted text layer to a specified file path before any validation checks execute. According to lines 97-106 in [`tools/verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/verify_pdf.py), this diagnostic output allows developers to inspect exactly what content the ATS would see, even when verification fails due to page count or content requirements. The dump occurs prior to validation logic, ensuring raw extraction results are always preserved for debugging.

## Integration and Security

The verification system integrates deeply with the framework's architecture. The `verify_pdf` function serves as the primary entry point for both command-line and programmatic validation, beginning with a file existence check at lines 91-94. Unit tests in [`tests/test_verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tests/test_verify_pdf.py) mock both extraction backends to validate error handling without dependencies on external binaries. Additionally, [`tools/security_guards.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/security_guards.py) explicitly lists `verify_pdf` as a permitted command within the project's sandboxed security model, allowing automated pipelines to execute verification safely.

## Usage Examples

Verify a generated CV from the command line, requiring exactly two pages and specific content:

```bash
python -m tools.verify_pdf \
    output/cv.pdf \
    --pages 2 \
    --contains "Professional Experience"

```

Dump extracted text for debugging purposes:

```bash
python -m tools.verify_pdf \
    output/cv.pdf \
    --dump-text tmp/cv_extracted.txt

```

Integrate verification programmatically within the framework:

```python
from tools.verify_pdf import verify_pdf, VerificationError

try:
    extractor, text, pages = verify_pdf(
        pdf_path="output/cv.pdf",
        expected_pages=2,
        min_chars=100,
        required_text=("Professional Experience", "Education"),
        dump_text="debug/cv.txt",
    )
    print(f"PDF verified (extractor={extractor}, pages={pages})")
except VerificationError as err:
    print(f"PDF verification failed: {err}")

```

## Summary

- The framework validates ATS parseability through [`tools/verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/verify_pdf.py), which extracts text using pypdf first, falling back to Poppler's pdftotext if the primary method fails or returns empty content.
- Verification enforces page count accuracy via `expected_pages`, minimum character thresholds via `min_chars`, and the presence of required text strings to ensure machine-readability.
- The `--dump-text` diagnostic flag preserves extraction results for debugging at lines 97-106, while comprehensive unit tests in [`tests/test_verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tests/test_verify_pdf.py) ensure reliability across extraction backends.

## Frequently Asked Questions

### What happens if neither pypdf nor pdftotext can extract text from the PDF?

If both extraction methods fail or return empty normalized text, the `verify_pdf` function raises a `VerificationError`, indicating that the PDF likely contains only raster images and lacks the searchable text layer required by ATS systems. This prevents "image-only" resumes from passing validation.

### Can I verify multiple required phrases in a single PDF check?

Yes, the `required_text` argument accepts a tuple or list of strings (lines 23-27), and the verification logic confirms that all specified phrases are present in the normalized extracted text before returning successfully. If any required string is missing, the tool raises `VerificationError` immediately.

### Does the framework verify PDFs generated by external tools, or only its own output?

While designed for the framework's AI-generated CVs, the `verify_pdf` utility functions as a standalone tool capable of validating any PDF file path passed to it, as implemented in [`tools/verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/verify_pdf.py). The file existence check at lines 91-94 ensures the path is valid before extraction begins.

### How does the security model handle the verify_pdf command?

According to [`tools/security_guards.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/security_guards.py), `verify_pdf` is explicitly whitelisted as a permitted command within the project's sandboxed execution environment. This allows automated pipelines to run verification without triggering security restrictions, ensuring that PDF validation can occur safely within the framework's guarded execution context.