How the `verify_pdf.py` Tool Validates ATS Parseability in PDF Resumes
The verify_pdf.py tool ensures ATS parseability by extracting the PDF's text layer through a priority fallback chain (pypdf → pdftotext), normalizing whitespace, and enforcing configurable checks for page count, minimum character density, and mandatory content strings.
The ai-job-search repository by MadsLorentzen includes a specialized verification utility to ensure generated resumes remain machine-readable. Located at tools/verify_pdf.py, this tool validates that a PDF contains a searchable text layer rather than just scanned images, which is critical for Applicant Tracking System (ATS) compatibility. Understanding how this verification works helps developers debug resume generation pipelines and prevent submission failures.
Dual-Path Text Extraction Strategy
The verification process begins with robust text extraction designed to handle various PDF generation methods. According to the source code in tools/verify_pdf.py, the tool implements a priority-based fallback mechanism to ensure text recovery even when pure-Python libraries fail.
Primary Extraction via pypdf
First, the tool attempts extraction using the pure-Python library pypdf through the _extract_pypdf function (lines 55-69). This method reads the PDF content programmatically without requiring external system dependencies, making it ideal for portable environments.
Fallback to pdftotext
If pypdf is unavailable, raises an error, or returns empty content after whitespace normalization, the tool automatically falls back to the Poppler utility pdftotext via _extract_pdftotext (lines 72-77). This external binary handles complex PDF structures—such as those with embedded fonts or unusual compression—that might trip up pure-Python parsers, ensuring maximum compatibility across different resume generators.
Normalization and Validation Checks
Once extracted, the text undergoes normalization and a series of validation gates that ensure ATS readability.
Whitespace Normalization
The normalize_text helper (lines 50-52) collapses any whitespace sequence into a single space. This creates a consistent baseline for character counting and string matching, eliminating formatting discrepancies that could trigger false negatives during verification.
Page Count Verification
When the --pages argument is supplied, the tool compares the extracted page count against the expected value using parse_page_count (lines 11-13). A mismatch immediately raises a VerificationError (implemented at lines 111-115), preventing multi-page overflows or missing content from reaching recruiters.
Minimum Character Threshold
The --min-chars parameter (default: 1) ensures the text layer contains sufficient machine-readable content. The implementation at lines 16-21 checks if the normalized text length falls below the threshold, raising an error if the PDF lacks a substantive text layer—a strong indicator of image-only scanned documents that ATS cannot parse.
Required Content Validation
Using the --contains flag (or required_text argument), the tool validates mandatory strings such as candidate names or section headers. The loop at lines 23-28 normalizes each required string and searches for it within the normalized PDF text. Missing strings trigger a VerificationError, ensuring critical information is actually present in the extractable layer.
Debug Text Extraction
For troubleshooting generation issues, the tool supports raw text inspection. When --dump-text is provided, the extracted layer is written to a specified file before any verification checks execute (lines 97-107). This preserves debugging information even when validation fails, allowing developers to inspect exactly what content is visible to ATS software.
Command-Line and Programmatic Usage
Command-Line Verification
python -m tools.verify_pdf my_resume.pdf --pages 2 --min-chars 100 --contains "John Doe" --dump-text ./tmp/extracted.txt
This command validates that my_resume.pdf contains exactly two pages, at least 100 normalized characters, the phrase "John Doe", and saves the extracted text to ./tmp/extracted.txt for inspection.
Programmatic Integration
from tools.verify_pdf import verify_pdf, VerificationError
try:
extractor, text, pages = verify_pdf(
"candidate.pdf",
expected_pages=1,
min_chars=50,
required_text=["Experience", "Education"],
)
print(f"✅ PDF is ATS-parseable (extracted by {extractor})")
except VerificationError as err:
print(f"❌ PDF verification failed: {err}")
This pattern allows CI/CD pipelines to automatically reject resumes that fail ATS compatibility standards, returning the name of the successful extractor (pypdf or pdftotext) along with the extracted content and page count upon success.
Summary
- Dual extraction strategy: The tool prioritizes pypdf (lines 55-69) and falls back to pdftotext (lines 72-77) to maximize PDF compatibility across different generation methods.
- Whitespace normalization: The
normalize_textfunction (lines 50-52) ensures consistent text processing by collapsing whitespace sequences before validation. - Configurable validation: Parameters for page count, minimum characters, and required strings provide granular control over ATS requirements.
- Debug support: The
--dump-textoption (lines 97-107) extracts raw text before validation for troubleshooting extraction failures. - Clear error reporting: All failures raise descriptive
VerificationErrorexceptions indicating exactly which check failed, enabling rapid pipeline adjustments.
Frequently Asked Questions
Why does verify_pdf.py use two different extraction methods?
The tool uses pypdf as the primary extractor because it is a pure-Python dependency that is easy to install in most environments. However, some PDF generation methods create complex structures that pypdf cannot parse. The fallback to pdftotext (lines 72-77) ensures that even problematic PDFs can be validated for ATS parseability without requiring manual intervention or additional Python packages.
What happens if my PDF contains scanned images without text?
If the PDF lacks a searchable text layer, both extraction methods will return empty or near-empty strings. When this occurs, the normalize_text function (lines 50-52) will produce minimal output, triggering the --min-chars validation check (lines 16-21) and raising a VerificationError. This prevents image-based resumes from passing verification and failing silently in ATS systems.
Can I verify multiple required phrases at once?
Yes. The --contains flag accepts multiple values, and the implementation at lines 23-28 iterates through each required string in the required_text list. Each phrase is normalized and searched independently within the normalized PDF content. If any required phrase is missing, the tool immediately raises a VerificationError specifying which content was not found.
How does the tool handle PDFs with unusual whitespace formatting?
The normalize_text helper collapses any sequence of whitespace characters—including tabs, newlines, and multiple consecutive spaces—into a single space. This normalization occurs before character counting and string matching, ensuring that formatting differences between PDF generators do not cause false verification failures while preserving the actual text content for ATS parsing.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →