ATS Text-Layer Verification Process: How `tools/verify_pdf.py` Ensures PDF Compatibility
The tools/verify_pdf.py script in the MadsLorentzen/ai-job-search repository validates that PDFs contain an extractable text layer and correct page counts, ensuring compatibility with applicant tracking systems (ATS) that cannot index image-based documents.
The MadsLorentzen/ai-job-search repository includes a specialized utility for validating PDF output before submission to candidate tracking platforms. The ATS text-layer verification process implemented in tools/verify_pdf.py performs automated checks to confirm that generated resumes contain machine-readable text rather than rasterized images, while also verifying pagination requirements.
What is ATS Text-Layer Verification?
Applicant tracking systems require two fundamental properties from uploaded PDFs: the correct number of pages and an embedded text layer that allows indexing and keyword searching. The verify_pdf.py tool validates both properties programmatically, preventing rejected applications caused by unscannable image-based PDFs.
Architecture of the Verification Pipeline
Command-Line Interface with build_parser()
The script begins by collecting validation constraints through build_parser(), which constructs an argparse.ArgumentParser instance. This function defines arguments for PDF path, expected page count, minimum character thresholds, required text snippets, and optional text dump locations.
Dual-Strategy Text Extraction
The verification engine implements a resilient fallback mechanism for text extraction. The _extract_pypdf() function attempts to extract text using the pure-Python pypdf library first, returning None on any failure. When this approach fails or returns empty content, the system automatically falls back to _extract_pdftotext(), which invokes the Poppler pdftotext command-line utility. This design prioritizes speed and portability while ensuring compatibility across diverse CI environments.
Page Count Retrieval via pdfinfo
When the Poppler-based extraction path activates, the script retrieves page counts by executing pdfinfo and parsing the output through parse_page_count(). This function specifically scans for the Pages: line in the utility's output, converting the value to an integer for validation against user-specified requirements.
Validation Logic in verify_pdf()
The core verify_pdf() function orchestrates all verification checks, comparing extracted text against minimum character counts, searching for required string snippets, and validating page counts. Any constraint violation raises a custom VerificationError exception containing specific failure details. The function optionally writes extracted text to a dump file for debugging purposes.
Robust Error Handling with run_tool()
All subprocess invocations route through run_tool(), which wraps subprocess.run calls to normalize exceptions and non-zero exit codes into descriptive VerificationError instances. This abstraction ensures consistent error reporting whether tools are missing or commands fail during execution.
Command-Line Usage
Verify that a resume PDF contains exactly two pages and includes specific professional keywords:
python -m tools.verify_pdf resume.pdf --pages 2 --contains "Software Engineer"
The script exits with status code zero on success or prints an error message and returns non-zero on validation failure.
Programmatic Integration
Embed verification directly into Python workflows:
from pathlib import Path
from tools.verify_pdf import verify_pdf, VerificationError
pdf_path = Path("resume.pdf")
try:
extractor, text, pages = verify_pdf(
pdf_path,
expected_pages=2,
min_chars=10,
required_text=("Software Engineer", "Python"),
dump_text=Path("resume_text.txt"),
)
print(f"✅ Verified with {extractor}: {pages} pages")
except VerificationError as err:
print(f"❌ Verification failed: {err}")
This pattern captures VerificationError to handle validation failures gracefully while optionally persisting extracted text for manual inspection.
Test Coverage
The verification logic includes comprehensive unit tests in tests/test_verify_pdf.py, which exercise the extractor fallback mechanism, page-count validation edge cases, and required-text matching functionality. These tests ensure the dual-extraction strategy and error handling paths function correctly across different environments.
Summary
- The
tools/verify_pdf.pyscript validates PDFs for ATS compatibility by checking page counts and text extractability. - The implementation uses a primary
pypdfextraction method with an automatic fallback to Poppler'spdftotextutility. - The
verify_pdf()function enforces constraints through theVerificationErrorexception hierarchy. - Command-line and programmatic interfaces support both CI/CD pipelines and custom Python workflows.
- Test coverage in
tests/test_verify_pdf.pyvalidates the dual-strategy extraction and error handling logic.
Frequently Asked Questions
Why does the script use both pypdf and pdftotext?
The dual-strategy approach prioritizes the pure-Python pypdf library for speed and reduced dependencies, but automatically falls back to the Poppler pdftotext tool when the primary method fails. This ensures the ATS text-layer verification process works reliably across different operating systems and CI environments without requiring specific binary installations.
What happens if a PDF contains scanned images instead of text?
If the PDF contains no extractable text layer, both _extract_pypdf() and _extract_pdftotext() will return empty or minimal content. The verify_pdf() function then compares this result against the min_chars parameter (defaulting to the user-specified threshold) and raises a VerificationError indicating the document lacks the required machine-readable content.
How does the script validate page counts?
When using the Poppler fallback path, the script executes pdfinfo via the parse_page_count() function to parse the Pages: field from the metadata. This integer is compared against the expected_pages argument passed to verify_pdf(), triggering a VerificationError if the actual count differs from the requirement.
Can I use this verification in automated testing pipelines?
Yes. The script is designed for CI/CD integration through its command-line interface and non-zero exit codes on failure. Alternatively, import verify_pdf and VerificationError directly into Python test suites to validate PDF generation as part of automated resume-building workflows.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →