Understanding Parseability Checks in ATS Verification: A Technical Deep Dive
The ai-job-search repository performs six systematic parseability checks in ATS verification through tools/verify_pdf.py, ensuring that generated résumé PDFs contain machine-readable text layers, correct page counts, and required content before submission to Applicant Tracking Systems.
The MadsLorentzen/ai-job-search toolkit automates the creation of ATS-friendly résumés, but creating a PDF is only half the battle. These parseability checks in ATS verification confirm that the document can be successfully ingested by employer parsing software, preventing your application from being rejected due to unreadable formatting or missing text layers.
What Are Parseability Checks in ATS Verification?
Parseability checks are automated validations that verify a PDF file's internal structure and content accessibility. Unlike visual formatting checks, these inspect the document's text layer, metadata, and character encoding to ensure ATS software can extract meaningful data. In the ai-job-search codebase, these checks are implemented as a pipeline that raises a VerificationError immediately upon detecting any condition that would cause an ATS to fail reading the résumé.
The Six Parseability Checks Explained
The verification logic in tools/verify_pdf.py implements six distinct checks that execute sequentially during the validation process.
1. Page-Count Extraction
Before analyzing content, the system verifies the document's physical structure. The parse_page_count function (lines 44‑48) retrieves the total number of pages using either the pdfinfo command-line tool or the page count metadata from pypdf. This establishes a baseline for subsequent structural validations.
2. ATS-Readable Text Extraction
The most critical check ensures the PDF contains an extractable text layer rather than image-only content. The system attempts extraction via _extract_pypdf (lines 54‑69) using the pypdf library first. If this returns empty or unavailable content, it automatically falls back to _extract_pdftotext (lines 72‑77), which invokes Poppler’s pdftotext utility. This dual-method approach maximizes compatibility across different PDF generation methods.
3. Text Normalisation
Raw extracted text often contains irregular whitespace that can interfere with string matching. The normalize_text function (lines 50‑52) collapses multiple whitespace characters into single spaces, ensuring that subsequent checks for required content are independent of formatting variations or line-break inconsistencies.
4. Exact Page-Count Validation
When the caller supplies an expected page count via the --pages argument, the verify_pdf function (lines 11‑15) performs an exact equality check against the extracted page count. Any discrepancy triggers a VerificationError, preventing submission of résumés that exceed single-page limits or fail to meet minimum length requirements.
5. Minimum Character Count
To prevent submission of scanned images or corrupted PDFs that contain no readable text, the system enforces a minimum character threshold. The verify_pdf function (lines 16‑21) counts non-whitespace characters in the normalized text and compares this against a configurable minimum (defaulting to 1). This check catches documents where the text layer is technically present but effectively empty.
6. Required-Text Presence
Finally, the system verifies that specific keywords or phrases appear in the extracted content. The verify_pdf function (lines 23‑27) checks that all strings supplied via the --contains argument exist within the normalized text. This ensures critical sections—such as contact information, skills, or job titles—were successfully embedded in the PDF rather than rendered as unsearchable graphics.
The Verification Pipeline Flow
The parseability checks execute in a strict four-step sequence:
- Extract the text layer via
extract_text_layer, which orchestrates the pypdf-to-pdftotext fallback chain. - Optionally dump raw text to a specified file path when
--dump-textis provided, enabling manual inspection of what the ATS will actually see. - Validate extracted data against configured thresholds for pages, character counts, and required strings.
- Return metadata including the extractor name (
"pypdf"or"pdftotext"), the normalized text content, and the page count upon successful verification.
If any validation step fails, the pipeline immediately raises a VerificationError with a descriptive message indicating which parseability condition was not satisfied.
Command-Line Interface for ATS Verification
You can run the complete verification suite directly from the terminal to validate a résumé before submission:
# Verify that résumé.pdf has exactly 2 pages, at least 100 characters,
# and contains the phrase "Software Engineer"
python -m tools.verify_pdf résumé.pdf \
--pages 2 \
--min-chars 100 \
--contains "Software Engineer" \
--dump-text extracted.txt
This command executes all parseability checks and writes the extracted text to extracted.txt for manual review if the verification passes.
Programmatic ATS Verification in Python
For integration into automated workflows, import the verify_pdf function directly from tools.verify_pdf:
from tools.verify_pdf import verify_pdf, VerificationError
pdf_path = "résumé.pdf"
try:
extractor, text, pages = verify_pdf(
pdf_path,
expected_pages=2,
min_chars=100,
required_text=("Software Engineer", "Python"),
dump_text="extracted.txt",
)
print(f"✅ PDF is ATS‑parseable (extracted via {extractor}, {pages} pages).")
except VerificationError as exc:
print(f"❌ ATS verification failed: {exc}")
The function returns a tuple containing the extraction method used, the normalized text string, and the integer page count, allowing downstream processes to log exactly how the content was parsed.
Unit Testing Parseability Checks
The repository includes comprehensive tests in tests/test_verify_pdf.py that validate each parseability condition. Here is an example demonstrating the minimum character check:
@patch("tools.verify_pdf._extract_pypdf", return_value=("Hello ATS body", 1))
def test_min_chars(self, mock_extract):
# The extracted text has enough characters, so verification passes
extractor, text, pages = verify_pdf("dummy.pdf", min_chars=5)
self.assertEqual(extractor, "pypdf")
self.assertIn("Hello ATS body", text)
These tests ensure that the fallback logic, text normalization, and validation thresholds behave correctly across different PDF structures.
Summary
- Six core checks define parseability in ATS verification: page-count extraction, text-layer extraction (with pypdf/pdftotext fallback), text normalization, exact page validation, minimum character thresholds, and required-string verification.
tools/verify_pdf.pyimplements the complete pipeline, raisingVerificationErrorimmediately upon detecting any ATS-incompatible condition.- Dual extraction methods maximize compatibility: the system prefers
pypdfbut automatically falls back topdftotextwhen necessary. - Configurable thresholds allow customization of minimum character counts and required content strings via CLI arguments or Python parameters.
- Comprehensive testing in
tests/test_verify_pdf.pyensures reliable behavior across different PDF generation scenarios.
Frequently Asked Questions
What happens if a PDF fails the parseability checks?
The verify_pdf function raises a VerificationError with a specific message indicating which check failed—whether it was a page-count mismatch, insufficient characters, or missing required text. This allows the calling code to handle the failure appropriately, such as regenerating the PDF or alerting the user to fix the source document.
Which text extraction method does the tool prefer?
According to the source code in tools/verify_pdf.py, the system first attempts extraction via _extract_pypdf (lines 54‑69) using the pypdf library. Only if this returns empty content does it fall back to _extract_pdftotext (lines 72‑77), which uses Poppler’s pdftotext command-line tool. The returned extractor name indicates which method successfully parsed the document.
Can I customize the minimum character count for ATS verification?
Yes. The min_chars parameter in the verify_pdf function (lines 16‑21) accepts any integer value to override the default threshold of 1. When using the CLI, pass --min-chars followed by your desired number to ensure the résumé contains sufficient content for ATS parsing.
How do I debug text extraction issues?
Use the --dump-text CLI argument or the dump_text Python parameter to write the extracted and normalized text to a file. This allows you to inspect exactly what content the ATS will see, helping identify whether issues stem from image-based PDFs, encoding problems, or missing text layers before they cause application rejections.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →