# Understanding Parseability Checks in ATS Verification: A Technical Deep Dive

> Explore ATS verification parseability checks in ai-job-search. Learn how PDFs ensure machine-readable text, accurate page counts, and essential content for systems.

- Repository: [Mads Lorentzen/ai-job-search](https://github.com/MadsLorentzen/ai-job-search)
- Tags: deep-dive
- Published: 2026-08-28

---

**The `ai-job-search` repository performs six systematic parseability checks in ATS verification through [`tools/verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/verify_pdf.py), ensuring that generated résumé PDFs contain machine-readable text layers, correct page counts, and required content before submission to Applicant Tracking Systems.**

The MadsLorentzen/ai-job-search toolkit automates the creation of ATS-friendly résumés, but creating a PDF is only half the battle. These parseability checks in ATS verification confirm that the document can be successfully ingested by employer parsing software, preventing your application from being rejected due to unreadable formatting or missing text layers.

## What Are Parseability Checks in ATS Verification?

Parseability checks are automated validations that verify a PDF file's internal structure and content accessibility. Unlike visual formatting checks, these inspect the document's **text layer**, **metadata**, and **character encoding** to ensure ATS software can extract meaningful data. In the `ai-job-search` codebase, these checks are implemented as a pipeline that raises a `VerificationError` immediately upon detecting any condition that would cause an ATS to fail reading the résumé.

## The Six Parseability Checks Explained

The verification logic in [`tools/verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/verify_pdf.py) implements six distinct checks that execute sequentially during the validation process.

### 1. Page-Count Extraction

Before analyzing content, the system verifies the document's physical structure. The `parse_page_count` function (lines 44‑48) retrieves the total number of pages using either the `pdfinfo` command-line tool or the page count metadata from `pypdf`. This establishes a baseline for subsequent structural validations.

### 2. ATS-Readable Text Extraction

The most critical check ensures the PDF contains an extractable text layer rather than image-only content. The system attempts extraction via `_extract_pypdf` (lines 54‑69) using the `pypdf` library first. If this returns empty or unavailable content, it automatically falls back to `_extract_pdftotext` (lines 72‑77), which invokes Poppler’s `pdftotext` utility. This dual-method approach maximizes compatibility across different PDF generation methods.

### 3. Text Normalisation

Raw extracted text often contains irregular whitespace that can interfere with string matching. The `normalize_text` function (lines 50‑52) collapses multiple whitespace characters into single spaces, ensuring that subsequent checks for required content are independent of formatting variations or line-break inconsistencies.

### 4. Exact Page-Count Validation

When the caller supplies an expected page count via the `--pages` argument, the `verify_pdf` function (lines 11‑15) performs an exact equality check against the extracted page count. Any discrepancy triggers a `VerificationError`, preventing submission of résumés that exceed single-page limits or fail to meet minimum length requirements.

### 5. Minimum Character Count

To prevent submission of scanned images or corrupted PDFs that contain no readable text, the system enforces a **minimum character threshold**. The `verify_pdf` function (lines 16‑21) counts non-whitespace characters in the normalized text and compares this against a configurable minimum (defaulting to 1). This check catches documents where the text layer is technically present but effectively empty.

### 6. Required-Text Presence

Finally, the system verifies that specific keywords or phrases appear in the extracted content. The `verify_pdf` function (lines 23‑27) checks that all strings supplied via the `--contains` argument exist within the normalized text. This ensures critical sections—such as contact information, skills, or job titles—were successfully embedded in the PDF rather than rendered as unsearchable graphics.

## The Verification Pipeline Flow

The parseability checks execute in a strict four-step sequence:

1. **Extract the text layer** via `extract_text_layer`, which orchestrates the pypdf-to-pdftotext fallback chain.
2. **Optionally dump raw text** to a specified file path when `--dump-text` is provided, enabling manual inspection of what the ATS will actually see.
3. **Validate extracted data** against configured thresholds for pages, character counts, and required strings.
4. **Return metadata** including the extractor name (`"pypdf"` or `"pdftotext"`), the normalized text content, and the page count upon successful verification.

If any validation step fails, the pipeline immediately raises a `VerificationError` with a descriptive message indicating which parseability condition was not satisfied.

## Command-Line Interface for ATS Verification

You can run the complete verification suite directly from the terminal to validate a résumé before submission:

```bash

# Verify that résumé.pdf has exactly 2 pages, at least 100 characters,

# and contains the phrase "Software Engineer"

python -m tools.verify_pdf résumé.pdf \
    --pages 2 \
    --min-chars 100 \
    --contains "Software Engineer" \
    --dump-text extracted.txt

```

This command executes all parseability checks and writes the extracted text to [`extracted.txt`](https://github.com/MadsLorentzen/ai-job-search/blob/main/extracted.txt) for manual review if the verification passes.

## Programmatic ATS Verification in Python

For integration into automated workflows, import the `verify_pdf` function directly from `tools.verify_pdf`:

```python
from tools.verify_pdf import verify_pdf, VerificationError

pdf_path = "résumé.pdf"
try:
    extractor, text, pages = verify_pdf(
        pdf_path,
        expected_pages=2,
        min_chars=100,
        required_text=("Software Engineer", "Python"),
        dump_text="extracted.txt",
    )
    print(f"✅ PDF is ATS‑parseable (extracted via {extractor}, {pages} pages).")
except VerificationError as exc:
    print(f"❌ ATS verification failed: {exc}")

```

The function returns a tuple containing the extraction method used, the normalized text string, and the integer page count, allowing downstream processes to log exactly how the content was parsed.

## Unit Testing Parseability Checks

The repository includes comprehensive tests in [`tests/test_verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tests/test_verify_pdf.py) that validate each parseability condition. Here is an example demonstrating the minimum character check:

```python
@patch("tools.verify_pdf._extract_pypdf", return_value=("Hello ATS body", 1))
def test_min_chars(self, mock_extract):
    # The extracted text has enough characters, so verification passes

    extractor, text, pages = verify_pdf("dummy.pdf", min_chars=5)
    self.assertEqual(extractor, "pypdf")
    self.assertIn("Hello ATS body", text)

```

These tests ensure that the fallback logic, text normalization, and validation thresholds behave correctly across different PDF structures.

## Summary

- **Six core checks** define parseability in ATS verification: page-count extraction, text-layer extraction (with pypdf/pdftotext fallback), text normalization, exact page validation, minimum character thresholds, and required-string verification.
- **[`tools/verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/verify_pdf.py)** implements the complete pipeline, raising `VerificationError` immediately upon detecting any ATS-incompatible condition.
- **Dual extraction methods** maximize compatibility: the system prefers `pypdf` but automatically falls back to `pdftotext` when necessary.
- **Configurable thresholds** allow customization of minimum character counts and required content strings via CLI arguments or Python parameters.
- **Comprehensive testing** in [`tests/test_verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tests/test_verify_pdf.py) ensures reliable behavior across different PDF generation scenarios.

## Frequently Asked Questions

### What happens if a PDF fails the parseability checks?

The `verify_pdf` function raises a `VerificationError` with a specific message indicating which check failed—whether it was a page-count mismatch, insufficient characters, or missing required text. This allows the calling code to handle the failure appropriately, such as regenerating the PDF or alerting the user to fix the source document.

### Which text extraction method does the tool prefer?

According to the source code in [`tools/verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/verify_pdf.py), the system first attempts extraction via `_extract_pypdf` (lines 54‑69) using the `pypdf` library. Only if this returns empty content does it fall back to `_extract_pdftotext` (lines 72‑77), which uses Poppler’s `pdftotext` command-line tool. The returned extractor name indicates which method successfully parsed the document.

### Can I customize the minimum character count for ATS verification?

Yes. The `min_chars` parameter in the `verify_pdf` function (lines 16‑21) accepts any integer value to override the default threshold of 1. When using the CLI, pass `--min-chars` followed by your desired number to ensure the résumé contains sufficient content for ATS parsing.

### How do I debug text extraction issues?

Use the `--dump-text` CLI argument or the `dump_text` Python parameter to write the extracted and normalized text to a file. This allows you to inspect exactly what content the ATS will see, helping identify whether issues stem from image-based PDFs, encoding problems, or missing text layers before they cause application rejections.