# How ATS Parseability Is Verified for Generated Documents in ai-job-search

> Verify ATS parseability for generated documents by extracting text layers and validating against configurable rules for content, character, and page counts.

- Repository: [Mads Lorentzen/ai-job-search](https://github.com/MadsLorentzen/ai-job-search)
- Tags: how-to-guide
- Published: 2026-08-30

---

**ATS parseability is verified by extracting the text layer from generated PDF résumés and cover letters, then validating against configurable rules for character count, page count, and required content phrases.**

The `MadsLorentzen/ai-job-search` repository automates job applications with AI-generated documents, but those documents are worthless if Applicant Tracking Systems cannot read them. The codebase includes a dedicated verification pipeline in [`tools/verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/verify_pdf.py) that ensures every generated PDF contains a clean, machine-readable text layer before submission.

## PDF Text Extraction Strategy

The verification system uses a **dual-backend extraction strategy** to maximize reliability across different environments.

In [`tools/verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/verify_pdf.py), the `_extract_text` function attempts extraction in this order:

1. **pypdf** (`_extract_pypdf`) — The pure-Python library is tried first for portability
2. **pdftotext** (`_extract_pdftotext`) — Falls back to the Poppler command-line utility if pypdf is unavailable, fails, or returns empty content

The function returns the successful extractor name, pages detected, and extracted text:

```python

# From tools/verify_pdf.py lines 54-86

def _extract_text(pdf_path: str | Path) -> tuple[str, int, str]:
    """Extract text from PDF. Returns (extractor_name, num_pages, text)."""
    for extractor_func, name in (
        (_extract_pypdf, "pypdf"),
        (_extract_pdftotext, "pdftotext"),
    ):
        try:
            text, pages = extractor_func(pdf_path)
            if text.strip():  # Ensure we got actual content

                return name, pages, text
        except Exception:
            continue
    raise VerificationError(f"Could not extract text from {pdf_path}")

```

This fallback design ensures ATS parseability checks work whether the environment has only Python dependencies or requires system tools.

## ATS-Readable Text Validation Rules

Once extracted, the text undergoes **normalization and multi-layer validation** in `verify_pdf()`:

| Check | Parameter | Default | Purpose |
|-------|-----------|---------|---------|
| Minimum characters | `min_chars` | 1 | Confirms PDF isn't image-only or corrupted |
| Exact page count | `expected_pages` | `None` | Validates document length expectations |
| Required phrases | `required_text` | `()` | Ensures critical sections (contact info, experience headers) are present |

The validation implementation (lines 16-27 of [`tools/verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/verify_pdf.py)) normalizes whitespace then applies each rule:

```python
def verify_pdf(
    pdf_path: str | Path,
    *,
    expected_pages: int | None = None,
    min_chars: int = 1,
    required_text: tuple[str, ...] = (),
) -> tuple[str, int, str]:
    """Verify PDF is ATS-parseable. Returns (extractor, pages, normalized_text)."""
    extractor, pages, raw_text = _extract_text(pdf_path)
    text = normalize_text(raw_text)
    
    if expected_pages is not None and pages != expected_pages:
        raise VerificationError(
            f"Expected {expected_pages} pages, got {pages} (extractor: {extractor})"
        )
    if len(text) < min_chars:
        raise VerificationError(
            f"Text too short: {len(text)} chars < {min_chars} (extractor: {extractor})"
        )
    for phrase in required_text:
        if phrase not in text:
            raise VerificationError(
                f"Required text not found: {phrase!r} (extractor: {extractor})"
            )
    return extractor, pages, text

```

All failures include the extractor used, enabling quick diagnosis of whether pypdf or pdftotext produced problematic results.

## Integration in the Application Workflow

The ATS parseability check executes during **step 5d of the `/apply` workflow**, as documented in [`tools/security_guards.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/security_guards.py) (lines 60-64). This security guard context confirms that:

- Generated CVs pass verification before submission
- The extracted text layer is persisted as `cv/*.txt` files for later inspection
- Failures block the application pipeline

This integration prevents visually polished but machine-unreadable documents from reaching employers.

## Command-Line and Programmatic Usage

### Command-Line Interface

The [`verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/verify_pdf.py) module exposes a full CLI for manual checks and CI/CD pipelines:

```bash
python tools/verify_pdf.py \
    example.pdf \
    --pages 2 \
    --min-chars 200 \
    --contains "Professional Experience" \
    --contains "[your.email@example.com]" \
    --dump-text extracted.txt

```

- `--pages` — Enforce exact page count
- `--min-chars` — Set minimum character threshold
- `--contains` — Add required phrases (repeatable)
- `--dump-text` — Save normalized text to file for inspection

### Python API

For programmatic integration, import `verify_pdf` and handle `VerificationError`:

```python
from tools.verify_pdf import verify_pdf, VerificationError
from pathlib import Path

pdf_path = Path("cv/main_example.pdf")

try:
    extractor, text, pages = verify_pdf(
        pdf_path,
        expected_pages=2,
        min_chars=30,
        required_text=("Professional Experience", "your.email@example.com"),
    )
    print(f"ATS-parseable with {extractor}; {pages} pages, {len(text)} chars")
except VerificationError as exc:
    print(f"Verification failed: {exc}")
    # Block submission, alert user, or trigger regeneration

```

## Test Coverage

The test suite in [`tests/test_verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tests/test_verify_pdf.py) validates the verification system itself across lines 34-94:

- **Page count parsing** — Confirms correct detection regardless of extractor used
- **Missing file handling** — Ensures clear errors for absent PDFs
- **Fallback chain** — Verifies pypdf → pdftotext progression when pypdf fails
- **Validation thresholds** — Tests `min_chars` and `required_text` enforcement

This coverage guarantees that the ATS parseability verification remains reliable as the codebase evolves.

## Summary

- **Dual extraction** (pypdf → pdftotext) ensures text layer access across environments
- **Configurable validation** (pages, characters, required phrases) adapts to document requirements
- **Clear diagnostics** include extractor name in all error messages
- **Workflow integration** at `/apply` step 5d blocks unparseable submissions
- **Persistent text dumps** enable forensic inspection of extracted content

## Frequently Asked Questions

### What makes a PDF "ATS-parseable" according to this system?

An ATS-parseable PDF contains a **searchable text layer** that can be extracted programmatically. The system verifies this by extracting raw text, confirming it meets minimum length requirements, and validating that expected content (headers, contact info) appears in normalized form. Image-only PDFs or corrupted text layers fail verification.

### Why does the system fall back from pypdf to pdftotext?

**pypdf** is preferred for being pure-Python and dependency-free, but it can fail on complex PDFs or return empty strings for image-heavy documents. **pdftotext** (Poppler) handles more PDF variants reliably. The fallback ensures verification works in minimal environments while maximizing accuracy where system tools are available.

### How can I debug a failing ATS parseability check?

The `VerificationError` message includes which **extractor succeeded or failed**, making it easy to identify whether pypdf or pdftotext produced problematic results. Use `--dump-text` in CLI mode or capture the returned text in Python to inspect exactly what characters were extracted. Common fixes include regenerating the PDF with embedded fonts or adjusting `min_chars`/`required_text` for document-specific content.

### Where is the extracted text stored for inspection?

During the `/apply` workflow, verified CV text is saved as **`.txt` files in the `cv/` directory** alongside the PDF, as noted in [`tools/security_guards.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/security_guards.py). These files persist after verification, allowing manual review of what ATS software will actually ingest versus the visual appearance of the PDF.