How ATS Parseability Is Verified for Generated Documents in ai-job-search

ATS parseability is verified by extracting the text layer from generated PDF résumés and cover letters, then validating against configurable rules for character count, page count, and required content phrases.

The MadsLorentzen/ai-job-search repository automates job applications with AI-generated documents, but those documents are worthless if Applicant Tracking Systems cannot read them. The codebase includes a dedicated verification pipeline in tools/verify_pdf.py that ensures every generated PDF contains a clean, machine-readable text layer before submission.

PDF Text Extraction Strategy

The verification system uses a dual-backend extraction strategy to maximize reliability across different environments.

In tools/verify_pdf.py, the _extract_text function attempts extraction in this order:

  1. pypdf (_extract_pypdf) — The pure-Python library is tried first for portability
  2. pdftotext (_extract_pdftotext) — Falls back to the Poppler command-line utility if pypdf is unavailable, fails, or returns empty content

The function returns the successful extractor name, pages detected, and extracted text:


# From tools/verify_pdf.py lines 54-86

def _extract_text(pdf_path: str | Path) -> tuple[str, int, str]:
    """Extract text from PDF. Returns (extractor_name, num_pages, text)."""
    for extractor_func, name in (
        (_extract_pypdf, "pypdf"),
        (_extract_pdftotext, "pdftotext"),
    ):
        try:
            text, pages = extractor_func(pdf_path)
            if text.strip():  # Ensure we got actual content

                return name, pages, text
        except Exception:
            continue
    raise VerificationError(f"Could not extract text from {pdf_path}")

This fallback design ensures ATS parseability checks work whether the environment has only Python dependencies or requires system tools.

ATS-Readable Text Validation Rules

Once extracted, the text undergoes normalization and multi-layer validation in verify_pdf():

Check Parameter Default Purpose
Minimum characters min_chars 1 Confirms PDF isn't image-only or corrupted
Exact page count expected_pages None Validates document length expectations
Required phrases required_text () Ensures critical sections (contact info, experience headers) are present

The validation implementation (lines 16-27 of tools/verify_pdf.py) normalizes whitespace then applies each rule:

def verify_pdf(
    pdf_path: str | Path,
    *,
    expected_pages: int | None = None,
    min_chars: int = 1,
    required_text: tuple[str, ...] = (),
) -> tuple[str, int, str]:
    """Verify PDF is ATS-parseable. Returns (extractor, pages, normalized_text)."""
    extractor, pages, raw_text = _extract_text(pdf_path)
    text = normalize_text(raw_text)
    
    if expected_pages is not None and pages != expected_pages:
        raise VerificationError(
            f"Expected {expected_pages} pages, got {pages} (extractor: {extractor})"
        )
    if len(text) < min_chars:
        raise VerificationError(
            f"Text too short: {len(text)} chars < {min_chars} (extractor: {extractor})"
        )
    for phrase in required_text:
        if phrase not in text:
            raise VerificationError(
                f"Required text not found: {phrase!r} (extractor: {extractor})"
            )
    return extractor, pages, text

All failures include the extractor used, enabling quick diagnosis of whether pypdf or pdftotext produced problematic results.

Integration in the Application Workflow

The ATS parseability check executes during step 5d of the /apply workflow, as documented in tools/security_guards.py (lines 60-64). This security guard context confirms that:

  • Generated CVs pass verification before submission
  • The extracted text layer is persisted as cv/*.txt files for later inspection
  • Failures block the application pipeline

This integration prevents visually polished but machine-unreadable documents from reaching employers.

Command-Line and Programmatic Usage

Command-Line Interface

The verify_pdf.py module exposes a full CLI for manual checks and CI/CD pipelines:

python tools/verify_pdf.py \
    example.pdf \
    --pages 2 \
    --min-chars 200 \
    --contains "Professional Experience" \
    --contains "[your.email@example.com]" \
    --dump-text extracted.txt
  • --pages — Enforce exact page count
  • --min-chars — Set minimum character threshold
  • --contains — Add required phrases (repeatable)
  • --dump-text — Save normalized text to file for inspection

Python API

For programmatic integration, import verify_pdf and handle VerificationError:

from tools.verify_pdf import verify_pdf, VerificationError
from pathlib import Path

pdf_path = Path("cv/main_example.pdf")

try:
    extractor, text, pages = verify_pdf(
        pdf_path,
        expected_pages=2,
        min_chars=30,
        required_text=("Professional Experience", "your.email@example.com"),
    )
    print(f"ATS-parseable with {extractor}; {pages} pages, {len(text)} chars")
except VerificationError as exc:
    print(f"Verification failed: {exc}")
    # Block submission, alert user, or trigger regeneration

Test Coverage

The test suite in tests/test_verify_pdf.py validates the verification system itself across lines 34-94:

  • Page count parsing — Confirms correct detection regardless of extractor used
  • Missing file handling — Ensures clear errors for absent PDFs
  • Fallback chain — Verifies pypdf → pdftotext progression when pypdf fails
  • Validation thresholds — Tests min_chars and required_text enforcement

This coverage guarantees that the ATS parseability verification remains reliable as the codebase evolves.

Summary

  • Dual extraction (pypdf → pdftotext) ensures text layer access across environments
  • Configurable validation (pages, characters, required phrases) adapts to document requirements
  • Clear diagnostics include extractor name in all error messages
  • Workflow integration at /apply step 5d blocks unparseable submissions
  • Persistent text dumps enable forensic inspection of extracted content

Frequently Asked Questions

What makes a PDF "ATS-parseable" according to this system?

An ATS-parseable PDF contains a searchable text layer that can be extracted programmatically. The system verifies this by extracting raw text, confirming it meets minimum length requirements, and validating that expected content (headers, contact info) appears in normalized form. Image-only PDFs or corrupted text layers fail verification.

Why does the system fall back from pypdf to pdftotext?

pypdf is preferred for being pure-Python and dependency-free, but it can fail on complex PDFs or return empty strings for image-heavy documents. pdftotext (Poppler) handles more PDF variants reliably. The fallback ensures verification works in minimal environments while maximizing accuracy where system tools are available.

How can I debug a failing ATS parseability check?

The VerificationError message includes which extractor succeeded or failed, making it easy to identify whether pypdf or pdftotext produced problematic results. Use --dump-text in CLI mode or capture the returned text in Python to inspect exactly what characters were extracted. Common fixes include regenerating the PDF with embedded fonts or adjusting min_chars/required_text for document-specific content.

Where is the extracted text stored for inspection?

During the /apply workflow, verified CV text is saved as .txt files in the cv/ directory alongside the PDF, as noted in tools/security_guards.py. These files persist after verification, allowing manual review of what ATS software will actually ingest versus the visual appearance of the PDF.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →