ATS Verification Steps in the AI-Job-Search Framework: Technical Implementation Guide

The AI-Job-Search framework guarantees mechanical ATS readability through a rigorous three-phase verification pipeline that extracts text layers, enforces parseability constraints, and validates keyword coverage before any job submission.

The MadsLorentzen/ai-job-search repository automates job applications with built-in ATS verification steps that prevent visually perfect PDFs from failing machine parsing. These steps validate that compiled CVs and cover letters contain extractable text, meet structural requirements, and include critical keywords from job postings.

The Three-Phase ATS Verification Pipeline

The framework implements mechanical readability checks in tools/verify_pdf.py through a tightly-coupled sequence that fails fast on any ATS compatibility issue.

Phase 1: Text Layer Extraction with Dual-Engine Fallback

The extract_text_layer() function (lines 81-88) implements a resilient extraction strategy using pypdf as the primary engine. If pypdf is unavailable, raises an error, or returns an empty string, the tool automatically falls back to the Poppler utility pdftotext with flags -layout -enc UTF-8. The function returns the extractor name (pypdf or pdftotext) for audit reporting, ensuring you know exactly how the text was retrieved.

Phase 2: Mechanical Parseability Validation

Following extraction, the verify_pdf() function (lines 90-128) conducts rigorous structural validation through configurable constraints:

  • File existence: Verifies the PDF path resolves to an actual file
  • Page count enforcement: Validates exact page counts when --pages is specified
  • Character density checks: Ensures extracted text exceeds the --min-chars threshold to catch image-only PDFs
  • Content verification: Confirms presence of required strings via --contains

Any failed condition raises a VerificationError immediately, halting the workflow before a broken document reaches recruiters.

Phase 3: Debugging Artifacts and CLI Reporting

The verification pipeline supports forensic analysis through the --dump-text parameter. When provided, the extracted text writes to the specified path before any validation error raises, guaranteeing a traceable artifact for debugging glyph corruption or encoding issues. Upon success, the CLI prints a concise confirmation line including the extractor used and total page count.

Workflow Integration and Keyword Coverage

Apply Command Orchestration

In .claude/commands/apply.md, step 5d ("ATS & keyword verification") orchestrates the verification flow. The framework invokes the verification CLI automatically during the Apply workflow, records the extractor name in the final Step-6 report, and falls back to visual keyword-coverage review with an explicit degraded-mode notation if both extraction engines are unavailable.

Keyword Coverage Analysis

Post-extraction, the framework scans the text layer against required and desired keywords from the job posting. The analysis reports coverage in four categories: covered, synonym-only, missing-have-it, and missing-gap. This ensures the ATS will actually index the keywords your CV advertises, not just display them visually.

Practical Implementation Examples

CLI-Based Verification


# Verify a compiled CV with strict constraints and debug output

python tools/verify_pdf.py cv/main_acme_engineer.pdf \
    --pages 2 \
    --min-chars 100 \
    --contains "Professional Experience" \
    --dump-text cv/main_acme_engineer.txt

The command prints Verified cv/main_acme_engineer.pdf (extractor: pypdf, pages: 2) or raises a clear VerificationError detailing which constraint failed.

Programmatic Python API

from tools.verify_pdf import verify_pdf, VerificationError

pdf_path = "cv/main_acme_engineer.pdf"
try:
    extractor, text, pages = verify_pdf(
        pdf_path,
        expected_pages=2,
        min_chars=100,
        required_text=("Professional Experience", "Education"),
        dump_text="cv/main_acme_engineer.txt",
    )
    print(f"✓ ATS-ready ({extractor}, {pages} pages)")
except VerificationError as err:
    print(f"✗ ATS verification failed: {err}")

CI/CD Pipeline Integration


# In .github/workflows/ci.yml

python3 tools/verify_pdf.py cv/main_example.pdf --min-chars 100
python3 tools/verify_pdf.py cover_letters/cover_example.pdf --min-chars 100

The workflow aborts if either PDF fails ATS verification, preventing broken submissions from reaching applicant tracking systems.

Summary

  • Three-phase validation: Text extraction, parseability checks, and keyword coverage analysis ensure mechanical ATS compatibility
  • Dual-engine extraction: Primary pypdf with pdftotext fallback prevents single-point-of-failure during text layer retrieval
  • Configurable constraints: Exact page counts (--pages), minimum character thresholds (--min-chars), and required string presence (--contains) adapt to specific employer requirements
  • Forensic debugging: The --dump-text parameter generates inspectable artifacts to resolve encoding or glyph issues
  • Workflow integration: Automatic invocation during step 5d of the Apply command with extractor audit trails in final reports

Frequently Asked Questions

What happens if both pypdf and pdftotext are unavailable?

According to .claude/commands/apply.md, the framework falls back to a visual keyword-coverage review and explicitly notes the degraded mode in the final Step-6 report, though mechanical verification is bypassed.

How does the framework handle PDFs containing only images?

If extract_text_layer() returns an empty string—common with image-only or scanned PDFs—the verification fails the --min-chars check and raises a VerificationError before any submission occurs, preventing invisible CVs.

Can I verify cover letters using the same ATS verification steps?

Yes. The framework applies identical verification logic to both CVs and cover letters, as demonstrated in the CI pipeline examples and the Apply workflow documentation, using the same verify_pdf() API and CLI interface.

Where are the unit tests for ATS verification located?

The comprehensive test suite resides in tests/test_verify_pdf.py, covering scenarios including missing PDFs, insufficient character counts, required text validation, and fallback extraction handling to ensure the verification pipeline remains robust across environments.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →