ATS Text Layer Verification: 4 Essential Checks in ai-job-search

The ATS Text Layer Verification process in the ai-job-search repository validates CV PDFs through four critical checks—literal contact details, clean text extraction, logical reading order, and exact keyword matching—ensuring compatibility with Applicant Tracking Systems that parse embedded text layers rather than rendered pages.

The ai-job-search repository automates job application workflows through a sophisticated /apply command that includes rigorous ATS Text Layer Verification. This process ensures that generated CV PDFs contain machine-readable text layers that ATS software can accurately parse and index. Performing these checks during the document compilation phase prevents automatic rejections caused by formatting issues that hide critical candidate information from HR parsing systems.

The Four Core Checks of ATS Text Layer Verification

1. Contact Details as Literal Text

This check confirms that email addresses and phone numbers appear as plain text in the PDF's embedded layer rather than as icons or hyperlinks. The verification logic in tools/verify_pdf.py specifically flags FontAwesome glyphs and linked contact data, which extract as empty strings or metadata rather than readable strings. If contact data fails this check, the ATS cannot index the candidate's coordinates, effectively rendering the application invisible to recruiters regardless of qualifications.

2. No Garbled Output Detection

The system scans extracted text for replacement characters (``) and CID markers ((cid:NNN)), which indicate missing Unicode mappings in embedded fonts. These artifacts signal that the PDF contains subsetted fonts without proper ToUnicode tables, causing ATS parsers to receive corrupted gibberish instead of searchable content. This validation ensures every character maps correctly, preserving the semantic integrity of the CV in downstream HR databases.

3. Correct Reading Order Verification

This validation ensures the extraction order matches the visual single-column flow of the document. Multi-column or sidebar layouts often cause extractors to interleave unrelated lines—for example, mixing sidebar skills with main content experience—scrambling the logical narrative. The check verifies that section headers precede their content and employment history flows chronologically, maintaining the semantic structure that ATS algorithms use for keyword context scoring.

4. Keyword Coverage Analysis

The verification compares the job posting's required and preferred terms against the extracted text, weighing exact matches more heavily than synonyms. Because ATS filters typically operate on literal string matching rather than natural language processing, the check calculates coverage percentages using the posting's own terminology. Maximizing exact keyword matches in the verification output directly correlates with higher relevance scores in automated filtering stages.

Implementation in tools/verify_pdf.py

The verification engine resides in tools/verify_pdf.py, which orchestrates text extraction through a resilient dual-engine strategy. The script first attempts extraction using the BSD-licensed pypdf library; if unavailable, it gracefully falls back to the pdftotext utility from Poppler. The verify_pdf function accepts path parameters and regex patterns for contact validation, returning a results dictionary containing extraction metadata and check outcomes.

If both extractors are missing, the process degrades gracefully to a visual keyword review mode, as documented in README.md (lines 68-78), ensuring the workflow never hard-fails due to missing dependencies.

from tools.verify_pdf import verify_pdf

# verify_pdf returns a dict with the checks' outcomes

result = verify_pdf(
    pdf_path="cv/main_company_role.pdf",
    contact_regex=r"\b[\w\.-]+@[\w\.-]+\b",   # email pattern

    phone_regex=r"\+?\d{2,3}[\s-]?\d{6,10}"    # phone pattern

)

print("Contact details OK:", result["contact_ok"])
print("No garbled text:", result["clean_text"])
print("Reading order OK:", result["reading_order_ok"])
print("Keyword coverage:", result["keyword_match_percent"], "%")

Integration in the /apply Workflow

The ATS Text Layer Verification constitutes step 5d of the /apply pipeline, defined in .claude/commands/apply.md (lines 257-282). The verification executes in four sequential stages:

  1. Compile CV: LaTeX templates generate the PDF document from <CV_EXT> source files.
  2. Extract Text Layer: tools/verify_pdf.py executes the extraction pipeline, preferring pypdf and falling back to pdftotext.
  3. Run Verification: The system examines the extracted string against the four core checks, scanning for CID markers and validating contact regex matches.
  4. Report Results: Specific failures surface to the user with remediation guidance, referencing template adjustments documented in .claude/skills/job-application-assistant/05-cv-templates.md.

When verification fails, users must edit the LaTeX source or select a different layout, then re-run steps 5a-5c before re-executing the ATS verification step.


# Extract and dump the PDF text layer for manual inspection

python tools/verify_pdf.py cv/main_company_role.pdf --dump-text cv/main_company_role.txt

Summary

  • Four validation criteria: The process verifies literal contact text, absence of Unicode replacement characters and CID markers, single-column reading order preservation, and exact keyword matching against job postings.
  • Dual extraction engines: The implementation leverages pypdf primarily, with automatic fallback to Poppler's pdftotext utility to ensure consistent text extraction across diverse system environments.
  • Workflow integration: Step 5d of the /apply command validates PDFs immediately after compilation, preventing ATS parsing failures that could filter out qualified candidates before human review.
  • Graceful degradation: Missing dependencies trigger a visual keyword review rather than hard failures, maintaining workflow continuity as specified in the repository README.

Frequently Asked Questions

What causes the "garbled output" check to fail?

The garbled output check fails when extracted text contains replacement characters (``) or CID markers formatted as (cid:NNN). These artifacts indicate that the PDF's embedded fonts lack proper Unicode ToUnicode tables, causing ATS systems to index unreadable character sequences instead of the candidate's actual content.

Why does ATS Text Layer Verification require literal text for contact details instead of icons?

Applicant Tracking Systems parse only the embedded text layer of PDFs, ignoring rendered glyphs and visual styling. When contact information appears only as icons—such as FontAwesome glyphs—or within hyperlink metadata, the extraction yields empty strings rather than email addresses or phone numbers, making the candidate unreachable for interview scheduling.

How does the verification handle multi-column CV layouts?

The reading order check validates that text extracts in a logical single-column flow matching the human reading pattern. Multi-column designs can cause extractors to interleave text from adjacent columns, resulting in scrambled narratives that reduce keyword relevance scores and confuse semantic parsing algorithms used by modern ATS platforms.

What happens if both pypdf and pdftotext are unavailable on the system?

If neither extraction library is present, the ATS Text Layer Verification degrades gracefully to a visual keyword review mode. As documented in README.md (lines 68-78), this fallback allows the /apply workflow to continue using alternative verification methods rather than terminating with a dependency error.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →