How the PDF Verification Loop Works in the AI Job Search /apply Step
The PDF verification loop in the /apply step validates CV documents through an iterative check system that enforces page count, character minimums, and required text phrases before allowing the job application to proceed.
The /apply command in the MadsLorentzen/ai-job-search repository automates job applications through a Markdown-driven pipeline. Before recording any application, the system runs a mandatory PDF verification loop to ensure CV integrity. This loop, defined in the command specification and implemented in Python validation scripts, prevents corrupted or incomplete documents from entering the application tracker.
Where the PDF Verification Loop Is Defined
The verification loop spans three critical files in the repository architecture:
.claude/commands/apply.md– The Markdown command specification that defines the/applyworkflow and orchestrates the verification steptools/verify_pdf.py– The Python implementation containing theverify_pdffunction that performs document extraction and validationtools/security_guards.py– The security configuration that registers the PDF verifier as a mandatory guard for the application process
How the PDF Verification Loop Validates Documents
The loop operates as a repeat-until-valid control structure that extracts PDF content and validates it against configurable constraints.
Text Extraction and Parsing
When invoked, the verify_pdf function in tools/verify_pdf.py first extracts raw text from the CV using the _extract_pypdf helper:
extractor, text, pages = _extract_pypdf(pdf_path)
This extraction uses PyPDF2 or a fallback parser to return the extractor object, concatenated text content, and total page count from the PDF file.
Validation Guards and Constraints
The verification loop applies three primary validation guards against the extracted content:
- Page Count Validation – If
expected_pagesis specified, the document must contain at least that many pages - Character Minimum Check – The
min_charsparameter (defaulting to 1) ensures the PDF is not empty or trivially short - Required Text Verification – The
required_texttuple specifies mandatory phrases (e.g., "Professional Experience") that must appear in the extracted text
Each failed guard appends a descriptive error message to an accumulator list.
Error Handling and Retry Logic
If any validation fails, the function streams error messages to stderr and returns None:
if errors:
for msg in errors:
print(msg, file=sys.stderr)
return None
return extractor, text, pages
The calling loop in apply.md detects this failure and prompts the user to correct the PDF before retrying, creating the iterative verification cycle.
Core Implementation: The verify_pdf Function
The validation logic resides in tools/verify_pdf.py and follows this signature:
def verify_pdf(pdf_path,
expected_pages: int | None = None,
min_chars: int = 1,
required_text: tuple[str, ...] = (),
dump_text: Path | None = None):
extractor, text, pages = _extract_pypdf(pdf_path)
errors = []
if expected_pages and pages < expected_pages:
errors.append(f"expected ≥{expected_pages} pages, got {pages}")
if len(text) < min_chars:
errors.append(f"expected ≥{min_chars} characters, got {len(text)}")
for phrase in required_text:
if phrase not in text:
errors.append(f"missing required phrase: {phrase!r}")
if errors:
for msg in errors:
print(msg, file=sys.stderr)
return None
return extractor, text, pages
This function acts as the gatekeeper: returning the extraction data only when all constraints are satisfied, or blocking progression by returning None and logging specific failure reasons.
Security Guard Integration
The verification loop is enforced through the security guards system defined in tools/security_guards.py. The PDF verifier is registered as a mandatory pre-condition:
SECURITY_GUARDS = [
"Bash(python tools/verify_pdf.py:*)",
# … other guards …
]
This registration ensures that verify_pdf.py executes automatically before the /apply step writes any tracker entries, guaranteeing that every recorded application references a validated CV document.
The Loop Control Structure in apply.md
The command specification implements the iterative loop using bash control flow within the Markdown command definition:
while true; do
result=$(python tools/verify_pdf.py "$CV_PATH" \
--expected-pages 2 \
--min-chars 500 \
--required-text "Professional Experience")
if [ $? -eq 0 ]; then
break
fi
echo "PDF verification failed – see above. Please fix the file and retry."
read -p "Press <Enter> after fixing the PDF..."
done
This structure creates a blocking loop that pauses the application process until the PDF passes all validation checks, providing immediate feedback to the user about document quality issues.
Summary
- The PDF verification loop is a mandatory gate in the
/applyworkflow that validates CV documents before application recording - Validation occurs in
tools/verify_pdf.py, where theverify_pdffunction checks page counts, character minimums, and required text phrases - The loop follows a repeat-until-valid pattern defined in
.claude/commands/apply.md, prompting users to fix documents that fail validation - Security guards in
tools/security_guards.pyenforce this verification as a prerequisite for the application process - Failed validations return
Noneand print diagnostic errors tostderr, while successful validations return extraction data and allow workflow progression
Frequently Asked Questions
What happens if my PDF fails the verification loop?
If the PDF fails validation, the verify_pdf function prints specific error messages to stderr indicating which constraints failed—such as insufficient pages, low character counts, or missing required phrases. The /apply command then traps these errors in a while loop, displays them to the user, and pauses execution until the user confirms they have fixed the document and presses Enter to retry.
Which validation parameters can I configure in the PDF verification loop?
According to the verify_pdf function signature in tools/verify_pdf.py, you can configure three main parameters. These include expected_pages for minimum page count, min_chars for minimum character count (defaulting to 1), and required_text for mandatory phrases. You pass these values as command-line arguments when the verification script is invoked within the /apply step.
Where does the PDF verification loop fit in the overall /apply workflow?
The loop executes during Step 0: Parse Input of the /apply command, immediately after the CV path is identified. Because tools/security_guards.py registers the verifier as a mandatory guard, this validation occurs before any application data is written to the tracker, ensuring data integrity at the entry point of the workflow.
How does the system extract text from PDFs for validation?
The verify_pdf function relies on an internal _extract_pypdf helper that uses PyPDF2 (with fallback mechanisms) to parse the PDF binary. This extractor returns the raw text content, page count, and an extractor object, which the verification loop then analyzes against the configured constraints.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →