How the AI Job Search Framework Parses Job Postings from URLs and Text

The AI Job Search Framework parses job postings in Step 0 of the /apply workflow, automatically detecting whether the input is a URL or raw text, fetching remote content with resilient fallback strategies, and extracting essential metadata while treating the posting as untrusted data.

The MadsLorentzen/ai-job-search repository automates job application workflows through a structured command system. When you invoke the /apply command, the framework immediately executes its parsing logic defined in /.claude/commands/apply.md to process your input. This step determines whether the system fetches remote content or processes local text, establishing the data foundation for all subsequent automation including fit evaluation and document generation.

URL Detection vs. Plain Text Input

The parsing logic begins with deterministic input classification. In /.claude/commands/apply.md at lines 21-27, the framework evaluates the /apply argument to determine if it matches a URL pattern.

  • URL pattern matched: The system treats the argument as a remote job posting and initiates the WebFetch tool.
  • No URL pattern detected: The framework assumes the argument contains the complete job description text and proceeds directly to metadata extraction.

This automatic detection allows users to paste either a link or the full posting without specifying the input type manually.

Remote Content Fetching and Error Handling

When processing URLs, the framework implements a resilient escalation policy to handle common web scraping obstacles. The primary fetch mechanism uses the built-in WebFetch tool to download the page content.

If the fetch fails with a 403 Forbidden error or encounters login walls and irrelevant aggregator listings, the parser falls back to the escalation policy defined in /.claude/skills/job-application-assistant/09-web-research.md:

  1. Retry with browser-like headers via curl to bypass basic bot detection.
  2. Locate the employer’s own careers page when aggregator sites block automated access.

This ensures the framework can retrieve posting content even when job boards implement aggressive anti-scraping measures.

Authority Preference for Employer Sources

The parser prioritizes data integrity by preferring authoritative sources over third-party aggregators. When the system detects that both an aggregator (e.g., LinkedIn) and the employer’s direct careers page are available, it selects the employer’s posting.

According to lines 25-26 in /.claude/commands/apply.md, employer pages retain crucial details often stripped by aggregators, including:

  • Requisition IDs
  • Seniority grades
  • Department hierarchies

This preference ensures the downstream evaluation and drafting steps operate on the most complete and accurate metadata available.

Security Model and Untrusted Data Boundaries

The framework implements strict trust boundaries when handling posting content. As specified in /.claude/commands/apply.md at lines 27-28, the parser never treats the job posting as trusted input:

  • No hidden instruction execution: The system ignores any embedded instructions or prompts within the posting body.
  • No secondary URL fetching: The parser does not fetch URLs that appear inside the posting text (with the sole exception of the original user-supplied URL).
  • No content injection: Posting-derived content is never directly injected into generated CVs or cover letters.

These security rules propagate to every subsequent step of the workflow, ensuring that malicious or malformed postings cannot compromise the output documents.

Metadata Extraction and Caching Strategy

Once the raw content is secured, the parser extracts six essential fields defined at lines 28-30 in /.claude/commands/apply.md:

  • Company name
  • Role title
  • Department (if specified)
  • Location
  • Application deadline (if present)
  • Language of the posting (Danish or English)

The framework retains the full posting text verbatim for archiving and verification purposes. This raw text is stored in the command context and passed to /.claude/skills/job-application-assistant/04-job-evaluation.md for the fit assessment step. The unmodified posting is later archived during Step 6b of the workflow, creating an immutable record of the original job description.

Practical Usage Examples

The /apply command accepts input through both methods demonstrated below:


# Fetch and parse a remote job posting URL

/apply https://jobindex.dk/job/1234567

# Parse pasted job description text directly

/apply <<EOF
Senior Data Engineer – Acme Corp
Location: Copenhagen, Denmark
We seek a candidate with 5+ years of Python, Spark, and cloud experience...
EOF

In both cases, the framework automatically executes the parsing logic, applies the security trust boundaries, extracts the required metadata fields, and proceeds to the evaluation and drafting phases.

Summary

  • Step 0 parsing in /.claude/commands/apply.md automatically detects URL versus text input using pattern matching.
  • Resilient fetching leverages the WebFetch tool with escalation policies defined in 09-web-research.md to handle 403 errors and login walls.
  • Source authority prioritizes employer career pages over aggregators to preserve requisition IDs and seniority details.
  • Security boundaries treat all posting content as untrusted data, preventing hidden instruction execution and content injection.
  • Six core fields are extracted and cached in command context for downstream use by the evaluation skill and archived in Step 6b for compliance.

Frequently Asked Questions

How does the AI Job Search Framework handle blocked job board URLs?

When the WebFetch tool encounters a 403 error or login wall, the framework implements the escalation policy from the web-research skill. It first retries the request using curl with browser-like headers to bypass basic bot detection. If that fails, the system attempts to locate and fetch the job posting directly from the employer’s careers page rather than the aggregator site.

Can I use the framework with pasted job descriptions instead of URLs?

Yes. If the /apply argument does not match a URL pattern, the framework automatically treats the entire input as the job posting text. This allows you to paste content from emails, PDFs, or internal job boards directly into the command, bypassing the WebFetch step entirely while still extracting the same metadata fields.

Why does the parser prefer employer websites over LinkedIn or other aggregators?

The parser prioritizes employer sources because aggregator sites often strip critical metadata such as requisition IDs, hiring manager details, and precise seniority grades. According to the logic in /.claude/commands/apply.md, these fields are essential for accurate fit evaluation and tailoring application materials to the specific role requirements.

Where is the job posting data stored during the workflow execution?

Extracted metadata is stored in the command context and passed to subsequent steps including /.claude/skills/job-application-assistant/04-job-evaluation.md. The full raw posting text is retained verbatim and archived during Step 6b of the workflow, creating an immutable reference for later verification or auditing purposes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →