The Six Execution Steps of the /scrape Workflow in AI-Job-Search
The /scrape workflow executes six sequential steps—Load State, Search, Fetch & Parse, Quick Fit Assessment, Deduplicate & Store, and Present Results—to automate job discovery while avoiding duplicate postings.
The /scrape command is implemented as the job-scraper skill in the MadsLorentzen/ai-job-search repository. Defined in .claude/skills/job-scraper/SKILL.md, this deterministic pipeline aggregates postings from multiple job portals, evaluates candidate fit, and surfaces new opportunities without reprocessing historical data.
Overview of the /scrape Workflow
The workflow is designed for idempotent execution. Each run consults persistent state files to avoid redundant API calls and duplicate entries. The six core steps—numbered Step 0 through Step 5—handle everything from initialization to final presentation, with an optional Step 6 ("Update Tracker") available only when manually recording external applications.
Step-by-Step Breakdown
Step 0 – Load State
The scraper initializes its knowledge base by loading three critical data sources:
job_scraper/seen_jobs.json— A persistent cache of every job previously encountered. If the file is missing, the system initializes it with an empty{ "seen": {} }structure.job_search_tracker.csv— The application tracker used to identify companies and roles already applied to.search-queries.md— User-defined search strategies specifying which query categories to prioritize.
This initialization ensures that subsequent steps only process novel opportunities.
Step 1 – Search
This step executes the actual job-board queries according to the strategy defined in search-queries.md:
- Category Selection — By default, the system runs the top three query categories. Users can override this with
broad(all categories) or a focused category argument. - CLI Preference — The system prefers installed portal-specific CLI tools, falling back to WebSearch when a CLI is unavailable or when bun is missing from the environment.
- Parallel Execution — Sub-step 1b runs each enabled portal's CLI in parallel, respecting individual portal
enabledflags and recency settings.
If a portal fails or lacks CLI support, sub-step 1c triggers a WebSearch fallback to ensure coverage.
Step 2 – Fetch & Parse
For each promising result identified in Step 1, the system retrieves comprehensive posting details:
- Data Retrieval — Invokes the portal's
detailcommand for CLI-driven portals, or uses WebFetch for WebSearch-derived URLs. - Field Extraction — Captures
title,company,location,posting_date,deadline, a short description, and anisActiveflag (specifically for LinkedIn postings). - Duplicate Filtering — Immediately skips entries existing in
seen_jobs.jsonor the tracker CSV, preventing redundant processing.
Step 3 – Quick Fit Assessment
Each new job receives a lightweight classification based on predefined criteria:
- Match Rating — Jobs are categorized as high, medium, or low match according to skill overlap analysis.
- Language Gate — A hard filter overrides the rating if the posting requires a language the candidate has not declared proficiency in.
This ensures that only relevant opportunities progress to the final presentation stage.
Step 4 – Deduplicate & Store
Persistence occurs before presentation to maintain state integrity:
- Schema Writing — Every fetched job (including skipped duplicates) is written to
seen_jobs.jsonwith a rich schema containingtitle,company,url,first_seen,posted_date,deadline,fit,status,portal, andsource. - Tracker Consultation — Only jobs absent from
job_search_tracker.csvare flagged for presentation in the next step.
This write-ahead approach prevents data loss if the process interrupts before completion.
Step 5 – Present Results
The final step renders a human-readable summary:
- Sorted Table — Generates a markdown table sorted by fit rating (high → low), displaying only new opportunities not present in the application tracker.
- Diagnostics — Appends metadata including disabled portals, WebSearch fallback notifications, and health-check warnings (from Step 4.75).
- Next Actions — Prompts the user to proceed with deeper evaluation via
/rankor/applycommands.
Key Files and Data Sources
The workflow relies on specific files for deterministic execution:
| File | Role in the /scrape Workflow |
|---|---|
.claude/skills/job-scraper/SKILL.md |
Source of truth defining all six steps and sub-steps (1a, 1b, 1c). |
job_scraper/seen_jobs.json |
Persistent cache preventing reprocessing; updated in Step 4. |
job_search_tracker.csv |
Application history consulted in Steps 0 and 5. |
search-queries.md |
User-editable search strategy read during Steps 0 and 1. |
.agents/skills/*/SKILL.md |
Portal-specific skill definitions driving Step 1b CLI invocations. |
CLI Usage Examples
Invoke the workflow through the repository's thin-pointer CLI:
# Basic scrape – uses default top-3 query categories
/scrape
# Broad scrape – runs all query categories
/scrape broad
# Focused scrape – prioritize a specific area (e.g., data-science)
/scrape data science
# Health-check only – runs portal diagnostics without searching
/scrape health
Summary
- The
/scrapecommand implements a six-step deterministic pipeline defined in.claude/skills/job-scraper/SKILL.md. - Step 0 loads historical state from JSON and CSV files to prevent duplicate processing.
- Step 1 executes parallel portal queries with WebSearch fallback for missing CLI tools.
- Step 2 fetches full posting details and normalizes data fields while filtering existing entries.
- Step 3 applies lightweight fit classification with language-gate overrides.
- Step 4 persists all encountered jobs to
seen_jobs.jsonusing a comprehensive schema. - Step 5 presents a sorted markdown table of new opportunities and prompts for next actions.
Frequently Asked Questions
What happens if seen_jobs.json is missing when running /scrape?
The system automatically initializes an empty JSON structure with { "seen": {} } during Step 0 – Load State, ensuring the workflow continues without manual file creation.
How does the system handle portals without CLI tools installed?
During Step 1 – Search, sub-step 1c activates a WebSearch fallback whenever a portal CLI is unavailable or when bun is missing from the environment, maintaining coverage across all configured job boards.
What is the difference between seen_jobs.json and job_search_tracker.csv?
seen_jobs.json caches every job the scraper has ever encountered (whether applied to or not), while job_search_tracker.csv specifically tracks applications submitted by the user. The scraper consults both to avoid presenting previously seen or already-applied positions.
Can I run a focused search on a specific job category?
Yes. Instead of the default top-three categories or a broad search, you can pass a specific category name as an argument (e.g., /scrape data science) to prioritize that particular search strategy defined in search-queries.md.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →