# The Six Execution Steps of the /scrape Workflow in AI-Job-Search

> Discover the six execution steps of the /scrape workflow in AI-Job-Search: Load State, Search, Fetch & Parse, Quick Fit Assessment, Deduplicate & Store, and Present Results. Automate job discovery efficiently.

- Repository: [Mads Lorentzen/ai-job-search](https://github.com/MadsLorentzen/ai-job-search)
- Tags: internals
- Published: 2026-08-31

---

**The `/scrape` workflow executes six sequential steps—Load State, Search, Fetch & Parse, Quick Fit Assessment, Deduplicate & Store, and Present Results—to automate job discovery while avoiding duplicate postings.**

The `/scrape` command is implemented as the *job-scraper* skill in the [MadsLorentzen/ai-job-search](https://github.com/MadsLorentzen/ai-job-search) repository. Defined in [`.claude/skills/job-scraper/SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/.claude/skills/job-scraper/SKILL.md), this deterministic pipeline aggregates postings from multiple job portals, evaluates candidate fit, and surfaces new opportunities without reprocessing historical data.

## Overview of the /scrape Workflow

The workflow is designed for **idempotent execution**. Each run consults persistent state files to avoid redundant API calls and duplicate entries. The six core steps—numbered Step 0 through Step 5—handle everything from initialization to final presentation, with an optional Step 6 ("Update Tracker") available only when manually recording external applications.

## Step-by-Step Breakdown

### Step 0 – Load State

The scraper initializes its knowledge base by loading three critical data sources:

- **[`job_scraper/seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/job_scraper/seen_jobs.json)** — A persistent cache of every job previously encountered. If the file is missing, the system initializes it with an empty `{ "seen": {} }` structure.
- **`job_search_tracker.csv`** — The application tracker used to identify companies and roles already applied to.
- **[`search-queries.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/search-queries.md)** — User-defined search strategies specifying which query categories to prioritize.

This initialization ensures that subsequent steps only process novel opportunities.

### Step 1 – Search

This step executes the actual job-board queries according to the strategy defined in [`search-queries.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/search-queries.md):

1. **Category Selection** — By default, the system runs the top three query categories. Users can override this with `broad` (all categories) or a focused category argument.
2. **CLI Preference** — The system prefers installed portal-specific CLI tools, falling back to **WebSearch** when a CLI is unavailable or when **bun** is missing from the environment.
3. **Parallel Execution** — Sub-step **1b** runs each enabled portal's CLI in parallel, respecting individual portal `enabled` flags and recency settings.

If a portal fails or lacks CLI support, sub-step **1c** triggers a WebSearch fallback to ensure coverage.

### Step 2 – Fetch & Parse

For each promising result identified in Step 1, the system retrieves comprehensive posting details:

- **Data Retrieval** — Invokes the portal's `detail` command for CLI-driven portals, or uses **WebFetch** for WebSearch-derived URLs.
- **Field Extraction** — Captures `title`, `company`, `location`, `posting_date`, `deadline`, a short description, and an `isActive` flag (specifically for LinkedIn postings).
- **Duplicate Filtering** — Immediately skips entries existing in [`seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/seen_jobs.json) or the tracker CSV, preventing redundant processing.

### Step 3 – Quick Fit Assessment

Each new job receives a lightweight classification based on predefined criteria:

- **Match Rating** — Jobs are categorized as **high**, **medium**, or **low** match according to skill overlap analysis.
- **Language Gate** — A hard filter overrides the rating if the posting requires a language the candidate has not declared proficiency in.

This ensures that only relevant opportunities progress to the final presentation stage.

### Step 4 – Deduplicate & Store

Persistence occurs before presentation to maintain state integrity:

- **Schema Writing** — Every fetched job (including skipped duplicates) is written to [`seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/seen_jobs.json) with a rich schema containing `title`, `company`, `url`, `first_seen`, `posted_date`, `deadline`, `fit`, `status`, `portal`, and `source`.
- **Tracker Consultation** — Only jobs absent from `job_search_tracker.csv` are flagged for presentation in the next step.

This write-ahead approach prevents data loss if the process interrupts before completion.

### Step 5 – Present Results

The final step renders a human-readable summary:

- **Sorted Table** — Generates a markdown table sorted by fit rating (high → low), displaying only new opportunities not present in the application tracker.
- **Diagnostics** — Appends metadata including disabled portals, WebSearch fallback notifications, and health-check warnings (from Step 4.75).
- **Next Actions** — Prompts the user to proceed with deeper evaluation via `/rank` or `/apply` commands.

## Key Files and Data Sources

The workflow relies on specific files for deterministic execution:

| File | Role in the /scrape Workflow |
|------|------------------------------|
| [`.claude/skills/job-scraper/SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/.claude/skills/job-scraper/SKILL.md) | Source of truth defining all six steps and sub-steps (1a, 1b, 1c). |
| [`job_scraper/seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/job_scraper/seen_jobs.json) | Persistent cache preventing reprocessing; updated in Step 4. |
| `job_search_tracker.csv` | Application history consulted in Steps 0 and 5. |
| [`search-queries.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/search-queries.md) | User-editable search strategy read during Steps 0 and 1. |
| `.agents/skills/*/SKILL.md` | Portal-specific skill definitions driving Step 1b CLI invocations. |

## CLI Usage Examples

Invoke the workflow through the repository's thin-pointer CLI:

```bash

# Basic scrape – uses default top-3 query categories

/scrape

# Broad scrape – runs all query categories

/scrape broad

# Focused scrape – prioritize a specific area (e.g., data-science)

/scrape data science

# Health-check only – runs portal diagnostics without searching

/scrape health

```

## Summary

- The `/scrape` command implements a six-step deterministic pipeline defined in [`.claude/skills/job-scraper/SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/.claude/skills/job-scraper/SKILL.md).
- **Step 0** loads historical state from JSON and CSV files to prevent duplicate processing.
- **Step 1** executes parallel portal queries with WebSearch fallback for missing CLI tools.
- **Step 2** fetches full posting details and normalizes data fields while filtering existing entries.
- **Step 3** applies lightweight fit classification with language-gate overrides.
- **Step 4** persists all encountered jobs to [`seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/seen_jobs.json) using a comprehensive schema.
- **Step 5** presents a sorted markdown table of new opportunities and prompts for next actions.

## Frequently Asked Questions

### What happens if [`seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/seen_jobs.json) is missing when running /scrape?

The system automatically initializes an empty JSON structure with `{ "seen": {} }` during **Step 0 – Load State**, ensuring the workflow continues without manual file creation.

### How does the system handle portals without CLI tools installed?

During **Step 1 – Search**, sub-step **1c** activates a **WebSearch** fallback whenever a portal CLI is unavailable or when **bun** is missing from the environment, maintaining coverage across all configured job boards.

### What is the difference between [`seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/seen_jobs.json) and `job_search_tracker.csv`?

[`seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/seen_jobs.json) caches every job the scraper has ever encountered (whether applied to or not), while `job_search_tracker.csv` specifically tracks applications submitted by the user. The scraper consults both to avoid presenting previously seen or already-applied positions.

### Can I run a focused search on a specific job category?

Yes. Instead of the default top-three categories or a `broad` search, you can pass a specific category name as an argument (e.g., `/scrape data science`) to prioritize that particular search strategy defined in [`search-queries.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/search-queries.md).