# How the /scrape Command Orchestrates Job Searches in the AI Job Search Framework

> Discover how the /scrape command orchestrates AI job searches via a seven-step pipeline. Learn about its parallel search, listing parsing, and referral link generation.

- Repository: [Mads Lorentzen/ai-job-search](https://github.com/MadsLorentzen/ai-job-search)
- Tags: how-to-guide
- Published: 2026-09-01

---

**The `/scrape` command executes a seven-step pipeline—load state, parallel search with automatic portal discovery, fetch and parse listings, detect mass postings, assess fit, deduplicate, and generate referral links—coordinating CLI calls via the Agent tool while maintaining WebSearch fallback for resilience.**

The `/scrape` command serves as the central orchestration engine for the AI Job Search Framework, automating the discovery and initial filtering of job opportunities across multiple portals. Implemented as a skill defined in [`/.claude/skills/job-scraper/SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main//.claude/skills/job-scraper/SKILL.md), it drives the entire job-search pipeline by executing a series of well-defined steps that handle everything from runtime detection to duplicate prevention. This command integrates seamlessly with the framework's modular architecture, automatically discovering portal-specific CLI tools while ensuring robust fallback mechanisms remain available.

## The Seven-Step Orchestration Pipeline

### Step 0 – Load Previous State

The pipeline begins by reading three critical data sources: the previous scraper state stored in [`job_scraper/seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/job_scraper/seen_jobs.json), the application tracker at `job_search_tracker.csv`, and the search strategy defined in [`search-queries.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/search-queries.md). If the JSON state file is missing, the command automatically initializes it with an empty `"seen"` map, ensuring the deduplication engine has a clean slate to work from.

### Step 1 – Parallel Search Across Portals

The search phase implements a sophisticated three-tier strategy:

**Runtime Validation** – The command first checks for the `bun` runtime availability. If `bun` is not present, the entire search operation falls back to **WebSearch** to prevent execution failures.

**Portal Discovery and Execution** – When `bun` is available, the skill discovers every portal-CLI skill located under `.agents/skills/*/SKILL.md`. It respects each portal's `enabled` flag, translates queries from [`search-queries.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/search-queries.md) into portal-specific CLI flags, scopes results to the last **14 days**, caps results at approximately 20 per portal, and forces `--format json` for machine-readable output. All CLI calls execute in parallel via the **Agent** tool, maximizing throughput across multiple job boards simultaneously.

**WebSearch Fallback** – For portals lacking CLI implementations or when `bun` is unavailable, the command automatically routes queries through the **WebSearch** tool using raw query strings from [`search-queries.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/search-queries.md).

### Step 2 – Fetch and Parse Listings

For each promising result, the skill retrieves detailed job information through two methods. It either calls the portal's **detail** command (as defined in the portal's own [`SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SKILL.md)) to extract title, company, location, posting date, deadline, and description, or uses **WebFetch** on the URL with fallback headers when the result originates from WebSearch. The system detects closed-source postings (such as `linkedin-search` returning `isActive: false`) and records these as `"expired"` in the state file to prevent future processing.

### Step 2.5 – Mass-Posting Detection

The command identifies clusters of near-identical postings from the same employer, consolidating them into single entries with annotations like "posted in 6 cities." This prevents clutter in the results while maintaining awareness of multi-location opportunities.

### Step 3 – Quick Fit Assessment

Each job receives a coarse fit label—**high**, **medium**, or **low**—based on alignment with the candidate's core skills defined in [`CLAUDE.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/CLAUDE.md). The assessment applies an overriding rule blocking jobs that require languages not declared in the candidate's profile, ensuring language requirements act as hard filters regardless of other skill matches.

### Step 4 – Deduplication and Persistent Storage

Every fetched job, whether new, skipped, or expired, writes to [`job_scraper/seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/job_scraper/seen_jobs.json) with a rich schema including `first_seen`, `posted_date`, `deadline`, `fit`, `status`, `portal`, and `source` (CLI vs. WebSearch). Before presenting any job to the user, the system checks both the JSON state and the tracker CSV to eliminate duplicates and avoid resurfacing previously applied positions.

### Steps 4.5 and 4.75 – Referral Links and Health Monitoring

For **high** and **medium**-fit jobs, the command generates two LinkedIn search URLs (recruiter and peer links) to facilitate networking, though these remain links only and are never fetched programmatically.

Following the search run, the command executes a **portal health check**, examining each enabled portal for degradation indicators such as missing titles or empty company fields. If a portal appears broken, a sentinel probe performs a single search with optional retry to confirm status, with results reported in the final diagnostic table.

### Step 5 – Results Presentation

The command emits a markdown table sorted by fit level, with optional sections for disabled portals, fallback portals, and health diagnostics. Each row includes the posting link, fit level, and flags for mass postings or language gates. After presentation, the system prompts the user to request deeper evaluation via `/rank` or detailed analysis commands.

### Step 6 – State Management

While `/scrape` initializes the state and prevents duplicate scraping, it defers tracker updates to downstream commands when the user actually applies to positions, maintaining clean separation between discovery and application tracking.

## Data-Driven Discovery and Resilient Architecture

The orchestration follows a **discover-run-fallback-parse-dedupe-store-report** pattern that remains completely data-driven. New portal skills are automatically included because Step 1b scans every [`SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SKILL.md) under `.agents/skills/`. The `enabled` flag in each portal's configuration allows users to disable sources without removing code, while the **WebSearch** fallback ensures missing or broken CLIs never halt the workflow. Detailed provenance tracking in [`seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/seen_jobs.json) enables downstream commands (`/rank`, `/apply`, `/outcome`) to rely on a single source of truth regarding job origins and discovery dates.

## Command Usage Examples

The `/scrape` command supports several invocation patterns to control search scope:

```bash

# Basic scrape – finds new jobs using all enabled portal CLIs

/scrape

# Scrape a specific focus area (e.g., data science)

/scrape data science

# Run a broad search across all query categories

/scrape broad

# Health-check only (no searching, just portal diagnostics)

/scrape health

```

These invocations trigger the full pipeline described above, with the skill parsing arguments to adjust query selection and execution mode. The command is also runnable via the generic agent interface (e.g., `agent run /scrape`) because the skill advertises `allowed-tools: … Agent …` in its definition.

## Core Implementation Files

The `/scrape` command relies on several key files that define its behavior and maintain state:

- [`/.claude/skills/job-scraper/SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main//.claude/skills/job-scraper/SKILL.md) – Core workflow definition containing steps, flags, and the output schema
- [`/.claude/skills/job-scraper/search-queries.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main//.claude/skills/job-scraper/search-queries.md) – Prioritized search query categories used in Step 1
- `.agents/skills/*/SKILL.md` – Individual portal-CLI specifications that `/scrape` discovers and executes dynamically
- [`tests/test_scrape_contract.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tests/test_scrape_contract.py) – CI test enforcing Step 2 data contract (presence of `date`, `company`, etc.) across all portals
- [`tests/test_scrape_provenance.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tests/test_scrape_provenance.py) – Tests verifying correct provenance recording (source, `first_seen`) by `/scrape`
- [`job_scraper/seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/job_scraper/seen_jobs.json) – Runtime JSON store of all scraped jobs for deduplication and provenance tracking
- `job_search_tracker.csv` – CSV file tracking previously applied jobs, consulted during deduplication

## Summary

- The `/scrape` command implements a seven-step pipeline defined in [`/.claude/skills/job-scraper/SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main//.claude/skills/job-scraper/SKILL.md) that orchestrates job discovery, parsing, and initial filtering.
- It automatically discovers portal CLI tools under `.agents/skills/*/SKILL.md` while respecting `enabled` flags, executing searches in parallel via the **Agent** tool.
- The system maintains resilience through **WebSearch** fallback when the `bun` runtime is unavailable or portal CLIs fail.
- **Mass-posting detection** consolidates duplicate employer listings, while **fit assessment** applies hard filters for undeclared language requirements.
- Comprehensive provenance tracking in [`job_scraper/seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/job_scraper/seen_jobs.json) includes `first_seen` dates, sources (CLI vs. WebSearch), and portal metadata to support downstream workflow commands.
- **Portal health checks** with sentinel probes ensure degraded data sources are identified and reported without breaking the search workflow.

## Frequently Asked Questions

### What happens if the bun runtime is not installed when running /scrape?

If the `bun` runtime is unavailable, the entire search operation automatically falls back to **WebSearch** using the raw query strings from [`search-queries.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/search-queries.md). This ensures the workflow continues uninterrupted even when JavaScript/TypeScript CLI tools cannot execute, though results may vary in structure compared to native portal CLIs.

### How does /scrape prevent showing the same job multiple times?

The command implements duplicate detection at multiple stages. Before presentation, it checks both [`job_scraper/seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/job_scraper/seen_jobs.json) (the persistent state of all previously seen jobs) and `job_search_tracker.csv` (jobs the user has already applied to). Every job—whether new, skipped, or expired—is recorded in [`seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/seen_jobs.json) with a unique identifier, preventing future rescraping regardless of which portal originally sourced the listing.

### Can I disable specific job portals without removing their code?

Yes. The `/scrape` command respects the `enabled` flag defined in each portal's [`SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SKILL.md) file under `.agents/skills/`. Setting `enabled: false` in a portal's configuration excludes it from the automatic discovery and execution process in Step 1b, allowing you to maintain the skill files while excluding them from active searches.

### What information is stored about each job in seen_jobs.json?

The JSON schema captures rich metadata including `first_seen` (timestamp), `posted_date`, `deadline`, `fit` level (high/medium/low), `status` (active/expired), originating `portal`, and `source` type (CLI vs. WebSearch). This provenance data enables downstream commands like `/rank` and `/apply` to make informed decisions without rescraping original sources.