# How the /scrape Command Discovers and Fetches Job Postings

> Learn how the /scrape command discovers and fetches job postings by executing portal CLIs via Bun, fetching details, and storing results in a deduplicated JSON registry.

- Repository: [Mads Lorentzen/ai-job-search](https://github.com/MadsLorentzen/ai-job-search)
- Tags: internals
- Published: 2026-08-31

---

**The `/scrape` command discovers job postings by enumerating installed portal-search CLIs defined in individual [`SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SKILL.md) files, executes them via the Bun runtime with query-specific flags, and fetches detailed posting data through portal-specific detail commands or a WebFetch fallback, storing all results in a deduplicated JSON registry.**

The `ai-job-search` repository by MadsLorentzen implements an intelligent job scraping system through a deterministic pipeline defined in [`.claude/skills/job-scraper/SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/.claude/skills/job-scraper/SKILL.md). When invoked, the command orchestrates the discovery, retrieval, and parsing of job postings across multiple portals without requiring code changes to integrate new sources.

## Skill Architecture and State Initialization

The `/scrape` skill is implemented as a structured Markdown file located at [`.claude/skills/job-scraper/SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/.claude/skills/job-scraper/SKILL.md). Before executing any searches, the command loads persistent state from three critical files:

- [`job_scraper/seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/job_scraper/seen_jobs.json) – Tracks previously encountered postings to prevent duplicates
- `job_search_tracker.csv` – Records companies and roles the user has already applied to
- [`search-queries.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/search-queries.md) – Contains the query strings and categories that drive each portal's search

This initialization occurs at lines 39‑44 of the skill definition, ensuring the system maintains context across runs and respects the user's existing application history.

## How Portal Discovery Works

The command discovers available job portals through a dynamic enumeration process rather than hardcoded lists. At lines 51‑69, the skill executes the following discovery logic:

1. **Runtime Verification** – Checks if `bun` (the JavaScript runtime) is installed, as portal CLIs depend on it
2. **Glob Pattern Matching** – Searches for all portal definitions using the pattern `.agents/skills/*/SKILL.md`
3. **Enabled Status Filtering** – Parses each discovered [`SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SKILL.md) and skips any portal where `enabled: false` is set
4. **Command Signature Extraction** – Reads the exact `bun run …` invocation and supported flags from each portal's skill file

This architecture allows new job portals to be integrated simply by adding a new directory under `.agents/skills/` with a properly configured [`SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SKILL.md) file.

## Executing Searches with Dynamic Flag Translation

Once portals are discovered, the command translates queries from [`search-queries.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/search-queries.md) into portal-specific CLI flags. The skill applies several constraints to optimize the search:

- **Recency Filtering** – Uses a 14‑day filter, leveraging the portal's native flag when available or filtering client‑side otherwise
- **Result Limiting** – Caps each call at approximately 20 results to manage API load and processing time
- **Format Specification** – Requests JSON output for structured parsing

If `bun` is missing, a CLI fails, or a portal lacks a command‑line interface, the skill automatically falls back to WebSearch using portal‑specific query strings (lines 77‑86).

## Fetching Full Posting Details

The fetch stage differentiates between CLI‑sourced and WebSearch‑sourced results, as detailed at lines 90‑107.

**For CLI results**, the data already contains `title`, `company`, `location`, `date`, and `url`. The skill then invokes the portal‑specific **detail** command (defined in that portal's [`SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SKILL.md)) to retrieve *requirements*, *application deadline*, and a short description.

**For WebSearch results**, the skill uses the generic `WebFetch` tool to download the page content, then manually extracts the same fields. During this process, the system detects expired postings (such as LinkedIn listings marked "No longer accepting applications") and records them with `status: "expired"`.

## Deduplication and State Management

After fetching, the system performs a lightweight fit assessment (lines 29‑36) scoring each posting as high, medium, or low fit before the full ranking step. All fetched jobs—whether new or skipped—are written to [`job_scraper/seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/job_scraper/seen_jobs.json) with provenance fields including:

- `first_seen` – Timestamp of initial discovery
- `portal` – Source portal name
- `source` – CLI or WebSearch origin
- `title`, `company`, `url`, `posted_date`, `deadline`, `fit`, `status`

This deduplication mechanism prevents re‑scraping identical postings across multiple runs and enables longitudinal tracking of job status changes.

## Portal Health Monitoring

Following each execution, the skill analyzes returned data to detect silent failures in portal CLIs (lines 97‑99). If a portal returns empty results or malformed data consistently, the system flags this for investigation, ensuring the user maintains awareness of which sources are functioning correctly.

## Practical Usage Examples

### Running the scrape command

```bash
/scrape                # default – run top‑3 query categories

/scrape broad          # search all query categories

/scrape data science   # focus on "data science" queries

/scrape health         # run only the portal‑health check (no search)

```

### Adding a new portal without code changes

1. Create a new directory under `.agents/skills/<portal-name>/`
2. Add a [`SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SKILL.md) file defining the CLI flags, recency parameters, and detail command
3. Run `/scrape` – the command automatically discovers and uses the new portal

### Inspecting stored results

```python
import json
import pathlib

seen_path = pathlib.Path("job_scraper/seen_jobs.json")
with seen_path.open() as f:
    data = json.load(f)
    
print(json.dumps(data["seen"], indent=2))

```

This outputs the complete registry of encountered postings with all provenance fields and current status values.

## Summary

- The `/scrape` command operates from [`.claude/skills/job-scraper/SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/.claude/skills/job-scraper/SKILL.md) as a deterministic pipeline with six distinct stages
- Portal discovery uses glob patterns (`.agents/skills/*/SKILL.md`) to find CLI definitions dynamically, skipping any marked `enabled: false`
- Search execution requires the Bun runtime and applies 14‑day recency filters with ~20 result limits per query
- Data fetching uses portal‑specific detail commands for CLI results or generic `WebFetch` for WebSearch fallbacks
- All postings are deduplicated against [`job_scraper/seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/job_scraper/seen_jobs.json) and `job_search_tracker.csv` with full provenance tracking
- Portal health monitoring detects silent CLI failures after each run

## Frequently Asked Questions

### How does the scrape command discover new job portals automatically?

The command searches for all [`SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SKILL.md) files under `.agents/skills/*/SKILL.md` at runtime. Each file defines the portal's CLI invocation, flags, and enabled status. When you add a new portal directory with a properly configured skill file, `/scrape` discovers it immediately on the next run without requiring changes to the core codebase.

### What happens if a job portal CLI fails or Bun is not installed?

If Bun is missing, a CLI returns an error, or a portal lacks a CLI implementation, the skill falls back to WebSearch (lines 77‑86). It uses the portal‑specific query strings from [`search-queries.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/search-queries.md) to perform web searches, then fetches individual posting details using the generic `WebFetch` tool. Results are tagged as *WebSearch‑sourced* to distinguish them from CLI‑sourced data.

### Where are the scraped job postings stored and what data is kept?

All postings are stored in [`job_scraper/seen_jobs.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/job_scraper/seen_jobs.json) with fields including `title`, `company`, `url`, `first_seen`, `posted_date`, `deadline`, `fit` (scoring), `status` (active/expired), `portal`, and `source`. This file prevents duplicate scraping and maintains a historical record of all encountered opportunities.

### How does the system handle expired or closed job listings?

During the fetch and parse stage (lines 90‑107), the system detects expired postings—such as LinkedIn listings marked "No longer accepting applications"—and records them with `status: "expired"`. This prevents users from wasting time on opportunities that are no longer available while maintaining the record for deduplication purposes.