How the /scrape Command Discovers and Fetches Job Postings

The /scrape command discovers job postings by enumerating installed portal-search CLIs defined in individual SKILL.md files, executes them via the Bun runtime with query-specific flags, and fetches detailed posting data through portal-specific detail commands or a WebFetch fallback, storing all results in a deduplicated JSON registry.

The ai-job-search repository by MadsLorentzen implements an intelligent job scraping system through a deterministic pipeline defined in .claude/skills/job-scraper/SKILL.md. When invoked, the command orchestrates the discovery, retrieval, and parsing of job postings across multiple portals without requiring code changes to integrate new sources.

Skill Architecture and State Initialization

The /scrape skill is implemented as a structured Markdown file located at .claude/skills/job-scraper/SKILL.md. Before executing any searches, the command loads persistent state from three critical files:

  • job_scraper/seen_jobs.json – Tracks previously encountered postings to prevent duplicates
  • job_search_tracker.csv – Records companies and roles the user has already applied to
  • search-queries.md – Contains the query strings and categories that drive each portal's search

This initialization occurs at lines 39‑44 of the skill definition, ensuring the system maintains context across runs and respects the user's existing application history.

How Portal Discovery Works

The command discovers available job portals through a dynamic enumeration process rather than hardcoded lists. At lines 51‑69, the skill executes the following discovery logic:

  1. Runtime Verification – Checks if bun (the JavaScript runtime) is installed, as portal CLIs depend on it
  2. Glob Pattern Matching – Searches for all portal definitions using the pattern .agents/skills/*/SKILL.md
  3. Enabled Status Filtering – Parses each discovered SKILL.md and skips any portal where enabled: false is set
  4. Command Signature Extraction – Reads the exact bun run … invocation and supported flags from each portal's skill file

This architecture allows new job portals to be integrated simply by adding a new directory under .agents/skills/ with a properly configured SKILL.md file.

Executing Searches with Dynamic Flag Translation

Once portals are discovered, the command translates queries from search-queries.md into portal-specific CLI flags. The skill applies several constraints to optimize the search:

  • Recency Filtering – Uses a 14‑day filter, leveraging the portal's native flag when available or filtering client‑side otherwise
  • Result Limiting – Caps each call at approximately 20 results to manage API load and processing time
  • Format Specification – Requests JSON output for structured parsing

If bun is missing, a CLI fails, or a portal lacks a command‑line interface, the skill automatically falls back to WebSearch using portal‑specific query strings (lines 77‑86).

Fetching Full Posting Details

The fetch stage differentiates between CLI‑sourced and WebSearch‑sourced results, as detailed at lines 90‑107.

For CLI results, the data already contains title, company, location, date, and url. The skill then invokes the portal‑specific detail command (defined in that portal's SKILL.md) to retrieve requirements, application deadline, and a short description.

For WebSearch results, the skill uses the generic WebFetch tool to download the page content, then manually extracts the same fields. During this process, the system detects expired postings (such as LinkedIn listings marked "No longer accepting applications") and records them with status: "expired".

Deduplication and State Management

After fetching, the system performs a lightweight fit assessment (lines 29‑36) scoring each posting as high, medium, or low fit before the full ranking step. All fetched jobs—whether new or skipped—are written to job_scraper/seen_jobs.json with provenance fields including:

  • first_seen – Timestamp of initial discovery
  • portal – Source portal name
  • source – CLI or WebSearch origin
  • title, company, url, posted_date, deadline, fit, status

This deduplication mechanism prevents re‑scraping identical postings across multiple runs and enables longitudinal tracking of job status changes.

Portal Health Monitoring

Following each execution, the skill analyzes returned data to detect silent failures in portal CLIs (lines 97‑99). If a portal returns empty results or malformed data consistently, the system flags this for investigation, ensuring the user maintains awareness of which sources are functioning correctly.

Practical Usage Examples

Running the scrape command

/scrape                # default – run top‑3 query categories

/scrape broad          # search all query categories

/scrape data science   # focus on "data science" queries

/scrape health         # run only the portal‑health check (no search)

Adding a new portal without code changes

  1. Create a new directory under .agents/skills/<portal-name>/
  2. Add a SKILL.md file defining the CLI flags, recency parameters, and detail command
  3. Run /scrape – the command automatically discovers and uses the new portal

Inspecting stored results

import json
import pathlib

seen_path = pathlib.Path("job_scraper/seen_jobs.json")
with seen_path.open() as f:
    data = json.load(f)
    
print(json.dumps(data["seen"], indent=2))

This outputs the complete registry of encountered postings with all provenance fields and current status values.

Summary

  • The /scrape command operates from .claude/skills/job-scraper/SKILL.md as a deterministic pipeline with six distinct stages
  • Portal discovery uses glob patterns (.agents/skills/*/SKILL.md) to find CLI definitions dynamically, skipping any marked enabled: false
  • Search execution requires the Bun runtime and applies 14‑day recency filters with ~20 result limits per query
  • Data fetching uses portal‑specific detail commands for CLI results or generic WebFetch for WebSearch fallbacks
  • All postings are deduplicated against job_scraper/seen_jobs.json and job_search_tracker.csv with full provenance tracking
  • Portal health monitoring detects silent CLI failures after each run

Frequently Asked Questions

How does the scrape command discover new job portals automatically?

The command searches for all SKILL.md files under .agents/skills/*/SKILL.md at runtime. Each file defines the portal's CLI invocation, flags, and enabled status. When you add a new portal directory with a properly configured skill file, /scrape discovers it immediately on the next run without requiring changes to the core codebase.

What happens if a job portal CLI fails or Bun is not installed?

If Bun is missing, a CLI returns an error, or a portal lacks a CLI implementation, the skill falls back to WebSearch (lines 77‑86). It uses the portal‑specific query strings from search-queries.md to perform web searches, then fetches individual posting details using the generic WebFetch tool. Results are tagged as WebSearch‑sourced to distinguish them from CLI‑sourced data.

Where are the scraped job postings stored and what data is kept?

All postings are stored in job_scraper/seen_jobs.json with fields including title, company, url, first_seen, posted_date, deadline, fit (scoring), status (active/expired), portal, and source. This file prevents duplicate scraping and maintains a historical record of all encountered opportunities.

How does the system handle expired or closed job listings?

During the fetch and parse stage (lines 90‑107), the system detects expired postings—such as LinkedIn listings marked "No longer accepting applications"—and records them with status: "expired". This prevents users from wasting time on opportunities that are no longer available while maintaining the record for deduplication purposes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →