How the /scrape Command Discovers and Fetches Job Postings
The /scrape command discovers job postings by enumerating installed portal-search CLIs defined in individual SKILL.md files, executes them via the Bun runtime with query-specific flags, and fetches detailed posting data through portal-specific detail commands or a WebFetch fallback, storing all results in a deduplicated JSON registry.
The ai-job-search repository by MadsLorentzen implements an intelligent job scraping system through a deterministic pipeline defined in .claude/skills/job-scraper/SKILL.md. When invoked, the command orchestrates the discovery, retrieval, and parsing of job postings across multiple portals without requiring code changes to integrate new sources.
Skill Architecture and State Initialization
The /scrape skill is implemented as a structured Markdown file located at .claude/skills/job-scraper/SKILL.md. Before executing any searches, the command loads persistent state from three critical files:
job_scraper/seen_jobs.json– Tracks previously encountered postings to prevent duplicatesjob_search_tracker.csv– Records companies and roles the user has already applied tosearch-queries.md– Contains the query strings and categories that drive each portal's search
This initialization occurs at lines 39‑44 of the skill definition, ensuring the system maintains context across runs and respects the user's existing application history.
How Portal Discovery Works
The command discovers available job portals through a dynamic enumeration process rather than hardcoded lists. At lines 51‑69, the skill executes the following discovery logic:
- Runtime Verification – Checks if
bun(the JavaScript runtime) is installed, as portal CLIs depend on it - Glob Pattern Matching – Searches for all portal definitions using the pattern
.agents/skills/*/SKILL.md - Enabled Status Filtering – Parses each discovered
SKILL.mdand skips any portal whereenabled: falseis set - Command Signature Extraction – Reads the exact
bun run …invocation and supported flags from each portal's skill file
This architecture allows new job portals to be integrated simply by adding a new directory under .agents/skills/ with a properly configured SKILL.md file.
Executing Searches with Dynamic Flag Translation
Once portals are discovered, the command translates queries from search-queries.md into portal-specific CLI flags. The skill applies several constraints to optimize the search:
- Recency Filtering – Uses a 14‑day filter, leveraging the portal's native flag when available or filtering client‑side otherwise
- Result Limiting – Caps each call at approximately 20 results to manage API load and processing time
- Format Specification – Requests JSON output for structured parsing
If bun is missing, a CLI fails, or a portal lacks a command‑line interface, the skill automatically falls back to WebSearch using portal‑specific query strings (lines 77‑86).
Fetching Full Posting Details
The fetch stage differentiates between CLI‑sourced and WebSearch‑sourced results, as detailed at lines 90‑107.
For CLI results, the data already contains title, company, location, date, and url. The skill then invokes the portal‑specific detail command (defined in that portal's SKILL.md) to retrieve requirements, application deadline, and a short description.
For WebSearch results, the skill uses the generic WebFetch tool to download the page content, then manually extracts the same fields. During this process, the system detects expired postings (such as LinkedIn listings marked "No longer accepting applications") and records them with status: "expired".
Deduplication and State Management
After fetching, the system performs a lightweight fit assessment (lines 29‑36) scoring each posting as high, medium, or low fit before the full ranking step. All fetched jobs—whether new or skipped—are written to job_scraper/seen_jobs.json with provenance fields including:
first_seen– Timestamp of initial discoveryportal– Source portal namesource– CLI or WebSearch origintitle,company,url,posted_date,deadline,fit,status
This deduplication mechanism prevents re‑scraping identical postings across multiple runs and enables longitudinal tracking of job status changes.
Portal Health Monitoring
Following each execution, the skill analyzes returned data to detect silent failures in portal CLIs (lines 97‑99). If a portal returns empty results or malformed data consistently, the system flags this for investigation, ensuring the user maintains awareness of which sources are functioning correctly.
Practical Usage Examples
Running the scrape command
/scrape # default – run top‑3 query categories
/scrape broad # search all query categories
/scrape data science # focus on "data science" queries
/scrape health # run only the portal‑health check (no search)
Adding a new portal without code changes
- Create a new directory under
.agents/skills/<portal-name>/ - Add a
SKILL.mdfile defining the CLI flags, recency parameters, and detail command - Run
/scrape– the command automatically discovers and uses the new portal
Inspecting stored results
import json
import pathlib
seen_path = pathlib.Path("job_scraper/seen_jobs.json")
with seen_path.open() as f:
data = json.load(f)
print(json.dumps(data["seen"], indent=2))
This outputs the complete registry of encountered postings with all provenance fields and current status values.
Summary
- The
/scrapecommand operates from.claude/skills/job-scraper/SKILL.mdas a deterministic pipeline with six distinct stages - Portal discovery uses glob patterns (
.agents/skills/*/SKILL.md) to find CLI definitions dynamically, skipping any markedenabled: false - Search execution requires the Bun runtime and applies 14‑day recency filters with ~20 result limits per query
- Data fetching uses portal‑specific detail commands for CLI results or generic
WebFetchfor WebSearch fallbacks - All postings are deduplicated against
job_scraper/seen_jobs.jsonandjob_search_tracker.csvwith full provenance tracking - Portal health monitoring detects silent CLI failures after each run
Frequently Asked Questions
How does the scrape command discover new job portals automatically?
The command searches for all SKILL.md files under .agents/skills/*/SKILL.md at runtime. Each file defines the portal's CLI invocation, flags, and enabled status. When you add a new portal directory with a properly configured skill file, /scrape discovers it immediately on the next run without requiring changes to the core codebase.
What happens if a job portal CLI fails or Bun is not installed?
If Bun is missing, a CLI returns an error, or a portal lacks a CLI implementation, the skill falls back to WebSearch (lines 77‑86). It uses the portal‑specific query strings from search-queries.md to perform web searches, then fetches individual posting details using the generic WebFetch tool. Results are tagged as WebSearch‑sourced to distinguish them from CLI‑sourced data.
Where are the scraped job postings stored and what data is kept?
All postings are stored in job_scraper/seen_jobs.json with fields including title, company, url, first_seen, posted_date, deadline, fit (scoring), status (active/expired), portal, and source. This file prevents duplicate scraping and maintains a historical record of all encountered opportunities.
How does the system handle expired or closed job listings?
During the fetch and parse stage (lines 90‑107), the system detects expired postings—such as LinkedIn listings marked "No longer accepting applications"—and records them with status: "expired". This prevents users from wasting time on opportunities that are no longer available while maintaining the record for deduplication purposes.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →