How the /scrape Command Orchestrates Job Searches in the AI Job Search Framework

The /scrape command executes a seven-step pipeline—load state, parallel search with automatic portal discovery, fetch and parse listings, detect mass postings, assess fit, deduplicate, and generate referral links—coordinating CLI calls via the Agent tool while maintaining WebSearch fallback for resilience.

The /scrape command serves as the central orchestration engine for the AI Job Search Framework, automating the discovery and initial filtering of job opportunities across multiple portals. Implemented as a skill defined in /.claude/skills/job-scraper/SKILL.md, it drives the entire job-search pipeline by executing a series of well-defined steps that handle everything from runtime detection to duplicate prevention. This command integrates seamlessly with the framework's modular architecture, automatically discovering portal-specific CLI tools while ensuring robust fallback mechanisms remain available.

The Seven-Step Orchestration Pipeline

Step 0 – Load Previous State

The pipeline begins by reading three critical data sources: the previous scraper state stored in job_scraper/seen_jobs.json, the application tracker at job_search_tracker.csv, and the search strategy defined in search-queries.md. If the JSON state file is missing, the command automatically initializes it with an empty "seen" map, ensuring the deduplication engine has a clean slate to work from.

Step 1 – Parallel Search Across Portals

The search phase implements a sophisticated three-tier strategy:

Runtime Validation – The command first checks for the bun runtime availability. If bun is not present, the entire search operation falls back to WebSearch to prevent execution failures.

Portal Discovery and Execution – When bun is available, the skill discovers every portal-CLI skill located under .agents/skills/*/SKILL.md. It respects each portal's enabled flag, translates queries from search-queries.md into portal-specific CLI flags, scopes results to the last 14 days, caps results at approximately 20 per portal, and forces --format json for machine-readable output. All CLI calls execute in parallel via the Agent tool, maximizing throughput across multiple job boards simultaneously.

WebSearch Fallback – For portals lacking CLI implementations or when bun is unavailable, the command automatically routes queries through the WebSearch tool using raw query strings from search-queries.md.

Step 2 – Fetch and Parse Listings

For each promising result, the skill retrieves detailed job information through two methods. It either calls the portal's detail command (as defined in the portal's own SKILL.md) to extract title, company, location, posting date, deadline, and description, or uses WebFetch on the URL with fallback headers when the result originates from WebSearch. The system detects closed-source postings (such as linkedin-search returning isActive: false) and records these as "expired" in the state file to prevent future processing.

Step 2.5 – Mass-Posting Detection

The command identifies clusters of near-identical postings from the same employer, consolidating them into single entries with annotations like "posted in 6 cities." This prevents clutter in the results while maintaining awareness of multi-location opportunities.

Step 3 – Quick Fit Assessment

Each job receives a coarse fit label—high, medium, or low—based on alignment with the candidate's core skills defined in CLAUDE.md. The assessment applies an overriding rule blocking jobs that require languages not declared in the candidate's profile, ensuring language requirements act as hard filters regardless of other skill matches.

Step 4 – Deduplication and Persistent Storage

Every fetched job, whether new, skipped, or expired, writes to job_scraper/seen_jobs.json with a rich schema including first_seen, posted_date, deadline, fit, status, portal, and source (CLI vs. WebSearch). Before presenting any job to the user, the system checks both the JSON state and the tracker CSV to eliminate duplicates and avoid resurfacing previously applied positions.

For high and medium-fit jobs, the command generates two LinkedIn search URLs (recruiter and peer links) to facilitate networking, though these remain links only and are never fetched programmatically.

Following the search run, the command executes a portal health check, examining each enabled portal for degradation indicators such as missing titles or empty company fields. If a portal appears broken, a sentinel probe performs a single search with optional retry to confirm status, with results reported in the final diagnostic table.

Step 5 – Results Presentation

The command emits a markdown table sorted by fit level, with optional sections for disabled portals, fallback portals, and health diagnostics. Each row includes the posting link, fit level, and flags for mass postings or language gates. After presentation, the system prompts the user to request deeper evaluation via /rank or detailed analysis commands.

Step 6 – State Management

While /scrape initializes the state and prevents duplicate scraping, it defers tracker updates to downstream commands when the user actually applies to positions, maintaining clean separation between discovery and application tracking.

Data-Driven Discovery and Resilient Architecture

The orchestration follows a discover-run-fallback-parse-dedupe-store-report pattern that remains completely data-driven. New portal skills are automatically included because Step 1b scans every SKILL.md under .agents/skills/. The enabled flag in each portal's configuration allows users to disable sources without removing code, while the WebSearch fallback ensures missing or broken CLIs never halt the workflow. Detailed provenance tracking in seen_jobs.json enables downstream commands (/rank, /apply, /outcome) to rely on a single source of truth regarding job origins and discovery dates.

Command Usage Examples

The /scrape command supports several invocation patterns to control search scope:


# Basic scrape – finds new jobs using all enabled portal CLIs

/scrape

# Scrape a specific focus area (e.g., data science)

/scrape data science

# Run a broad search across all query categories

/scrape broad

# Health-check only (no searching, just portal diagnostics)

/scrape health

These invocations trigger the full pipeline described above, with the skill parsing arguments to adjust query selection and execution mode. The command is also runnable via the generic agent interface (e.g., agent run /scrape) because the skill advertises allowed-tools: … Agent … in its definition.

Core Implementation Files

The /scrape command relies on several key files that define its behavior and maintain state:

Summary

  • The /scrape command implements a seven-step pipeline defined in /.claude/skills/job-scraper/SKILL.md that orchestrates job discovery, parsing, and initial filtering.
  • It automatically discovers portal CLI tools under .agents/skills/*/SKILL.md while respecting enabled flags, executing searches in parallel via the Agent tool.
  • The system maintains resilience through WebSearch fallback when the bun runtime is unavailable or portal CLIs fail.
  • Mass-posting detection consolidates duplicate employer listings, while fit assessment applies hard filters for undeclared language requirements.
  • Comprehensive provenance tracking in job_scraper/seen_jobs.json includes first_seen dates, sources (CLI vs. WebSearch), and portal metadata to support downstream workflow commands.
  • Portal health checks with sentinel probes ensure degraded data sources are identified and reported without breaking the search workflow.

Frequently Asked Questions

What happens if the bun runtime is not installed when running /scrape?

If the bun runtime is unavailable, the entire search operation automatically falls back to WebSearch using the raw query strings from search-queries.md. This ensures the workflow continues uninterrupted even when JavaScript/TypeScript CLI tools cannot execute, though results may vary in structure compared to native portal CLIs.

How does /scrape prevent showing the same job multiple times?

The command implements duplicate detection at multiple stages. Before presentation, it checks both job_scraper/seen_jobs.json (the persistent state of all previously seen jobs) and job_search_tracker.csv (jobs the user has already applied to). Every job—whether new, skipped, or expired—is recorded in seen_jobs.json with a unique identifier, preventing future rescraping regardless of which portal originally sourced the listing.

Can I disable specific job portals without removing their code?

Yes. The /scrape command respects the enabled flag defined in each portal's SKILL.md file under .agents/skills/. Setting enabled: false in a portal's configuration excludes it from the automatic discovery and execution process in Step 1b, allowing you to maintain the skill files while excluding them from active searches.

What information is stored about each job in seen_jobs.json?

The JSON schema captures rich metadata including first_seen (timestamp), posted_date, deadline, fit level (high/medium/low), status (active/expired), originating portal, and source type (CLI vs. WebSearch). This provenance data enables downstream commands like /rank and /apply to make informed decisions without rescraping original sources.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →