How Does the /scrape Command Discover New Skills in ai-job-search

The /scrape command discovers new skills by executing portal-specific CLIs to retrieve job postings, parsing structured JSON responses to extract normalized skill tokens from requirement fields, and persisting novel skills to the candidate profile with full audit provenance.

The /scrape command operates as the data ingestion engine for the job-scraper skill defined in the MadsLorentzen/ai-job-search repository. By continuously harvesting live job postings from configured portals, it identifies emerging skill requirements and automatically expands the candidate's capabilities file based on verified market demand.

The Skill Discovery Pipeline

The /scrape implementation follows a six-stage pipeline defined in .claude/skills/job-scraper/SKILL.md. Each stage transforms raw portal data into structured skill intelligence.

Portal CLI Invocation

The process begins by invoking portal-specific command-line interfaces located under .agents/skills/*/cli.md. These CLIs abstract the scraping logic for individual job portals (e.g., LinkedIn, Indeed) and return standardized JSON payloads containing posting metadata. The core orchestrator in tools/job_scraper.py automatically executes these CLIs for every configured portal, aggregating results that match the user's search query.

JSON Schema Parsing

Raw CLI responses conform to a common JSON schema with fields for title, description, requirements, and skills. The ScrapedJob model in tools/job_scraper.py maps these fields into typed internal structures. This abstraction layer ensures that portal-specific formatting differences are normalized before skill extraction begins.

Skill Token Extraction

For each posting, the scraper runs a lightweight NLP pipeline on the requirements and skills text. This pipeline:

  • Normalizes synonyms (e.g., mapping "JavaScript" ↔ "JS" to a canonical token)
  • Filters stop-words and irrelevant punctuation
  • Generates deduplicated skill tokens representing discrete capabilities

The resulting token set represents the raw skill demand extracted from the job market.

Persistence and Provenance Tracking

Newly extracted skills undergo a validation check against the existing candidate profile stored in .claude/skills/candidate/skills.md. If a token is absent from the current profile:

  1. It is written to a temporary "discoveries" list
  2. The /upskill command later merges this list into the canonical skill file
  3. Provenance metadata (source URL, portal name, timestamp) is recorded in job_scraper/provenance.json

This provenance tracking enables downstream commands like /rank to trace why specific skills were added and allows users to audit or prune false positives.

Gating and Feedback

The pipeline includes a feedback mechanism that checks gating fields (location, seniority level, employment type) defined in the job-scraper schema. If a posting fails these gates, it is excluded from skill extraction, preventing irrelevant skill noise from contaminating the candidate profile.

Core Implementation Files

The /scrape functionality is distributed across specification documents, implementation modules, and test suites:

Usage Examples

Invoke the scraper programmatically to extract skills from current market postings:

from tools.job_scraper import run_scrape

# Execute scrape across configured portals

scrape_results = run_scrape(search_query="machine learning engineer")

# Access discovered skills as a set of normalized strings

new_skills = scrape_results.discovered_skills
print("Newly discovered skills:", new_skills)

# Audit skill origins through provenance data

for skill, meta in scrape_results.provenance.items():
    print(f"{skill} → {meta['source_url']} @ {meta['timestamp']}")

Run via the command-line interface as defined in the portal CLI specifications:


# Trigger scrape for specific query

ai-job-search scrape --query "data scientist"

# Updates .claude/skills/candidate/skills.md with new discoveries

# and populates job_scraper/provenance.json with audit data

Summary

Frequently Asked Questions

Where does /scrape store newly discovered skills before they are added to my profile?

Newly discovered skills are initially written to a temporary discoveries list within the scraper's internal state. The /upskill command subsequently reviews this list and merges validated skills into .claude/skills/candidate/skills.md, ensuring that the canonical candidate profile only grows through intentional integration.

How does the scraper prevent duplicate skills from being recorded?

The scraper performs a deduplication check against the existing skill set in .claude/skills/candidate/skills.md during the persistence stage. Additionally, the NLP normalization layer maps variant terms (such as "React.js" and "React") to canonical tokens before the duplication check occurs, preventing semantic duplicates rather than just string matches.

What is skill provenance and why is it tracked?

Provenance refers to the metadata recording where each skill originated, including the job posting URL, portal name, and extraction timestamp. This data is stored in job_scraper/provenance.json to provide auditability, allowing users to verify that skills were derived from legitimate market postings and enabling the /rank command to weight skills based on source authority or recency.

How does /scrape filter out irrelevant skills from unrelated job postings?

The command implements a gating system that evaluates postings against configurable filters such as location, seniority level, and required experience. If a posting fails these gates—indicating it is outside the candidate's target scope—it is excluded from skill extraction entirely, ensuring the discovered skills remain relevant to the user's actual job search criteria.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →