How Does the /scrape Command Discover New Skills in ai-job-search
The /scrape command discovers new skills by executing portal-specific CLIs to retrieve job postings, parsing structured JSON responses to extract normalized skill tokens from requirement fields, and persisting novel skills to the candidate profile with full audit provenance.
The /scrape command operates as the data ingestion engine for the job-scraper skill defined in the MadsLorentzen/ai-job-search repository. By continuously harvesting live job postings from configured portals, it identifies emerging skill requirements and automatically expands the candidate's capabilities file based on verified market demand.
The Skill Discovery Pipeline
The /scrape implementation follows a six-stage pipeline defined in .claude/skills/job-scraper/SKILL.md. Each stage transforms raw portal data into structured skill intelligence.
Portal CLI Invocation
The process begins by invoking portal-specific command-line interfaces located under .agents/skills/*/cli.md. These CLIs abstract the scraping logic for individual job portals (e.g., LinkedIn, Indeed) and return standardized JSON payloads containing posting metadata. The core orchestrator in tools/job_scraper.py automatically executes these CLIs for every configured portal, aggregating results that match the user's search query.
JSON Schema Parsing
Raw CLI responses conform to a common JSON schema with fields for title, description, requirements, and skills. The ScrapedJob model in tools/job_scraper.py maps these fields into typed internal structures. This abstraction layer ensures that portal-specific formatting differences are normalized before skill extraction begins.
Skill Token Extraction
For each posting, the scraper runs a lightweight NLP pipeline on the requirements and skills text. This pipeline:
- Normalizes synonyms (e.g., mapping "JavaScript" ↔ "JS" to a canonical token)
- Filters stop-words and irrelevant punctuation
- Generates deduplicated skill tokens representing discrete capabilities
The resulting token set represents the raw skill demand extracted from the job market.
Persistence and Provenance Tracking
Newly extracted skills undergo a validation check against the existing candidate profile stored in .claude/skills/candidate/skills.md. If a token is absent from the current profile:
- It is written to a temporary "discoveries" list
- The
/upskillcommand later merges this list into the canonical skill file - Provenance metadata (source URL, portal name, timestamp) is recorded in
job_scraper/provenance.json
This provenance tracking enables downstream commands like /rank to trace why specific skills were added and allows users to audit or prune false positives.
Gating and Feedback
The pipeline includes a feedback mechanism that checks gating fields (location, seniority level, employment type) defined in the job-scraper schema. If a posting fails these gates, it is excluded from skill extraction, preventing irrelevant skill noise from contaminating the candidate profile.
Core Implementation Files
The /scrape functionality is distributed across specification documents, implementation modules, and test suites:
.claude/skills/job-scraper/SKILL.md– Canonical contract defining the scrape step's inputs, outputs, and schema constraints.agents/skills/*/cli.md– Portal-specific CLI wrappers that feed JSON data into the scrapertools/job_scraper.py– Core implementation containing theScrapedJobmodel andrun_scrape()orchestration function.claude/skills/candidate/skills.md– Persistent candidate skill list updated via the discovery mechanismjob_scraper/provenance.json– JSON storage for skill audit trails (source URLs and timestamps)tests/test_scrape_contract.py– Unit tests verifying compliance with the SKILL.md specificationtests/test_scrape_provenance.py– Unit tests validating provenance data integrity
Usage Examples
Invoke the scraper programmatically to extract skills from current market postings:
from tools.job_scraper import run_scrape
# Execute scrape across configured portals
scrape_results = run_scrape(search_query="machine learning engineer")
# Access discovered skills as a set of normalized strings
new_skills = scrape_results.discovered_skills
print("Newly discovered skills:", new_skills)
# Audit skill origins through provenance data
for skill, meta in scrape_results.provenance.items():
print(f"{skill} → {meta['source_url']} @ {meta['timestamp']}")
Run via the command-line interface as defined in the portal CLI specifications:
# Trigger scrape for specific query
ai-job-search scrape --query "data scientist"
# Updates .claude/skills/candidate/skills.md with new discoveries
# and populates job_scraper/provenance.json with audit data
Summary
- The
/scrapecommand implements a specification-driven pipeline defined in.claude/skills/job-scraper/SKILL.mdto harvest skills from live job data. - Portal-specific CLIs under
.agents/skills/retrieve postings, whiletools/job_scraper.pynormalizes and processes the JSON responses. - An NLP tokenization layer normalizes synonyms and filters noise to generate canonical skill identifiers.
- Discovered skills are staged for integration into
.claude/skills/candidate/skills.mdand tracked with full provenance injob_scraper/provenance.json. - Gating logic prevents irrelevant postings from polluting the candidate's skill profile.
- The implementation is validated by
tests/test_scrape_contract.pyandtests/test_scrape_provenance.pyto ensure specification compliance.
Frequently Asked Questions
Where does /scrape store newly discovered skills before they are added to my profile?
Newly discovered skills are initially written to a temporary discoveries list within the scraper's internal state. The /upskill command subsequently reviews this list and merges validated skills into .claude/skills/candidate/skills.md, ensuring that the canonical candidate profile only grows through intentional integration.
How does the scraper prevent duplicate skills from being recorded?
The scraper performs a deduplication check against the existing skill set in .claude/skills/candidate/skills.md during the persistence stage. Additionally, the NLP normalization layer maps variant terms (such as "React.js" and "React") to canonical tokens before the duplication check occurs, preventing semantic duplicates rather than just string matches.
What is skill provenance and why is it tracked?
Provenance refers to the metadata recording where each skill originated, including the job posting URL, portal name, and extraction timestamp. This data is stored in job_scraper/provenance.json to provide auditability, allowing users to verify that skills were derived from legitimate market postings and enabling the /rank command to weight skills based on source authority or recency.
How does /scrape filter out irrelevant skills from unrelated job postings?
The command implements a gating system that evaluates postings against configurable filters such as location, seniority level, and required experience. If a posting fails these gates—indicating it is outside the candidate's target scope—it is excluded from skill extraction entirely, ensuring the discovered skills remain relevant to the user's actual job search criteria.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →