# How Does the /scrape Command Discover New Skills in ai-job-search

> Discover how the /scrape command in ai-job-search finds new skills. It parses job postings, extracts skill tokens, and saves novel skills to candidate profiles.

- Repository: [Mads Lorentzen/ai-job-search](https://github.com/MadsLorentzen/ai-job-search)
- Tags: internals
- Published: 2026-08-28

---

**The /scrape command discovers new skills by executing portal-specific CLIs to retrieve job postings, parsing structured JSON responses to extract normalized skill tokens from requirement fields, and persisting novel skills to the candidate profile with full audit provenance.**

The `/scrape` command operates as the data ingestion engine for the **job-scraper skill** defined in the `MadsLorentzen/ai-job-search` repository. By continuously harvesting live job postings from configured portals, it identifies emerging skill requirements and automatically expands the candidate's capabilities file based on verified market demand.

## The Skill Discovery Pipeline

The `/scrape` implementation follows a six-stage pipeline defined in [`.claude/skills/job-scraper/SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/.claude/skills/job-scraper/SKILL.md). Each stage transforms raw portal data into structured skill intelligence.

### Portal CLI Invocation

The process begins by invoking portal-specific command-line interfaces located under `.agents/skills/*/cli.md`. These CLIs abstract the scraping logic for individual job portals (e.g., LinkedIn, Indeed) and return standardized JSON payloads containing posting metadata. The core orchestrator in [`tools/job_scraper.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/job_scraper.py) automatically executes these CLIs for every configured portal, aggregating results that match the user's search query.

### JSON Schema Parsing

Raw CLI responses conform to a common JSON schema with fields for **title**, **description**, **requirements**, and **skills**. The `ScrapedJob` model in [`tools/job_scraper.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/job_scraper.py) maps these fields into typed internal structures. This abstraction layer ensures that portal-specific formatting differences are normalized before skill extraction begins.

### Skill Token Extraction

For each posting, the scraper runs a lightweight NLP pipeline on the **requirements** and **skills** text. This pipeline:

- **Normalizes synonyms** (e.g., mapping "JavaScript" ↔ "JS" to a canonical token)
- **Filters stop-words** and irrelevant punctuation
- **Generates deduplicated skill tokens** representing discrete capabilities

The resulting token set represents the raw skill demand extracted from the job market.

### Persistence and Provenance Tracking

Newly extracted skills undergo a validation check against the existing candidate profile stored in [`.claude/skills/candidate/skills.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/.claude/skills/candidate/skills.md). If a token is absent from the current profile:

1. It is written to a temporary "discoveries" list
2. The `/upskill` command later merges this list into the canonical skill file
3. **Provenance metadata** (source URL, portal name, timestamp) is recorded in [`job_scraper/provenance.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/job_scraper/provenance.json)

This provenance tracking enables downstream commands like `/rank` to trace why specific skills were added and allows users to audit or prune false positives.

### Gating and Feedback

The pipeline includes a feedback mechanism that checks **gating fields** (location, seniority level, employment type) defined in the job-scraper schema. If a posting fails these gates, it is excluded from skill extraction, preventing irrelevant skill noise from contaminating the candidate profile.

## Core Implementation Files

The `/scrape` functionality is distributed across specification documents, implementation modules, and test suites:

- **[`.claude/skills/job-scraper/SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/.claude/skills/job-scraper/SKILL.md)** – Canonical contract defining the scrape step's inputs, outputs, and schema constraints
- **`.agents/skills/*/cli.md`** – Portal-specific CLI wrappers that feed JSON data into the scraper
- **[`tools/job_scraper.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/job_scraper.py)** – Core implementation containing the `ScrapedJob` model and `run_scrape()` orchestration function
- **[`.claude/skills/candidate/skills.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/.claude/skills/candidate/skills.md)** – Persistent candidate skill list updated via the discovery mechanism
- **[`job_scraper/provenance.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/job_scraper/provenance.json)** – JSON storage for skill audit trails (source URLs and timestamps)
- **[`tests/test_scrape_contract.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tests/test_scrape_contract.py)** – Unit tests verifying compliance with the SKILL.md specification
- **[`tests/test_scrape_provenance.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tests/test_scrape_provenance.py)** – Unit tests validating provenance data integrity

## Usage Examples

Invoke the scraper programmatically to extract skills from current market postings:

```python
from tools.job_scraper import run_scrape

# Execute scrape across configured portals

scrape_results = run_scrape(search_query="machine learning engineer")

# Access discovered skills as a set of normalized strings

new_skills = scrape_results.discovered_skills
print("Newly discovered skills:", new_skills)

# Audit skill origins through provenance data

for skill, meta in scrape_results.provenance.items():
    print(f"{skill} → {meta['source_url']} @ {meta['timestamp']}")

```

Run via the command-line interface as defined in the portal CLI specifications:

```bash

# Trigger scrape for specific query

ai-job-search scrape --query "data scientist"

# Updates .claude/skills/candidate/skills.md with new discoveries

# and populates job_scraper/provenance.json with audit data

```

## Summary

- The `/scrape` command implements a specification-driven pipeline defined in [`.claude/skills/job-scraper/SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/.claude/skills/job-scraper/SKILL.md) to harvest skills from live job data.
- **Portal-specific CLIs** under `.agents/skills/` retrieve postings, while [`tools/job_scraper.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/job_scraper.py) normalizes and processes the JSON responses.
- An **NLP tokenization layer** normalizes synonyms and filters noise to generate canonical skill identifiers.
- Discovered skills are staged for integration into [`.claude/skills/candidate/skills.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/.claude/skills/candidate/skills.md) and tracked with full **provenance** in [`job_scraper/provenance.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/job_scraper/provenance.json).
- **Gating logic** prevents irrelevant postings from polluting the candidate's skill profile.
- The implementation is validated by [`tests/test_scrape_contract.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tests/test_scrape_contract.py) and [`tests/test_scrape_provenance.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tests/test_scrape_provenance.py) to ensure specification compliance.

## Frequently Asked Questions

### Where does /scrape store newly discovered skills before they are added to my profile?

Newly discovered skills are initially written to a temporary discoveries list within the scraper's internal state. The `/upskill` command subsequently reviews this list and merges validated skills into [`.claude/skills/candidate/skills.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/.claude/skills/candidate/skills.md), ensuring that the canonical candidate profile only grows through intentional integration.

### How does the scraper prevent duplicate skills from being recorded?

The scraper performs a deduplication check against the existing skill set in [`.claude/skills/candidate/skills.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/.claude/skills/candidate/skills.md) during the persistence stage. Additionally, the NLP normalization layer maps variant terms (such as "React.js" and "React") to canonical tokens before the duplication check occurs, preventing semantic duplicates rather than just string matches.

### What is skill provenance and why is it tracked?

**Provenance** refers to the metadata recording where each skill originated, including the job posting URL, portal name, and extraction timestamp. This data is stored in [`job_scraper/provenance.json`](https://github.com/MadsLorentzen/ai-job-search/blob/main/job_scraper/provenance.json) to provide auditability, allowing users to verify that skills were derived from legitimate market postings and enabling the `/rank` command to weight skills based on source authority or recency.

### How does /scrape filter out irrelevant skills from unrelated job postings?

The command implements a **gating system** that evaluates postings against configurable filters such as location, seniority level, and required experience. If a posting fails these gates—indicating it is outside the candidate's target scope—it is excluded from skill extraction entirely, ensuring the discovered skills remain relevant to the user's actual job search criteria.