Portal Search Skills in AI Job Search: How Declarative Specifications Drive Modular Job Scraping

Portal search skills are self‑contained Markdown specifications that define how to query a specific job portal and transform its raw output into a standardized format, enabling the scraper orchestrator to support any portal without core code changes.

In the MadsLorentzen/ai-job-search repository, portal search skills implement a thin‑pointer design pattern that keeps agent runtimes synchronized through shared, declarative configuration files. This architecture separates portal‑specific query logic from the core scraping pipeline, making the system highly extensible.


What Portal Search Skills Are

A portal search skill is a lightweight specification file—typically named SKILL.md—that lives under the .agents/skills/<portal-name>/ directory. Each skill acts as a contract between the job search workflow and a particular job listing source (LinkedIn, Indeed, StackOverflow Jobs, etc.).

According to the repository's agent architecture documentation in [AGENTS.md](https://github.com/MadsLorentzen/ai-job-search/blob/master/AGENTS.md), these skills are organized under:

  • .agents/skills/ – Portable skill definitions in the "Agent Skills" format
  • .claude/skills/job-scraper/ – The canonical scraper workflow that consumes these skills

Every skill file contains four essential components:

  1. CLI command or script – The executable that performs the actual search
  2. Argument schema – Required and optional parameters (keywords, location, pagination)
  3. Output mapping – Rules for parsing the portal's response into the canonical job object schema
  4. Metadata – Rate limits, authentication requirements, and other operational constraints

This design ensures that all agent implementations (Claude, Codex, Antigravity, or others) share the same definition, eliminating duplication and drift between runtimes.


How Portal Search Skills Work

The scraper orchestrator interacts with portal search skills through a three‑stage pipeline:

Stage 1: Skill Discovery and Loading

When the workflow needs to search a specific portal, it locates the corresponding skill directory and parses the SKILL.md file. The skill identifier is typically the portal name (lowercase, hyphenated).

from pathlib import Path
import yaml

def load_skill(skill_name: str) -> dict:
    skill_path = Path(f".agents/skills/{skill_name}/SKILL.md")
    # Parse frontmatter and Markdown body

    return parse_skill_markdown(skill_path.read_text())

Stage 2: Command Execution with Parameter Binding

The orchestrator extracts the CLI template from the skill and binds the workflow's search parameters to create an executable command.

Given this skill definition in .agents/skills/linkedin/SKILL.md:


# LinkedIn Job Search Skill

## CLI

```bash
linkedin-search --query "{keywords}" --location "{location}" --page {page}

Arguments

  • keywords – Search term(s) (string, required)
  • location – City or region (string, required)
  • page – Pagination index (integer, default 1)

Output Mapping

{
  "jobs": "$.results[*]",
  "title": "$.results[*].title",
  "company": "$.results[*].companyName",
  "location": "$.results[*].location",
  "url": "$.results[*].applyUrl",
  "postedDate": "$.results[*].datePosted"
}

Rate Limit

  • 10 requests per minute
  • Requires LINKEDIN_SESSION environment variable

The orchestrator formats and executes:

```python
import subprocess
import json

def execute_skill(skill: dict, params: dict) -> list:
    # Bind parameters to CLI template

    command = skill["cli"].format(**params)
    
    result = subprocess.run(
        command,
        shell=True,
        capture_output=True,
        text=True,
        timeout=skill.get("timeout", 30)
    )
    
    if result.returncode != 0:
        raise SkillExecutionError(result.stderr)
    
    # Parse JSON response

    raw_data = json.loads(result.stdout)
    
    # Apply output mapping to transform to canonical schema

    return apply_output_mapping(raw_data, skill["output_mapping"])

Stage 3: Normalization and Pipeline Integration

The output mapping uses JSONPath or similar selectors to extract fields from the portal's native response structure. This produces job objects that conform to the repository's internal schema, regardless of source:

from jsonpath_ng import parse

def apply_output_mapping(raw_data: dict, mapping: dict) -> list:
    jobs = []
    
    # Extract array of job entries

    job_entries = [match.value for match in parse(mapping["jobs"]).find(raw_data)]
    
    for entry in job_entries:
        job = {
            "title": extract_field(entry, mapping["title"]),
            "company": extract_field(entry, mapping["company"]),
            "location": extract_field(entry, mapping["location"]),
            "url": extract_field(entry, mapping["url"]),
            "postedDate": extract_field(entry, mapping["postedDate"]),
            "source": skill["name"],  # e.g., "linkedin"

        }
        jobs.append(job)
    
    return jobs

The normalized jobs flow into downstream stages—ranking, filtering, and application generation—entirely agnostic to their original portal.


Adding a New Portal Search Skill

The thin‑pointer architecture enables portal support without modifying core scraper code:

  1. Create the skill directory: mkdir .agents/skills/remoteok
  2. Write SKILL.md with CLI, arguments, output mapping, and metadata
  3. Test locally: python -m job_scraper --skill remoteok --params '{"keywords": "python"}'
  4. Commit and deploy – All runtimes automatically recognize the new portal

# RemoteOK Job Search Skill

## CLI

```bash
curl -s "https://remoteok.com/api?tag={keywords}&page={page}"

Arguments

  • keywords – Job tag or search term (string, required)
  • page – Page offset (integer, default 1)

Output Mapping

{
  "jobs": "$[*].",
  "title": "$[*].position",
  "company": "$[*].company",
  "location": "$[*].location or 'Remote'",
  "url": "$[*].url",
  "postedDate": "$[*].date"
}

Rate Limit

  • 1 request per 2 seconds
  • Respect X-RateLimit-Remaining header

---

## Summary

- **Portal search skills** are declarative Markdown specifications in `.agents/skills/<portal>/SKILL.md` that define how to query and parse job listings from specific sources.
- The **scraper orchestrator** loads skills dynamically, binds workflow parameters to CLI templates, executes commands, and applies output mappings to produce normalized job objects.
- **Thin‑pointer design** ensures all agent runtimes share identical skill definitions, eliminating duplication across Claude, Codex, and other implementations.
- **Zero core code changes** are required to add new portals—simply create a new skill directory with the appropriate specification.

---

## Frequently Asked Questions

### What file format are portal search skills written in?

Portal search skills are written in **Markdown with YAML frontmatter**. The [`SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SKILL.md) files use standard Markdown structure with fenced code blocks for CLI templates and JSON mappings, making them human-readable and parseable by any runtime. This format choice aligns with the repository's goal of being "readable, editable, and cross-runtime compatible" as specified in [`AGENTS.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/AGENTS.md).

### How does the scraper handle rate limits defined in skills?

The scraper extracts **rate limit metadata** from each skill's [`SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SKILL.md) and enforces them through a throttling layer. For example, a skill specifying "10 requests per minute" causes the orchestrator to queue or delay subsequent calls to that portal. Some skills also define headers like `X-RateLimit-Remaining` for dynamic adjustment, which the scraper monitors during execution.

### Can multiple portal search skills run in parallel?

Yes, the **canonical workflow** in `.claude/skills/job-scraper/` can dispatch searches across multiple portals concurrently because each skill is **self-contained and stateless**. The orchestrator manages thread pools or async tasks per portal, respecting individual rate limits while maximizing throughput. Job objects from all sources converge into a single normalized stream for downstream processing.

### What happens if a portal changes its API response format?

Only the affected **portal search skill** requires updating—specifically the `Output Mapping` section of its [`SKILL.md`](https://github.com/MadsLorentzen/ai-job-search/blob/main/SKILL.md). The core scraper logic remains unchanged. This isolation is the primary benefit of the thin‑pointer architecture: portal-specific adaptations are localized to declarative configuration files rather than scattered across Python/TypeScript implementations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →