How /scrape Auto-Discovers Job Portal Skills in the AI Job Search Framework

/scrape dynamically discovers job portal search capabilities by scanning the .agents/skills/ directory for SKILL.md files, reading their YAML frontmatter to extract CLI commands and enabled status, and executing valid portals in parallel while falling back to generic WebSearch for unavailable ones.

The AI Job Search framework eliminates manual configuration for job portal integrations by implementing a data-driven discovery system. When invoked, the /scrape command automatically locates and executes search capabilities defined in modular skill definitions, ensuring the scraper stays synchronized with newly added portals without requiring code changes or registry updates.

Scanning the Skills Directory for Portal Definitions

The discovery process initiates with a filesystem glob targeting the pattern **/.agents/skills/*/SKILL.md. This pattern captures every portal skill definition within the repository's .agents/skills/ directory structure, where each subdirectory represents a distinct job portal such as linkedin-search.

// Discovery logic scans for portal definitions at runtime
const skillPaths = await glob('**/.agents/skills/*/SKILL.md');

for (const path of skillPaths) {
  const skill = parseYamlFrontmatter(path);
  if (skill.enabled === false) continue;
  
  // Execute the verbatim CLI command from frontmatter
  const results = await agent.run(skill.cli, searchArgs);
}

Because the scanner reads these definitions at runtime, any portal added via the /add-portal scaffolding command is recognized immediately without requiring framework restarts or configuration reloads.

Parsing SKILL.md Frontmatter and Validation

Each SKILL.md file contains YAML frontmatter that controls portal activation and execution. The frontmatter specifies an enabled boolean flag (defaulting to true) and the exact CLI invocation string that /scrape must execute.

Portals marked with enabled: false are excluded from the current scrape operation. For active portals, the cli field provides the verbatim command—typically formatted as bun run .agents/skills/{portal-name}/cli/src/cli.ts search. The /scrape implementation uses this string exactly as written, never inferring or modifying flags, ensuring consistent behavior across different portal implementations.

---
name: linkedin-search
description: Search LinkedIn job listings
enabled: true
cli: "bun run .agents/skills/linkedin-search/cli/src/cli.ts search"
---

Executing Portal CLIs with the JSON Contract

For each enabled portal, /scrape executes the extracted CLI command via the Agent tool, passing search query arguments as parameters. These commands run in parallel to maximize throughput across multiple job portals.

Every portal CLI must output JSON when invoked with the --format json flag. The output must include five mandatory fields: title, company, location, date, and url. This contract is enforced by the test suite in tests/test_scrape_contract.py, which verifies that each portal CLI emits the required schema. The test uses a regex pattern to assert that search outputs contain these specific fields, maintaining data integrity across the aggregation pipeline.

Fallback to WebSearch for Resilient Operation

The framework implements graceful degradation when portal-specific skills encounter issues. If a portal lacks a SKILL.md definition, if the extracted CLI command fails during execution, or if the bun runtime is not installed on the system, /scrape automatically falls back to a generic WebSearch implementation for that specific portal.

This fallback mechanism ensures continuous job search coverage even when individual portal integrations are incomplete, encounter technical errors, or run in environments missing required dependencies.

Extending the Framework with New Portals

The auto-discovery architecture enables zero-configuration extension of the scraping capabilities. When developers create a new portal skill, placing a SKILL.md file in .agents/skills/{new-portal}/ makes it immediately available to /scrape on the next invocation.

As documented in .claude/skills/job-scraper/SKILL.md, the system treats portal skills as data rather than compiled code. This design guarantees that the scraper stays in sync with newly scaffolded portal skills without requiring updates to a central registry or routing table.

Summary

  • /scrape uses glob patterns to scan .agents/skills/*/SKILL.md files for portal definitions at runtime.
  • Each SKILL.md contains YAML frontmatter with an enabled flag and verbatim CLI command strings extracted verbatim for execution.
  • Enabled portals execute in parallel via the Agent tool, with CLIs outputting JSON containing title, company, location, date, and url.
  • The system falls back to generic WebSearch when portal CLIs are missing, disabled, or fail to execute.
  • New portals are auto-discovered immediately upon file creation without registry updates or application restarts.

Frequently Asked Questions

What file pattern does /scrape use to find job portal skills?

/scrape searches for files matching the pattern **/.agents/skills/*/SKILL.md within the repository. This glob pattern captures all portal skill definitions located in subdirectories under .agents/skills/, where each subdirectory represents a specific job portal integration such as linkedin-search.

How does the framework determine which portals to skip during execution?

The framework parses the YAML frontmatter in each SKILL.md file and evaluates the enabled boolean field. If enabled is explicitly set to false, the portal is excluded from the current scrape operation. When the field is omitted, it defaults to true, ensuring portals are active unless explicitly disabled.

What data format must portal CLIs output to satisfy the scrape contract?

Each portal CLI must output JSON containing exactly five fields: title, company, location, date, and url. The CLI receives the --format json flag to trigger structured output mode. The test suite in tests/test_scrape_contract.py validates that all portal CLIs adhere to this contract before integration.

What happens if a portal CLI is not installed or fails during execution?

When a portal CLI fails to execute, returns non-zero exit codes, or when the bun runtime is unavailable on the host system, /scrape automatically falls back to a generic WebSearch implementation for that specific portal. This ensures the job search remains comprehensive even when specific portal integrations encounter technical difficulties or environment constraints.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →