How the AI Job Search Framework Auto-Discovers Installed Portal CLIs for Scraping
The AI Job Search Framework treats every job portal as a self-contained skill under .agents/skills/, automatically discovering installed CLIs by glob-matching SKILL.md files and parsing their frontmatter to extract entry points, flags, and enabled status without requiring manual registration.
The MadsLorentzen/ai-job-search repository implements a plug-and-play architecture for job scraping that eliminates boilerplate configuration. When users invoke the /scrape command, the framework's orchestrator dynamically locates and executes portal-specific CLIs by reading structured metadata files rather than relying on hardcoded registries. This auto-discovery mechanism enables seamless integration of new job portals without modifying core framework code.
The Skill-Based Architecture for Portal Integrations
Self-Contained Portal Skills
Each job portal integration resides as an isolated skill within the .agents/skills/ directory. Every skill includes a SKILL.md file that declares the CLI entry point, supported command-line flags, and an optional enabled: boolean flag that controls runtime visibility.
Core vs. Agent Skill Directories
The discovery mechanism searches across two distinct locations to build the complete skill inventory:
.claude/skills/— Core framework skills including thejob-scraperorchestrator itself.agents/skills/— User-installed or third-party portal integrations
How Auto-Discovery Works
Glob-Matching SKILL.md Files
When the /scrape command executes, the system initiates filesystem discovery by locating all SKILL.md manifests. According to tools/lint_skills.py at line 99, the framework builds the skill inventory using glob patterns that scan both core and agent directories:
skills = sorted(ROOT.glob(".claude/skills/*/SKILL.md")) + \
sorted(ROOT.glob(".agents/skills/*/SKILL.md"))
Parsing CLI Metadata and Enabled Flags
The scraper orchestrator reads each SKILL.md to extract the CLI invocation pattern and check the enabled status. As documented in .claude/skills/job-scraper/SKILL.md at line 61, the framework "discovers all installed portal CLI skills by reading every SKILL.md found under .agents/skills/*/SKILL.md."
Line 63 of the same file specifies that the framework "honors the enabled toggle," treating portals as enabled unless the frontmatter explicitly sets enabled: false. This allows developers to keep portal code installed while excluding it from execution.
Validating Against Security Guards
Before execution, discovered CLI patterns undergo validation against the allowed-tools registry defined in tools/security_guards.py (lines 46-51). This security layer ensures only explicitly permitted command patterns—such as bun run .agents/skills/linkedin-search/cli/src/cli.ts *—are eligible for subprocess execution.
Executing Discovered CLIs
For each enabled portal, the scraper constructs the final command using the bun run pattern defined in the portal's SKILL.md. The framework translates user search queries into portal-specific flags, executes the CLI via subprocess, and collects JSON output. Post-processing logic, such as client-side date filtering, applies as described in job-scraper/SKILL.md lines 68-73.
CI Pipeline Auto-Discovery
The discovery mechanism extends to continuous integration. The discover-clis job in .github/workflows/ci.yml (lines 185-211) automatically builds a test matrix by scanning for package.json files under each portal's cli/ directory. This ensures new portals undergo automated testing without requiring manual updates to the workflow configuration.
Implementation Examples
Listing Discovered Portal Skills
The /add-portal --list command leverages the same glob logic found in the linter to display available integrations:
from pathlib import Path
import yaml
ROOT = Path('.')
skill_files = sorted(ROOT.glob('.agents/skills/*/SKILL.md'))
for f in skill_files:
meta = yaml.safe_load(f.read_text())
print(f"{meta['name']}\t{meta.get('enabled', True)}")
Running a Discovered CLI
The scraper executes enabled portal CLIs by parsing the allowed-tools pattern from SKILL.md frontmatter:
import subprocess
import json
import pathlib
import yaml
def run_portal_cli(skill_md_path, query):
meta = yaml.safe_load(pathlib.Path(skill_md_path).read_text())
cli_pattern = meta['allowed-tools'][0]
# Extract command from pattern like "bun run path/to/cli.ts *"
cli_cmd = cli_pattern.split('(')[1].split(')')[0]
cmd = f"{cli_cmd} search --query \"{query}\" --format json".split()
result = subprocess.run(cmd, capture_output=True, text=True, check=True)
return json.loads(result.stdout)
# Aggregate postings from all enabled portals
for skill in discovered_skills:
if skill.get('enabled', True):
postings = run_portal_cli(skill['path'], user_query)
CI Matrix Generation
The GitHub Actions workflow dynamically constructs the test matrix using jq to scan for CLI entry points:
discover-clis:
runs-on: ubuntu-latest
outputs:
tools: ${{ steps.discover.outputs.tools }}
steps:
- name: Discover portal CLIs
id: discover
run: |
tools=$(jq -nc '
[paths(scalars) as $p | select($p | test("^\\.agents/skills/.+/cli/package\\.json$"))]
| map({"tool": "bun run " + ($p | join("/")) + " *"})')
echo "::set-output name=tools::$tools"
Key Files in the Discovery Pipeline
tools/lint_skills.py— Performs filesystem globbing for allSKILL.mdmanifests (line 99), establishing the canonical list of available skills..claude/skills/job-scraper/SKILL.md— Documents the auto-discovery protocol, theenabledflag behavior (lines 61-63), and post-processing requirements (lines 68-73).tools/security_guards.py— Validates CLI invocation patterns against an allowlist (lines 46-51), preventing unauthorized command execution..github/workflows/ci.yml— Contains thediscover-clisjob (lines 185-211) that auto-generates CI test matrices for portal CLIs..agents/skills/*/SKILL.md— Individual portal configurations (e.g.,linkedin-search/SKILL.md) defining entry points and capabilities.
Summary
- The framework implements zero-configuration portal discovery by glob-matching
SKILL.mdfiles across.claude/skills/and.agents/skills/directories. - Each portal's
SKILL.mdfrontmatter specifies the CLI entry point, supported flags, and an optionalenabledtoggle that defaults to true. - The
job-scraperskill orchestrates execution by validating commands againstsecurity_guards.pyand invoking CLIs via thebun runpattern. - Disabled portals are automatically skipped when their
SKILL.mdcontainsenabled: false, eliminating runtime overhead for inactive integrations. - The CI pipeline mirrors runtime discovery, automatically testing new portals by detecting
cli/package.jsonfiles without manual workflow updates.
Frequently Asked Questions
What file triggers the auto-discovery of a new portal CLI?
The presence of a SKILL.md file within an .agents/skills/<portal-name>/ directory triggers discovery. The framework specifically glob-matches **/*/SKILL.md patterns to identify candidate portals, then parses the YAML frontmatter to extract the CLI configuration.
How does the framework prevent disabled portals from running?
The scraper checks the enabled key in each SKILL.md's frontmatter before execution. As implemented in the discovery logic, portals default to enabled unless explicitly set to enabled: false, at which point they are omitted from the execution list entirely, incurring no runtime cost.
Can I add a new portal without modifying the core framework code?
Yes. Adding a new portal requires only creating a subdirectory under .agents/skills/ containing a properly formatted SKILL.md and the CLI implementation. No changes to tools/lint_skills.py, job-scraper/SKILL.md, or CI configuration are necessary—the framework discovers and integrates the new portal automatically upon the next /scrape invocation.
How does the CI pipeline know which portal CLIs to test?
The discover-clis job in .github/workflows/ci.yml independently scans the repository for package.json files within portal CLI directories. Using jq to construct a JSON matrix of discovered tools, the pipeline automatically includes new portals in the test suite without requiring manual updates to the workflow file.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →