Security Checks Performed Before Running a Third-Party Job Portal Skill

The AI-Job-Search framework enforces a mandatory two-layer validation system—verifying supply-chain integrity through tools/security_guards.py and web scraping permissions through tools/robots_check.py—before any third-party job portal skill executes.

The MadsLorentzen/ai-job-search repository implements rigorous security checks performed before running a third-party job portal skill to mitigate supply-chain attacks and ensure ethical web scraping. These safeguards analyze repository permissions, .gitignore rules, npm manifest files, and target site robots.txt policies. Only when both validation layers return exit code zero does the framework permit the actual scraper code to run.

Supply-Chain Integrity Validation

The first security layer resides in tools/security_guards.py, which validates three critical attack vectors before trusting any third-party code.

Permission Allowlist Verification

The script compares declared permissions in .claude/settings.json against a hard-coded ALLOWED_PERMISSIONS list. In lines 38-44 of security_guards.py, the check_permissions() function parses the settings file and aborts execution if any permission falls outside the pre-approved set. This prevents third-party skills from requesting unauthorized API access or filesystem privileges.

Gitignore Rule Enforcement

The framework mandates specific privacy protections through .gitignore. The check_gitignore() function (lines 49-84) iterates over REQUIRED_IGNORE_RULES to ensure sensitive personal data patterns remain excluded from version control. Additionally, lines 96-103 scan for dangerous negation rules (e.g., !sensitive-data.json) that could accidentally expose private information, requiring explicit approval for any exceptions.

Package Manifest Scrutiny

Every package.json under .agents/** undergoes inspection for forbidden lifecycle scripts. Lines 22-30 of the guard glob all manifests (excluding node_modules) and reject any containing preinstall, install, postinstall, or trustedDependencies entries. This blocks supply-chain attacks where malicious npm hooks might execute during installation.

Robots.txt Compliance Check

The second security layer, implemented in tools/robots_check.py, ensures the framework respects the target portal's crawling policies per RFC 9309.

RFC 9309 Parsing and Longest-Match Logic

The gate() function constructs the target robots.txt URL (line 14) and retrieves the file using curl to avoid urllib hangs. Lines 16-24 fetch the policy for two user-agents: Claude-User and a generic browser wildcard. The helper allowed() function (lines 99-108) applies a longest-match wins algorithm, giving precedence to Disallow directives when path lengths tie, ensuring conservative access controls prevail.

User-Agent Evaluation

The framework explicitly blocks execution if robots.txt contains a Disallow rule matching either the Claude-User specific agent or the universal * wildcard. Only explicit or implicit permission in the parsed policy allows the crawl to proceed. The script returns exit code 0 for allowed requests and exit code 1 for blocked attempts, with explanatory messages written to stdout (lines 30-34).

Execution Flow and Integration

When invoking a portal skill such as linkedin-search, the command dispatcher executes validation sequentially:

  1. security_guards.py validates repository integrity, permissions, and package manifests.
  2. robots_check.py verifies the target URL permits crawling under RFC 9309 rules.
  3. Skill execution proceeds only if both scripts return success codes.

If either guard detects violations—unapproved permissions, missing .gitignore rules, forbidden npm scripts, or restrictive robots.txt entries—the framework aborts immediately and prints diagnostic details. Contributors must remediate issues (e.g., adding permissions to ALLOWED_PERMISSIONS or respecting crawl delays) before retrying.

Practical Usage Examples

Run the supply-chain guard manually to validate repository integrity:

$ python tools/security_guards.py
security_guards: OK (permissions allowlist, hooks allowlist, gitignore rules, package manifests)

Verify crawling permissions for a specific job portal URL:

$ python tools/robots_check.py https://jobs.example.com/remote-developer
ALLOWED - robots.txt permits this path
$ echo $?
0  # Exit code 0 confirms safe execution

Execute a third-party skill only after both guards succeed:

$ python -m agents.skills.linkedin-search.cli.main "software engineer"

If robots_check.py returns exit code 1, the skill aborts before sending any HTTP requests to the target domain.

Summary

  • Dual-layer validation combines supply-chain verification (security_guards.py) and web ethics compliance (robots_check.py).
  • Permission allowlists in .claude/settings.json are hard-coded and validated against required patterns in lines 38-44 of the guard script.
  • Gitignore enforcement prevents accidental commits of personal data by checking required rules (lines 49-84) and flagging dangerous negations (lines 96-103).
  • Npm manifest scanning blocks packages with forbidden lifecycle scripts (preinstall, install, postinstall, trustedDependencies) across all .agents/**/package.json files.
  • RFC 9309 compliance ensures robots.txt respects longest-match logic and explicitly permits the Claude-User agent before any crawling occurs.
  • Zero-trust execution: Both guards must return exit code zero; any failure halts the job portal skill immediately.

Frequently Asked Questions

What happens if a third-party skill requests permissions not in the allowlist?

The check_permissions() function in tools/security_guards.py (lines 38-44) compares the requested permissions against ALLOWED_PERMISSIONS. If any disparity exists, the script exits with a non-zero status and prints the offending permissions, forcing manual review before execution resumes.

How does the framework prevent malicious npm packages from executing during skill installation?

The check_package_manifests() function scans every package.json under .agents/** (excluding node_modules) for forbidden lifecycle hooks. Lines 22-30 explicitly reject any manifest containing preinstall, install, postinstall, or trustedDependencies fields, preventing supply-chain attacks that rely on installation-time script execution.

Why does the robots.txt checker use curl instead of Python's urllib?

The gate() function in tools/robots_check.py invokes curl rather than urllib to avoid hangs and timeouts common with Python's standard library when handling malformed or slow-responding robots.txt endpoints. This ensures the security check completes rapidly without blocking the execution pipeline.

Can the security guards be bypassed when running a job portal skill locally?

No. According to the MadsLorentzen/ai-job-search source code, the command dispatcher internally invokes both security_guards.py and robots_check.py before loading any skill modules. Manual execution of the guards is provided for pre-flight checks, but the framework enforces these validations automatically regardless of how the skill is invoked.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →