How to Configure Company and Query Settings for the Career-Ops Job Scanner

Configure your job search by editing portals.yml: add entries to tracked_companies for direct company scraping and define queries in search_queries for web search discovery, then run node scan.mjs to execute.

The Career-Ops job scanner from the santifer/career-ops repository centralizes all search parameters in a single YAML configuration file. By modifying portals.yml (copied from templates/portals.example.yml), you control which companies are crawled directly and which job boards are searched via web queries. This guide walks through the specific configuration keys and file structures defined in the source code.

Understanding the portals.yml Structure

The scanner reads three main sections from portals.yml to determine its behavior. According to templates/portals.example.yml, the tracked_companies block (lines 24-31) handles direct career page scraping, while search_queries (lines 22-23) manages web search discovery. Post-fetch filters like title_filter (lines 86-93) and location_filter (lines 82-99) prune results based on your criteria after retrieval.

Configuring Tracked Companies for Direct Scanning

Basic Company Entry Structure

Each entry under the tracked_companies: list requires a name identifier and a careers_url pointing to the company's public careers page. The enabled boolean allows you to temporarily skip entries without deleting them.

tracked_companies:
  - name: Acme Corp
    careers_url: https://acme.com/careers
    enabled: true

Selecting the Scan Method

The scan_method parameter determines how the scanner interacts with the careers page. As implemented in the configuration schema, three options are available:

  • playwright (default): Uses browser automation to render JavaScript-heavy pages
  • local_parser: Executes a custom script you provide that outputs JSON
  • websearch: Falls back to the search_queries block for that company
tracked_companies:
  - name: Acme Corp
    careers_url: https://acme.com/careers
    scan_method: playwright
    enabled: true

Overriding Provider Detection

The scanner auto-detects the Applicant Tracking System (ATS) provider from the URL structure (logic defined around lines 66-71 in templates/portals.example.yml). If auto-detection fails for branded domains, force a specific provider using the provider key:

tracked_companies:
  - name: Acme Corp
    careers_url: https://jobs.acme.com
    provider: smartrecruiters
    enabled: true

The supported providers are listed in the provider modules under providers/*.mjs and referenced in the configuration comments at lines 66-91.

Implementing Local Parsers

When using scan_method: local_parser, you must specify a parser configuration under the parser: key. The parser must print a JSON array of objects containing title, url, and location fields to stdout (see the specification around lines 38-45).

tracked_companies:
  - name: Foo Ltd
    careers_url: https://foo.com/careers
    scan_method: local_parser
    parser:
      command: node
      script: scripts/parsers/foo-jobs.js
      format: jobs-json-v1
    enabled: true

Place your custom parser scripts in the scripts/parsers/ directory. The max_pages parameter (default 50) controls pagination depth for listings.

Setting Up Search Queries for Web Discovery

Query Structure and Syntax

The search_queries section starts around line 22 in the example file. Each entry requires three fields: name, query, and enabled. The query string uses standard Google search syntax, typically including a site: filter to limit results to specific job boards.

search_queries:
  - name: Wellfound — AI Engineer
    query: 'site:wellfound.com "AI Engineer" ("visa sponsorship" OR "H-1B")'
    enabled: true
  - name: Indeed — Data Science
    query: 'site:indeed.com "Data Scientist" remote'
    enabled: true

Enabling and Disabling Queries

Toggle individual queries using the enabled boolean. This allows you to maintain a library of search templates while activating only those relevant to your current search without modifying the query strings.

search_queries:
  - name: Kariyer.net — Backend Engineer
    query: 'site:kariyer.net "Backend Engineer" remote'
    enabled: false

How the Scanner Processes Your Configuration

The main scanner script scan.mjs reads portals.yml and executes in two distinct phases:

  1. Company-level scan: For each entry in tracked_companies, the scanner attempts to fetch listings using either the specified provider or the auto-detected provider from providers/*.mjs. If scan_method is set to local_parser, the configured script executes and must output valid JSON jobs.

  2. WebSearch fallback: After completing company scans, the scanner processes each enabled entry in search_queries, sending queries to the Google search endpoint and parsing results into job objects.

Finally, the scanner applies post-fetch filters defined in title_filter and location_filter (lines 82-99) to remove irrelevant postings before output.

Summary

  • Primary configuration happens in portals.yml, copied from templates/portals.example.yml
  • Direct company scanning uses the tracked_companies list with mandatory careers_url and optional scan_method and provider overrides
  • Job board discovery uses the search_queries list with Google-style site: filters
  • Custom parsers must output JSON arrays with title, url, and location fields when using scan_method: local_parser
  • Execution occurs via node scan.mjs, which reads the YAML and orchestrates providers and filters

Frequently Asked Questions

What file do I edit to configure the job scanner?

Edit portals.yml in the repository root. This file is a user-editable copy of templates/portals.example.yml, which serves as the reference documentation with detailed comments. The scanner reads portals.yml at runtime to determine which companies to crawl and which web searches to execute.

How do I add a company that uses a unique careers page URL?

Add an entry under tracked_companies with the careers_url pointing to the company's job listings page. If the scanner cannot auto-detect the ATS provider from the URL (as defined in lines 66-71 of the example file), explicitly set the provider key to one of the supported providers like greenhouse, lever, or workday.

Can I use a custom script to parse a company's job listings?

Yes. Set scan_method: local_parser and define a parser: block specifying the command, script path, and format. Your script must output a JSON array of job objects containing title, url, and location keys to stdout. Store these scripts in scripts/parsers/ and reference them from the portals.yml configuration.

Where does the scanner apply filters to remove unwanted jobs?

The scanner applies post-fetch filters after retrieving all jobs from both company scans and web searches. The title_filter (lines 86-93) and location_filter (lines 82-99) defined in portals.yml prune the results list before final output, allowing you to exclude postings by keywords or geographic restrictions.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →