How to Configure Company and Query Settings for the Career-Ops Job Scanner
Configure your job search by editing portals.yml: add entries to tracked_companies for direct company scraping and define queries in search_queries for web search discovery, then run node scan.mjs to execute.
The Career-Ops job scanner from the santifer/career-ops repository centralizes all search parameters in a single YAML configuration file. By modifying portals.yml (copied from templates/portals.example.yml), you control which companies are crawled directly and which job boards are searched via web queries. This guide walks through the specific configuration keys and file structures defined in the source code.
Understanding the portals.yml Structure
The scanner reads three main sections from portals.yml to determine its behavior. According to templates/portals.example.yml, the tracked_companies block (lines 24-31) handles direct career page scraping, while search_queries (lines 22-23) manages web search discovery. Post-fetch filters like title_filter (lines 86-93) and location_filter (lines 82-99) prune results based on your criteria after retrieval.
Configuring Tracked Companies for Direct Scanning
Basic Company Entry Structure
Each entry under the tracked_companies: list requires a name identifier and a careers_url pointing to the company's public careers page. The enabled boolean allows you to temporarily skip entries without deleting them.
tracked_companies:
- name: Acme Corp
careers_url: https://acme.com/careers
enabled: true
Selecting the Scan Method
The scan_method parameter determines how the scanner interacts with the careers page. As implemented in the configuration schema, three options are available:
playwright(default): Uses browser automation to render JavaScript-heavy pageslocal_parser: Executes a custom script you provide that outputs JSONwebsearch: Falls back to thesearch_queriesblock for that company
tracked_companies:
- name: Acme Corp
careers_url: https://acme.com/careers
scan_method: playwright
enabled: true
Overriding Provider Detection
The scanner auto-detects the Applicant Tracking System (ATS) provider from the URL structure (logic defined around lines 66-71 in templates/portals.example.yml). If auto-detection fails for branded domains, force a specific provider using the provider key:
tracked_companies:
- name: Acme Corp
careers_url: https://jobs.acme.com
provider: smartrecruiters
enabled: true
The supported providers are listed in the provider modules under providers/*.mjs and referenced in the configuration comments at lines 66-91.
Implementing Local Parsers
When using scan_method: local_parser, you must specify a parser configuration under the parser: key. The parser must print a JSON array of objects containing title, url, and location fields to stdout (see the specification around lines 38-45).
tracked_companies:
- name: Foo Ltd
careers_url: https://foo.com/careers
scan_method: local_parser
parser:
command: node
script: scripts/parsers/foo-jobs.js
format: jobs-json-v1
enabled: true
Place your custom parser scripts in the scripts/parsers/ directory. The max_pages parameter (default 50) controls pagination depth for listings.
Setting Up Search Queries for Web Discovery
Query Structure and Syntax
The search_queries section starts around line 22 in the example file. Each entry requires three fields: name, query, and enabled. The query string uses standard Google search syntax, typically including a site: filter to limit results to specific job boards.
search_queries:
- name: Wellfound — AI Engineer
query: 'site:wellfound.com "AI Engineer" ("visa sponsorship" OR "H-1B")'
enabled: true
- name: Indeed — Data Science
query: 'site:indeed.com "Data Scientist" remote'
enabled: true
Enabling and Disabling Queries
Toggle individual queries using the enabled boolean. This allows you to maintain a library of search templates while activating only those relevant to your current search without modifying the query strings.
search_queries:
- name: Kariyer.net — Backend Engineer
query: 'site:kariyer.net "Backend Engineer" remote'
enabled: false
How the Scanner Processes Your Configuration
The main scanner script scan.mjs reads portals.yml and executes in two distinct phases:
-
Company-level scan: For each entry in
tracked_companies, the scanner attempts to fetch listings using either the specifiedprovideror the auto-detected provider fromproviders/*.mjs. Ifscan_methodis set tolocal_parser, the configured script executes and must output valid JSON jobs. -
WebSearch fallback: After completing company scans, the scanner processes each enabled entry in
search_queries, sending queries to the Google search endpoint and parsing results into job objects.
Finally, the scanner applies post-fetch filters defined in title_filter and location_filter (lines 82-99) to remove irrelevant postings before output.
Summary
- Primary configuration happens in
portals.yml, copied fromtemplates/portals.example.yml - Direct company scanning uses the
tracked_companieslist with mandatorycareers_urland optionalscan_methodandprovideroverrides - Job board discovery uses the
search_querieslist with Google-stylesite:filters - Custom parsers must output JSON arrays with
title,url, andlocationfields when usingscan_method: local_parser - Execution occurs via
node scan.mjs, which reads the YAML and orchestrates providers and filters
Frequently Asked Questions
What file do I edit to configure the job scanner?
Edit portals.yml in the repository root. This file is a user-editable copy of templates/portals.example.yml, which serves as the reference documentation with detailed comments. The scanner reads portals.yml at runtime to determine which companies to crawl and which web searches to execute.
How do I add a company that uses a unique careers page URL?
Add an entry under tracked_companies with the careers_url pointing to the company's job listings page. If the scanner cannot auto-detect the ATS provider from the URL (as defined in lines 66-71 of the example file), explicitly set the provider key to one of the supported providers like greenhouse, lever, or workday.
Can I use a custom script to parse a company's job listings?
Yes. Set scan_method: local_parser and define a parser: block specifying the command, script path, and format. Your script must output a JSON array of job objects containing title, url, and location keys to stdout. Store these scripts in scripts/parsers/ and reference them from the portals.yml configuration.
Where does the scanner apply filters to remove unwanted jobs?
The scanner applies post-fetch filters after retrieving all jobs from both company scans and web searches. The title_filter (lines 86-93) and location_filter (lines 82-99) defined in portals.yml prune the results list before final output, allowing you to exclude postings by keywords or geographic restrictions.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →