# How to Configure Company and Query Settings for the Career-Ops Job Scanner

> Learn to configure career ops job scanner company and query settings by editing portals.yml. Add tracked companies and search queries then run node scan.mjs for efficient job searching.

- Repository: [Santiago Fernández de Valderrama/career-ops](https://github.com/santifer/career-ops)
- Tags: how-to-guide
- Published: 2026-08-28

---

**Configure your job search by editing [`portals.yml`](https://github.com/santifer/career-ops/blob/main/portals.yml): add entries to `tracked_companies` for direct company scraping and define queries in `search_queries` for web search discovery, then run `node scan.mjs` to execute.**

The Career-Ops job scanner from the `santifer/career-ops` repository centralizes all search parameters in a single YAML configuration file. By modifying [`portals.yml`](https://github.com/santifer/career-ops/blob/main/portals.yml) (copied from [`templates/portals.example.yml`](https://github.com/santifer/career-ops/blob/main/templates/portals.example.yml)), you control which companies are crawled directly and which job boards are searched via web queries. This guide walks through the specific configuration keys and file structures defined in the source code.

## Understanding the portals.yml Structure

The scanner reads three main sections from [`portals.yml`](https://github.com/santifer/career-ops/blob/main/portals.yml) to determine its behavior. According to [`templates/portals.example.yml`](https://github.com/santifer/career-ops/blob/main/templates/portals.example.yml), the **`tracked_companies`** block (lines 24-31) handles direct career page scraping, while **`search_queries`** (lines 22-23) manages web search discovery. Post-fetch filters like **`title_filter`** (lines 86-93) and **`location_filter`** (lines 82-99) prune results based on your criteria after retrieval.

## Configuring Tracked Companies for Direct Scanning

### Basic Company Entry Structure

Each entry under the `tracked_companies:` list requires a `name` identifier and a `careers_url` pointing to the company's public careers page. The `enabled` boolean allows you to temporarily skip entries without deleting them.

```yaml
tracked_companies:
  - name: Acme Corp
    careers_url: https://acme.com/careers
    enabled: true

```

### Selecting the Scan Method

The **`scan_method`** parameter determines how the scanner interacts with the careers page. As implemented in the configuration schema, three options are available:

- **`playwright`** (default): Uses browser automation to render JavaScript-heavy pages
- **`local_parser`**: Executes a custom script you provide that outputs JSON
- **`websearch`**: Falls back to the `search_queries` block for that company

```yaml
tracked_companies:
  - name: Acme Corp
    careers_url: https://acme.com/careers
    scan_method: playwright
    enabled: true

```

### Overriding Provider Detection

The scanner auto-detects the Applicant Tracking System (ATS) provider from the URL structure (logic defined around lines 66-71 in [`templates/portals.example.yml`](https://github.com/santifer/career-ops/blob/main/templates/portals.example.yml)). If auto-detection fails for branded domains, force a specific provider using the **`provider`** key:

```yaml
tracked_companies:
  - name: Acme Corp
    careers_url: https://jobs.acme.com
    provider: smartrecruiters
    enabled: true

```

The supported providers are listed in the provider modules under `providers/*.mjs` and referenced in the configuration comments at lines 66-91.

### Implementing Local Parsers

When using **`scan_method: local_parser`**, you must specify a parser configuration under the **`parser:`** key. The parser must print a JSON array of objects containing `title`, `url`, and `location` fields to stdout (see the specification around lines 38-45).

```yaml
tracked_companies:
  - name: Foo Ltd
    careers_url: https://foo.com/careers
    scan_method: local_parser
    parser:
      command: node
      script: scripts/parsers/foo-jobs.js
      format: jobs-json-v1
    enabled: true

```

Place your custom parser scripts in the `scripts/parsers/` directory. The `max_pages` parameter (default 50) controls pagination depth for listings.

## Setting Up Search Queries for Web Discovery

### Query Structure and Syntax

The **`search_queries`** section starts around line 22 in the example file. Each entry requires three fields: `name`, `query`, and `enabled`. The query string uses standard Google search syntax, typically including a `site:` filter to limit results to specific job boards.

```yaml
search_queries:
  - name: Wellfound — AI Engineer
    query: 'site:wellfound.com "AI Engineer" ("visa sponsorship" OR "H-1B")'
    enabled: true
  - name: Indeed — Data Science
    query: 'site:indeed.com "Data Scientist" remote'
    enabled: true

```

### Enabling and Disabling Queries

Toggle individual queries using the **`enabled`** boolean. This allows you to maintain a library of search templates while activating only those relevant to your current search without modifying the query strings.

```yaml
search_queries:
  - name: Kariyer.net — Backend Engineer
    query: 'site:kariyer.net "Backend Engineer" remote'
    enabled: false

```

## How the Scanner Processes Your Configuration

The main scanner script **`scan.mjs`** reads [`portals.yml`](https://github.com/santifer/career-ops/blob/main/portals.yml) and executes in two distinct phases:

1. **Company-level scan**: For each entry in `tracked_companies`, the scanner attempts to fetch listings using either the specified `provider` or the auto-detected provider from `providers/*.mjs`. If `scan_method` is set to `local_parser`, the configured script executes and must output valid JSON jobs.

2. **WebSearch fallback**: After completing company scans, the scanner processes each enabled entry in `search_queries`, sending queries to the Google search endpoint and parsing results into job objects.

Finally, the scanner applies post-fetch filters defined in `title_filter` and `location_filter` (lines 82-99) to remove irrelevant postings before output.

## Summary

- **Primary configuration** happens in [`portals.yml`](https://github.com/santifer/career-ops/blob/main/portals.yml), copied from [`templates/portals.example.yml`](https://github.com/santifer/career-ops/blob/main/templates/portals.example.yml)
- **Direct company scanning** uses the `tracked_companies` list with mandatory `careers_url` and optional `scan_method` and `provider` overrides
- **Job board discovery** uses the `search_queries` list with Google-style `site:` filters
- **Custom parsers** must output JSON arrays with `title`, `url`, and `location` fields when using `scan_method: local_parser`
- **Execution** occurs via `node scan.mjs`, which reads the YAML and orchestrates providers and filters

## Frequently Asked Questions

### What file do I edit to configure the job scanner?

Edit **[`portals.yml`](https://github.com/santifer/career-ops/blob/main/portals.yml)** in the repository root. This file is a user-editable copy of **[`templates/portals.example.yml`](https://github.com/santifer/career-ops/blob/main/templates/portals.example.yml)**, which serves as the reference documentation with detailed comments. The scanner reads [`portals.yml`](https://github.com/santifer/career-ops/blob/main/portals.yml) at runtime to determine which companies to crawl and which web searches to execute.

### How do I add a company that uses a unique careers page URL?

Add an entry under `tracked_companies` with the `careers_url` pointing to the company's job listings page. If the scanner cannot auto-detect the ATS provider from the URL (as defined in lines 66-71 of the example file), explicitly set the `provider` key to one of the supported providers like `greenhouse`, `lever`, or `workday`.

### Can I use a custom script to parse a company's job listings?

Yes. Set **`scan_method: local_parser`** and define a **`parser:`** block specifying the command, script path, and format. Your script must output a JSON array of job objects containing `title`, `url`, and `location` keys to stdout. Store these scripts in `scripts/parsers/` and reference them from the [`portals.yml`](https://github.com/santifer/career-ops/blob/main/portals.yml) configuration.

### Where does the scanner apply filters to remove unwanted jobs?

The scanner applies **post-fetch filters** after retrieving all jobs from both company scans and web searches. The **`title_filter`** (lines 86-93) and **`location_filter`** (lines 82-99) defined in [`portals.yml`](https://github.com/santifer/career-ops/blob/main/portals.yml) prune the results list before final output, allowing you to exclude postings by keywords or geographic restrictions.