# How to Set Up and Customize Portal Scanning for Greenhouse, Ashby, and Lever APIs

> Learn to set up and customize portal scanning for Greenhouse, Ashby, and Lever APIs with Career-Ops. Effortlessly fetch job data using the central portals.yml configuration.

- Repository: [Santiago Fernández de Valderrama/career-ops](https://github.com/santifer/career-ops)
- Tags: how-to-guide
- Published: 2026-08-19

---

**Career-Ops uses a plugin-based scanner that loads provider modules from `providers/*.mjs` to auto-detect and fetch jobs from Greenhouse, Ashby, and Lever APIs, with all customization controlled through a central [`portals.yml`](https://github.com/santifer/career-ops/blob/main/portals.yml) configuration file.**

The `career-ops` repository provides an open-source job aggregation system that discovers postings through a modular provider architecture. You can set up and customize portal scanning for Greenhouse, Ashby, and Lever APIs by editing the declarative [`portals.yml`](https://github.com/santifer/career-ops/blob/main/portals.yml) file and leveraging the built-in filtering system implemented in `scan.mjs`.

## Understanding the Provider Architecture

The core scanning logic resides in `scan.mjs`, which dynamically loads provider implementations from the `providers/` directory via `providers/_registry.mjs`. Each provider is a JavaScript module that exports an `id` and implements two required functions:

- `detect(entry)`: Analyzes the `careers_url` to determine the appropriate API endpoint
- `fetch(entry, ctx)`: Executes the HTTP request to the provider's API and returns normalized job objects with `{title, url, company, location, postedAt}`

This architecture allows the scanner to support multiple applicant tracking systems through a standardized interface while handling each provider's unique URL patterns and response formats.

## Configuring Portal Entries in portals.yml

All company configurations live in [`portals.yml`](https://github.com/santifer/career-ops/blob/main/portals.yml) at the repository root. Each entry specifies which provider handles the job board and how to access it.

### Basic Structure and Provider Selection

```yaml
- name: ExampleCo
  careers_url: https://exampleco.com/careers
  provider: greenhouse  # Options: greenhouse, ashby, lever

```

The `provider` field must match the `id` exported by the corresponding file in `providers/greenhouse.mjs`, `providers/ashby.mjs`, or `providers/lever.mjs`.

### Explicit API Endpoints

To bypass auto-detection, specify the API URL directly:

```yaml
- name: ExampleCo
  careers_url: https://exampleco.com/careers
  provider: greenhouse
  api: https://boards-api.greenhouse.io/v1/boards/exampleco/jobs

```

When the `api` field is present, the scanner skips the `detect()` function and uses the provided URL.

## Customizing Job Filters

The scanner provides three filtering layers defined in `scan.mjs` that process postings before storage.

### Title Filtering

The `title_filter` field uses the `compileKeyword` and `compilePositiveKeyword` helpers (defined in `scan.mjs` lines 4-11) to match role titles. Positive keywords support Boolean logic (e.g., `"engineer + senior"`), while negative keywords veto postings regardless of other matches.

```yaml
title_filter:
  positive: ["engineer + senior", "data scientist"]
  negative: ["intern", "contract"]

```

Short keywords (2-3 letters) automatically match on word boundaries to prevent false positives.

### Location Filtering

Location filters in `scan.mjs` (lines 38-62) process the `location_filter` block, which supports three rules:

- `always_allow`: Permits matching locations regardless of other rules
- `block`: Excludes locations containing these terms
- `allow`: Requires locations to contain at least one of these terms

The system uses look-around regex patterns so "india" does not match "Indianapolis".

```yaml
location_filter:
  always_allow: ["remote"]
  block: ["china", "russia"]
  allow: ["united states", "canada"]

```

### Content Filtering

For providers that return job descriptions, `content_filter` (implemented in `buildContentFilter`, lines 98-126 of `scan.mjs`) enables keyword matching within the posting body. You can also scope rules to specific title keywords using `by_title_keyword`.

```yaml
content_filter:
  positive: ["python", "aws"]
  negative: ["php"]
  by_title_keyword:
    "data scientist":
      positive: ["tensorflow", "pytorch"]
      negative: ["excel"]

```

## Provider-Specific Setup Details

Each ATS provider handles API discovery and data normalization differently.

### Greenhouse Configuration

The Greenhouse provider (`providers/greenhouse.mjs`) detects API URLs from board pages using `resolveApiUrl` (lines 28-37). It also handles location enrichment: when a posting lists only a work model like "Hybrid", the scanner fetches the `/offices` endpoint (via `isWorkModelOnly`, `officesUrlFor`, and `buildOfficeMap`, lines 48-77) to retrieve actual office locations.

### Ashby Configuration

The Ashby provider (`providers/ashby.mjs`) uses the public `/jobs` endpoint without requiring authentication. It follows the standard `detect()` → `fetch()` pattern without additional enrichment steps.

### Lever Configuration

The Lever provider (`providers/lever.mjs`) supports both explicit `api` URLs and automatic detection from Lever careers pages. It extracts the JSON job list and normalizes fields to match the standard schema.

## Running the Scanner

Execute the scanner from the repository root using Node.js:

```bash

# Scan all companies defined in portals.yml

node scan.mjs

# Test a single company

node scan.mjs --company ExampleCo

# Preview changes without writing to database

node scan.mjs --dry-run

```

CLI flag parsing is implemented at the top of `scan.mjs`.

## Example Configuration for All Three Providers

```yaml
- name: GreenTech
  careers_url: https://greentech.com/careers
  provider: greenhouse
  title_filter:
    positive: ["engineer + senior"]
  location_filter:
    allow: ["united states", "canada"]
  content_filter:
    positive: ["nodejs", "kubernetes"]
    negative: ["php"]

- name: AshbyAI
  careers_url: https://jobs.ashby.ai/ashbyai
  provider: ashby
  title_filter:
    positive: ["data scientist"]
    negative: ["intern"]
  location_filter:
    block: ["remote"]

- name: LeverCorp
  careers_url: https://jobs.lever.co/levercorp
  provider: lever
  title_filter:
    positive: ["product manager"]
  content_filter:
    positive: ["agile", "scrum"]

```

## Summary

- **Provider architecture**: `scan.mjs` loads modules from `providers/*.mjs` via `providers/_registry.mjs`, with each provider implementing `detect()` and `fetch()` functions.
- **Central configuration**: Define companies in [`portals.yml`](https://github.com/santifer/career-ops/blob/main/portals.yml) specifying `provider` (greenhouse, ashby, or lever) and optional explicit `api` endpoints.
- **Filtering system**: Use `title_filter` (with boundary-aware matching), `location_filter` (with regex look-arounds), and `content_filter` (with per-title scoping) to control which postings are retained.
- **Execution**: Run `node scan.mjs` with optional `--company` or `--dry-run` flags to control the scan scope.

## Frequently Asked Questions

### How does Career-Ops auto-detect the correct API URL for a job board?

The `detect(entry)` function in each provider analyzes the `careers_url` pattern. For Greenhouse, `resolveApiUrl` extracts the board identifier from URLs like `job-boards.greenhouse.io/exampleco` and constructs the API path. If detection fails or you need to override the endpoint, provide the `api` field directly in [`portals.yml`](https://github.com/santifer/career-ops/blob/main/portals.yml).

### Can I use multiple filters simultaneously for the same company?

Yes. The scanner processes filters sequentially: it first checks `title_filter`, then `location_filter`, and finally `content_filter`. A posting must pass all active filters to be stored. You can combine positive and negative keywords across all three filter types to create precise targeting rules.

### Why does the Greenhouse provider make additional HTTP requests for some jobs?

Greenhouse sometimes returns location strings that contain only work models like "Hybrid" or "Remote". The provider detects this via `isWorkModelOnly()` and calls the `/offices` endpoint (constructed by `officesUrlFor()`) to fetch the actual office locations, then maps them using `buildOfficeMap()` (see `greenhouse.mjs` lines 48-77).

### How do I test a new portal configuration without affecting my database?

Use the `--dry-run` flag when executing the scanner: `node scan.mjs --company NewCompany --dry-run`. This executes the detection, fetching, and filtering logic but prints results to stdout instead of persisting them, allowing you to verify your [`portals.yml`](https://github.com/santifer/career-ops/blob/main/portals.yml) configuration before production use.