How to Set Up and Customize Portal Scanning for Greenhouse, Ashby, and Lever APIs
Career-Ops uses a plugin-based scanner that loads provider modules from providers/*.mjs to auto-detect and fetch jobs from Greenhouse, Ashby, and Lever APIs, with all customization controlled through a central portals.yml configuration file.
The career-ops repository provides an open-source job aggregation system that discovers postings through a modular provider architecture. You can set up and customize portal scanning for Greenhouse, Ashby, and Lever APIs by editing the declarative portals.yml file and leveraging the built-in filtering system implemented in scan.mjs.
Understanding the Provider Architecture
The core scanning logic resides in scan.mjs, which dynamically loads provider implementations from the providers/ directory via providers/_registry.mjs. Each provider is a JavaScript module that exports an id and implements two required functions:
detect(entry): Analyzes thecareers_urlto determine the appropriate API endpointfetch(entry, ctx): Executes the HTTP request to the provider's API and returns normalized job objects with{title, url, company, location, postedAt}
This architecture allows the scanner to support multiple applicant tracking systems through a standardized interface while handling each provider's unique URL patterns and response formats.
Configuring Portal Entries in portals.yml
All company configurations live in portals.yml at the repository root. Each entry specifies which provider handles the job board and how to access it.
Basic Structure and Provider Selection
- name: ExampleCo
careers_url: https://exampleco.com/careers
provider: greenhouse # Options: greenhouse, ashby, lever
The provider field must match the id exported by the corresponding file in providers/greenhouse.mjs, providers/ashby.mjs, or providers/lever.mjs.
Explicit API Endpoints
To bypass auto-detection, specify the API URL directly:
- name: ExampleCo
careers_url: https://exampleco.com/careers
provider: greenhouse
api: https://boards-api.greenhouse.io/v1/boards/exampleco/jobs
When the api field is present, the scanner skips the detect() function and uses the provided URL.
Customizing Job Filters
The scanner provides three filtering layers defined in scan.mjs that process postings before storage.
Title Filtering
The title_filter field uses the compileKeyword and compilePositiveKeyword helpers (defined in scan.mjs lines 4-11) to match role titles. Positive keywords support Boolean logic (e.g., "engineer + senior"), while negative keywords veto postings regardless of other matches.
title_filter:
positive: ["engineer + senior", "data scientist"]
negative: ["intern", "contract"]
Short keywords (2-3 letters) automatically match on word boundaries to prevent false positives.
Location Filtering
Location filters in scan.mjs (lines 38-62) process the location_filter block, which supports three rules:
always_allow: Permits matching locations regardless of other rulesblock: Excludes locations containing these termsallow: Requires locations to contain at least one of these terms
The system uses look-around regex patterns so "india" does not match "Indianapolis".
location_filter:
always_allow: ["remote"]
block: ["china", "russia"]
allow: ["united states", "canada"]
Content Filtering
For providers that return job descriptions, content_filter (implemented in buildContentFilter, lines 98-126 of scan.mjs) enables keyword matching within the posting body. You can also scope rules to specific title keywords using by_title_keyword.
content_filter:
positive: ["python", "aws"]
negative: ["php"]
by_title_keyword:
"data scientist":
positive: ["tensorflow", "pytorch"]
negative: ["excel"]
Provider-Specific Setup Details
Each ATS provider handles API discovery and data normalization differently.
Greenhouse Configuration
The Greenhouse provider (providers/greenhouse.mjs) detects API URLs from board pages using resolveApiUrl (lines 28-37). It also handles location enrichment: when a posting lists only a work model like "Hybrid", the scanner fetches the /offices endpoint (via isWorkModelOnly, officesUrlFor, and buildOfficeMap, lines 48-77) to retrieve actual office locations.
Ashby Configuration
The Ashby provider (providers/ashby.mjs) uses the public /jobs endpoint without requiring authentication. It follows the standard detect() → fetch() pattern without additional enrichment steps.
Lever Configuration
The Lever provider (providers/lever.mjs) supports both explicit api URLs and automatic detection from Lever careers pages. It extracts the JSON job list and normalizes fields to match the standard schema.
Running the Scanner
Execute the scanner from the repository root using Node.js:
# Scan all companies defined in portals.yml
node scan.mjs
# Test a single company
node scan.mjs --company ExampleCo
# Preview changes without writing to database
node scan.mjs --dry-run
CLI flag parsing is implemented at the top of scan.mjs.
Example Configuration for All Three Providers
- name: GreenTech
careers_url: https://greentech.com/careers
provider: greenhouse
title_filter:
positive: ["engineer + senior"]
location_filter:
allow: ["united states", "canada"]
content_filter:
positive: ["nodejs", "kubernetes"]
negative: ["php"]
- name: AshbyAI
careers_url: https://jobs.ashby.ai/ashbyai
provider: ashby
title_filter:
positive: ["data scientist"]
negative: ["intern"]
location_filter:
block: ["remote"]
- name: LeverCorp
careers_url: https://jobs.lever.co/levercorp
provider: lever
title_filter:
positive: ["product manager"]
content_filter:
positive: ["agile", "scrum"]
Summary
- Provider architecture:
scan.mjsloads modules fromproviders/*.mjsviaproviders/_registry.mjs, with each provider implementingdetect()andfetch()functions. - Central configuration: Define companies in
portals.ymlspecifyingprovider(greenhouse, ashby, or lever) and optional explicitapiendpoints. - Filtering system: Use
title_filter(with boundary-aware matching),location_filter(with regex look-arounds), andcontent_filter(with per-title scoping) to control which postings are retained. - Execution: Run
node scan.mjswith optional--companyor--dry-runflags to control the scan scope.
Frequently Asked Questions
How does Career-Ops auto-detect the correct API URL for a job board?
The detect(entry) function in each provider analyzes the careers_url pattern. For Greenhouse, resolveApiUrl extracts the board identifier from URLs like job-boards.greenhouse.io/exampleco and constructs the API path. If detection fails or you need to override the endpoint, provide the api field directly in portals.yml.
Can I use multiple filters simultaneously for the same company?
Yes. The scanner processes filters sequentially: it first checks title_filter, then location_filter, and finally content_filter. A posting must pass all active filters to be stored. You can combine positive and negative keywords across all three filter types to create precise targeting rules.
Why does the Greenhouse provider make additional HTTP requests for some jobs?
Greenhouse sometimes returns location strings that contain only work models like "Hybrid" or "Remote". The provider detects this via isWorkModelOnly() and calls the /offices endpoint (constructed by officesUrlFor()) to fetch the actual office locations, then maps them using buildOfficeMap() (see greenhouse.mjs lines 48-77).
How do I test a new portal configuration without affecting my database?
Use the --dry-run flag when executing the scanner: node scan.mjs --company NewCompany --dry-run. This executes the detection, fetching, and filtering logic but prints results to stdout instead of persisting them, allowing you to verify your portals.yml configuration before production use.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →