How `scan.mjs` Performs Zero-Token Portal Scanning via Greenhouse, Ashby, and Lever APIs in Career-Ops
scan.mjs is a zero-token portal scanner that queries Greenhouse, Ashby, and Lever public APIs directly via HTTP requests without consuming any LLM tokens.
The Career-Ops project by santifer implements a lightweight, autonomous job discovery system that avoids expensive AI API calls. Instead of using language models to parse job descriptions, scan.mjs leverages standardized ATS (Applicant Tracking System) endpoints to fetch structured job data efficiently. This article explains exactly how the scanner works, from provider detection to deduplication.
Provider Plugin Architecture
The scanner loads modular providers through a registry pattern. At startup, scan.mjs imports the provider loader:
import { loadProviders, resolveProvider } from './providers/_registry.mjs';
This import appears at lines 45-48 of scan.mjs. The registry automatically discovers all *.mjs files in the providers/ directory (lines 94-96 of providers/_registry.mjs), enabling clean separation of concerns across ATS vendors.
Each provider exports a default object with three required properties:
id– A stable identifier stringdetect(entry)– Determines if this provider handles a given portal entryfetch(entry, ctx)– Executes the HTTP request and returns normalized job data
API URL Detection and Resolution
Greenhouse Detection
The Greenhouse provider in providers/greenhouse.mjs extracts API endpoints from careers page URLs using resolveApiUrl (lines 30-37):
// From careers_url like "https://job-boards.greenhouse.io/companyname"
// Extracts board name and constructs: https://job-boards.greenhouse.io/companyname/jobs
The function validates the hostname with assertGreenhouseUrl before parsing.
Lever Detection
The Lever provider in providers/lever.mjs performs equivalent URL transformation (lines 24-44). It converts public careers URLs:
https://jobs.lever.co/companyname → https://api.lever.co/v0/postings/companyname
The assertLeverUrl sandbox function guarantees only whitelisted *.lever.co domains are contacted.
Ashby Detection
providers/ashby.mjs follows the same pattern, detecting Ashby-hosted careers pages and mapping them to the corresponding Ashby GraphQL or REST endpoints.
Single-Request Job Fetching
All providers implement zero-token data retrieval through ctx.fetchJson(). No per-job requests are executed—each API call returns a complete job list in one payload.
Greenhouse fetch implementation (lines 36-38 of providers/greenhouse.mjs):
const json = await ctx.fetchJson(apiUrl, { redirect: 'error' });
return json.jobs.map(job => normalizeJob(job, ctx));
Lever fetch implementation (lines 82-84 of providers/lever.mjs):
const postings = await ctx.fetchJson(apiUrl, { redirect: 'error' });
return postings.map(post => normalizePost(post, ctx));
The redirect: 'error' option prevents unexpected URL hops that could leak to unintended domains.
Data Normalization and Enrichment
Providers transform raw API responses into a common schema with these fields: title, url, company, location, description (optional), postedAt (optional).
Greenhouse location enrichment: When a job lists only a work model (e.g., "Remote") without geography, the provider optionally fetches /offices to map office IDs to locations (lines 40-64 of providers/greenhouse.mjs).
Lever location merging: The Lever provider flattens location and allLocations arrays into a single searchable string (lines 46-63 of providers/lever.mjs):
const location = [
post.location,
...(post.allLocations || [])
].filter(Boolean).join('; ');
Filter Pipeline in scan.mjs
After normalization, jobs flow through a multi-stage filter pipeline defined entirely in scan.mjs. All filters execute locally without external API calls.
| Filter | Line Range | Purpose |
|---|---|---|
buildTitleFilter |
98-123 | Match positive keywords, exclude negative terms |
buildLocationFilter |
201-226 | Geographic constraints (country, city, remote status) |
buildContentFilter |
443-480 | Full-text search in descriptions |
| Salary filter | 82-93 | Numeric range matching on compensation fields |
| Visa filter | 120-123 | Sponsorship availability checks |
| Country eligibility | 151-165 | Work authorization requirements |
| Cooldown filter | — | Skip recently seen companies |
| Posting age filter | — | Maximum days since publication |
Each filter returns a predicate function applied sequentially. A job must pass all active filters to be output.
Deduplication Against Scan History
The scanner maintains persistent state to avoid re-processing identical postings:
loadSeenUrls()– Loads URL set fromdata/scan-history.tsvloadSeenCompanyRoles()– Loads composite keys of "company::normalized-title"
URL normalization (lines 1110-1118 of scan.mjs) preserves identity parameters like Greenhouse's gh_jid while stripping tracking parameters:
function normalizeUrlForDedup(url) {
const u = new URL(url);
// Keep gh_jid, remove utm_*, fbclid, etc.
return `${u.origin}${u.pathname}${preservedParams}`;
}
The main loop skips any offer whose dedup key exists (lines 4250-4258), ensuring idempotent scans.
Writing Results and Persistence
New offers append to data/pipeline.md via appendToPipeline (lines 1931-1945):
## ExampleCorp — Senior Platform Engineer
- **Location:** Berlin, Germany (Hybrid)
- **URL:** https://job-boards.greenhouse.io/example/jobs/12345
- **Posted:** 2024-01-15
- **Description:** [truncated]
Scan history appends to data/scan-history.tsv through appendToScanHistory (lines 1977-1990), creating an audit trail for future deduplication.
CLI Usage and Options
Invoke the scanner directly with Node.js:
# Scan all companies defined in portals.yml
node scan.mjs
# Scan single company (useful for debugging provider logic)
node scan.mjs --company Cohere
# Only consider postings from last N days
node scan.mjs --since 7
# Validate without writing files
node scan.mjs --dry-run
CLI argument parsing appears at lines 23-37 of scan.mjs, supporting both interactive and automation workflows.
Extending with Custom Providers
To add a new ATS (e.g., Workday, SmartRecruiters):
-
Create
providers/newats.mjsexporting:id: Provider identifierdetect(entry): Return API URL ornullfetch(entry, ctx): Return array of normalized job objects
-
The registry auto-loads your provider on next
scan.mjsexecution
Required output schema per job object:
{
title: string, // Job title
url: string, // Direct application URL
company: string, // Normalized company name
location: string, // Human-readable location
description?: string, // Optional full text
postedAt?: string // ISO 8601 date
}
Security and Sandboxing
Provider modules enforce hostname validation through assertion functions (assertGreenhouseUrl, assertLeverUrl, etc.). This prevents malicious portals.yml entries from exfiltrating data to untrusted endpoints. The redirect: 'error' fetch option adds defense-in-depth against open redirects.
Summary
scan.mjsperforms zero-token job scanning by querying Greenhouse, Ashby, and Lever public APIs directly- Provider plugins under
providers/*.mjshandle ATS-specific URL detection, fetching, and normalization - Single HTTP requests retrieve complete job listings—no per-job pagination or LLM processing
- Local filter pipeline in
scan.mjsapplies title, location, content, salary, and eligibility criteria without external dependencies - Deduplication system tracks seen URLs and company-role pairs across runs via
data/scan-history.tsv - Sandboxed HTTP with whitelist validation guarantees safe, reproducible execution
Frequently Asked Questions
Does scan.mjs require API keys or authentication?
No. The Greenhouse, Ashby, and Lever APIs used by scan.mjs are public endpoints that do not require authentication. The scanner sends unauthenticated HTTP GET requests to standardized job board URLs, and these ATS vendors intentionally expose this data for indexing and aggregation purposes.
How does the scanner avoid duplicate job postings across multiple runs?
scan.mjs implements two deduplication mechanisms. First, it normalizes and stores every scraped URL in data/scan-history.tsv. Second, it tracks composite "company::role" keys to catch reposted listings with different URLs. The normalizeUrlForDedup function carefully preserves essential identity parameters like gh_jid while removing cosmetic query strings that would otherwise fragment deduplication.
Can I filter jobs by salary or visa sponsorship without using AI?
Yes. The filter pipeline in scan.mjs (lines 82-93 for salary, 120-123 for visa) operates on structured data fields returned by the ATS APIs. When an API includes salary ranges or visa sponsorship flags, these filters match against them directly. No natural language processing or LLM calls are required for these constraints.
What happens if a careers page URL redirects to a different domain?
The scanner explicitly sets redirect: 'error' in all ctx.fetchJson() calls. This prevents silent redirection to unexpected domains, which could indicate a compromised link or tracking system. If a redirect occurs, the fetch throws an error and the provider skips that entry, maintaining the security boundary established by the hostname whitelists.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →