How `scan.mjs` Performs Zero-Token Portal Scanning via Greenhouse, Ashby, and Lever APIs in Career-Ops

scan.mjs is a zero-token portal scanner that queries Greenhouse, Ashby, and Lever public APIs directly via HTTP requests without consuming any LLM tokens.

The Career-Ops project by santifer implements a lightweight, autonomous job discovery system that avoids expensive AI API calls. Instead of using language models to parse job descriptions, scan.mjs leverages standardized ATS (Applicant Tracking System) endpoints to fetch structured job data efficiently. This article explains exactly how the scanner works, from provider detection to deduplication.

Provider Plugin Architecture

The scanner loads modular providers through a registry pattern. At startup, scan.mjs imports the provider loader:

import { loadProviders, resolveProvider } from './providers/_registry.mjs';

This import appears at lines 45-48 of scan.mjs. The registry automatically discovers all *.mjs files in the providers/ directory (lines 94-96 of providers/_registry.mjs), enabling clean separation of concerns across ATS vendors.

Each provider exports a default object with three required properties:

  • id – A stable identifier string
  • detect(entry) – Determines if this provider handles a given portal entry
  • fetch(entry, ctx) – Executes the HTTP request and returns normalized job data

API URL Detection and Resolution

Greenhouse Detection

The Greenhouse provider in providers/greenhouse.mjs extracts API endpoints from careers page URLs using resolveApiUrl (lines 30-37):

// From careers_url like "https://job-boards.greenhouse.io/companyname"
// Extracts board name and constructs: https://job-boards.greenhouse.io/companyname/jobs

The function validates the hostname with assertGreenhouseUrl before parsing.

Lever Detection

The Lever provider in providers/lever.mjs performs equivalent URL transformation (lines 24-44). It converts public careers URLs:


https://jobs.lever.co/companyname  →  https://api.lever.co/v0/postings/companyname

The assertLeverUrl sandbox function guarantees only whitelisted *.lever.co domains are contacted.

Ashby Detection

providers/ashby.mjs follows the same pattern, detecting Ashby-hosted careers pages and mapping them to the corresponding Ashby GraphQL or REST endpoints.

Single-Request Job Fetching

All providers implement zero-token data retrieval through ctx.fetchJson(). No per-job requests are executed—each API call returns a complete job list in one payload.

Greenhouse fetch implementation (lines 36-38 of providers/greenhouse.mjs):

const json = await ctx.fetchJson(apiUrl, { redirect: 'error' });
return json.jobs.map(job => normalizeJob(job, ctx));

Lever fetch implementation (lines 82-84 of providers/lever.mjs):

const postings = await ctx.fetchJson(apiUrl, { redirect: 'error' });
return postings.map(post => normalizePost(post, ctx));

The redirect: 'error' option prevents unexpected URL hops that could leak to unintended domains.

Data Normalization and Enrichment

Providers transform raw API responses into a common schema with these fields: title, url, company, location, description (optional), postedAt (optional).

Greenhouse location enrichment: When a job lists only a work model (e.g., "Remote") without geography, the provider optionally fetches /offices to map office IDs to locations (lines 40-64 of providers/greenhouse.mjs).

Lever location merging: The Lever provider flattens location and allLocations arrays into a single searchable string (lines 46-63 of providers/lever.mjs):

const location = [
  post.location,
  ...(post.allLocations || [])
].filter(Boolean).join('; ');

Filter Pipeline in scan.mjs

After normalization, jobs flow through a multi-stage filter pipeline defined entirely in scan.mjs. All filters execute locally without external API calls.

Filter Line Range Purpose
buildTitleFilter 98-123 Match positive keywords, exclude negative terms
buildLocationFilter 201-226 Geographic constraints (country, city, remote status)
buildContentFilter 443-480 Full-text search in descriptions
Salary filter 82-93 Numeric range matching on compensation fields
Visa filter 120-123 Sponsorship availability checks
Country eligibility 151-165 Work authorization requirements
Cooldown filter — Skip recently seen companies
Posting age filter — Maximum days since publication

Each filter returns a predicate function applied sequentially. A job must pass all active filters to be output.

Deduplication Against Scan History

The scanner maintains persistent state to avoid re-processing identical postings:

  • loadSeenUrls() – Loads URL set from data/scan-history.tsv
  • loadSeenCompanyRoles() – Loads composite keys of "company::normalized-title"

URL normalization (lines 1110-1118 of scan.mjs) preserves identity parameters like Greenhouse's gh_jid while stripping tracking parameters:

function normalizeUrlForDedup(url) {
  const u = new URL(url);
  // Keep gh_jid, remove utm_*, fbclid, etc.
  return `${u.origin}${u.pathname}${preservedParams}`;
}

The main loop skips any offer whose dedup key exists (lines 4250-4258), ensuring idempotent scans.

Writing Results and Persistence

New offers append to data/pipeline.md via appendToPipeline (lines 1931-1945):


## ExampleCorp — Senior Platform Engineer

- **Location:** Berlin, Germany (Hybrid)
- **URL:** https://job-boards.greenhouse.io/example/jobs/12345
- **Posted:** 2024-01-15
- **Description:** [truncated]

Scan history appends to data/scan-history.tsv through appendToScanHistory (lines 1977-1990), creating an audit trail for future deduplication.

CLI Usage and Options

Invoke the scanner directly with Node.js:


# Scan all companies defined in portals.yml

node scan.mjs

# Scan single company (useful for debugging provider logic)

node scan.mjs --company Cohere

# Only consider postings from last N days

node scan.mjs --since 7

# Validate without writing files

node scan.mjs --dry-run

CLI argument parsing appears at lines 23-37 of scan.mjs, supporting both interactive and automation workflows.

Extending with Custom Providers

To add a new ATS (e.g., Workday, SmartRecruiters):

  1. Create providers/newats.mjs exporting:

    • id: Provider identifier
    • detect(entry): Return API URL or null
    • fetch(entry, ctx): Return array of normalized job objects
  2. The registry auto-loads your provider on next scan.mjs execution

Required output schema per job object:

{
  title: string,      // Job title
  url: string,        // Direct application URL
  company: string,    // Normalized company name
  location: string,   // Human-readable location
  description?: string,  // Optional full text
  postedAt?: string     // ISO 8601 date
}

Security and Sandboxing

Provider modules enforce hostname validation through assertion functions (assertGreenhouseUrl, assertLeverUrl, etc.). This prevents malicious portals.yml entries from exfiltrating data to untrusted endpoints. The redirect: 'error' fetch option adds defense-in-depth against open redirects.

Summary

  • scan.mjs performs zero-token job scanning by querying Greenhouse, Ashby, and Lever public APIs directly
  • Provider plugins under providers/*.mjs handle ATS-specific URL detection, fetching, and normalization
  • Single HTTP requests retrieve complete job listings—no per-job pagination or LLM processing
  • Local filter pipeline in scan.mjs applies title, location, content, salary, and eligibility criteria without external dependencies
  • Deduplication system tracks seen URLs and company-role pairs across runs via data/scan-history.tsv
  • Sandboxed HTTP with whitelist validation guarantees safe, reproducible execution

Frequently Asked Questions

Does scan.mjs require API keys or authentication?

No. The Greenhouse, Ashby, and Lever APIs used by scan.mjs are public endpoints that do not require authentication. The scanner sends unauthenticated HTTP GET requests to standardized job board URLs, and these ATS vendors intentionally expose this data for indexing and aggregation purposes.

How does the scanner avoid duplicate job postings across multiple runs?

scan.mjs implements two deduplication mechanisms. First, it normalizes and stores every scraped URL in data/scan-history.tsv. Second, it tracks composite "company::role" keys to catch reposted listings with different URLs. The normalizeUrlForDedup function carefully preserves essential identity parameters like gh_jid while removing cosmetic query strings that would otherwise fragment deduplication.

Can I filter jobs by salary or visa sponsorship without using AI?

Yes. The filter pipeline in scan.mjs (lines 82-93 for salary, 120-123 for visa) operates on structured data fields returned by the ATS APIs. When an API includes salary ranges or visa sponsorship flags, these filters match against them directly. No natural language processing or LLM calls are required for these constraints.

What happens if a careers page URL redirects to a different domain?

The scanner explicitly sets redirect: 'error' in all ctx.fetchJson() calls. This prevents silent redirection to unexpected domains, which could indicate a compromised link or tracking system. If a redirect occurs, the fetch throws an error and the provider skips that entry, maintaining the security boundary established by the hostname whitelists.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →