# How `scan.mjs` Performs Zero-Token Portal Scanning via Greenhouse, Ashby, and Lever APIs in Career-Ops

> Learn how scan.mjs performs zero-token portal scanning using Greenhouse Ashby and Lever APIs directly via HTTP requests in the career ops repository. Discover efficient recruitment.

- Repository: [Santiago Fernández de Valderrama/career-ops](https://github.com/santifer/career-ops)
- Tags: how-to-guide
- Published: 2026-08-20

---

**`scan.mjs` is a zero-token portal scanner that queries Greenhouse, Ashby, and Lever public APIs directly via HTTP requests without consuming any LLM tokens.**

The Career-Ops project by santifer implements a lightweight, autonomous job discovery system that avoids expensive AI API calls. Instead of using language models to parse job descriptions, `scan.mjs` leverages standardized ATS (Applicant Tracking System) endpoints to fetch structured job data efficiently. This article explains exactly how the scanner works, from provider detection to deduplication.

## Provider Plugin Architecture

The scanner loads modular providers through a registry pattern. At startup, `scan.mjs` imports the provider loader:

```js
import { loadProviders, resolveProvider } from './providers/_registry.mjs';

```

This import appears at lines 45-48 of `scan.mjs`. The registry automatically discovers all `*.mjs` files in the `providers/` directory (lines 94-96 of `providers/_registry.mjs`), enabling clean separation of concerns across ATS vendors.

Each provider exports a default object with three required properties:

- **`id`** – A stable identifier string
- **`detect(entry)`** – Determines if this provider handles a given portal entry
- **`fetch(entry, ctx)`** – Executes the HTTP request and returns normalized job data

## API URL Detection and Resolution

### Greenhouse Detection

The Greenhouse provider in `providers/greenhouse.mjs` extracts API endpoints from careers page URLs using `resolveApiUrl` (lines 30-37):

```js
// From careers_url like "https://job-boards.greenhouse.io/companyname"
// Extracts board name and constructs: https://job-boards.greenhouse.io/companyname/jobs

```

The function validates the hostname with `assertGreenhouseUrl` before parsing.

### Lever Detection

The Lever provider in `providers/lever.mjs` performs equivalent URL transformation (lines 24-44). It converts public careers URLs:

```

https://jobs.lever.co/companyname  →  https://api.lever.co/v0/postings/companyname

```

The `assertLeverUrl` sandbox function guarantees only whitelisted `*.lever.co` domains are contacted.

### Ashby Detection

`providers/ashby.mjs` follows the same pattern, detecting Ashby-hosted careers pages and mapping them to the corresponding Ashby GraphQL or REST endpoints.

## Single-Request Job Fetching

All providers implement zero-token data retrieval through `ctx.fetchJson()`. No per-job requests are executed—each API call returns a complete job list in one payload.

**Greenhouse fetch implementation** (lines 36-38 of `providers/greenhouse.mjs`):

```js
const json = await ctx.fetchJson(apiUrl, { redirect: 'error' });
return json.jobs.map(job => normalizeJob(job, ctx));

```

**Lever fetch implementation** (lines 82-84 of `providers/lever.mjs`):

```js
const postings = await ctx.fetchJson(apiUrl, { redirect: 'error' });
return postings.map(post => normalizePost(post, ctx));

```

The `redirect: 'error'` option prevents unexpected URL hops that could leak to unintended domains.

## Data Normalization and Enrichment

Providers transform raw API responses into a common schema with these fields: `title`, `url`, `company`, `location`, `description` (optional), `postedAt` (optional).

**Greenhouse location enrichment**: When a job lists only a work model (e.g., "Remote") without geography, the provider optionally fetches `/offices` to map office IDs to locations (lines 40-64 of `providers/greenhouse.mjs`).

**Lever location merging**: The Lever provider flattens `location` and `allLocations` arrays into a single searchable string (lines 46-63 of `providers/lever.mjs`):

```js
const location = [
  post.location,
  ...(post.allLocations || [])
].filter(Boolean).join('; ');

```

## Filter Pipeline in scan.mjs

After normalization, jobs flow through a multi-stage filter pipeline defined entirely in `scan.mjs`. All filters execute locally without external API calls.

| Filter | Line Range | Purpose |
|--------|-----------|---------|
| `buildTitleFilter` | 98-123 | Match positive keywords, exclude negative terms |
| `buildLocationFilter` | 201-226 | Geographic constraints (country, city, remote status) |
| `buildContentFilter` | 443-480 | Full-text search in descriptions |
| Salary filter | 82-93 | Numeric range matching on compensation fields |
| Visa filter | 120-123 | Sponsorship availability checks |
| Country eligibility | 151-165 | Work authorization requirements |
| Cooldown filter | — | Skip recently seen companies |
| Posting age filter | — | Maximum days since publication |

Each filter returns a predicate function applied sequentially. A job must pass all active filters to be output.

## Deduplication Against Scan History

The scanner maintains persistent state to avoid re-processing identical postings:

- **`loadSeenUrls()`** – Loads URL set from `data/scan-history.tsv`
- **`loadSeenCompanyRoles()`** – Loads composite keys of "company::normalized-title"

**URL normalization** (lines 1110-1118 of `scan.mjs`) preserves identity parameters like Greenhouse's `gh_jid` while stripping tracking parameters:

```js
function normalizeUrlForDedup(url) {
  const u = new URL(url);
  // Keep gh_jid, remove utm_*, fbclid, etc.
  return `${u.origin}${u.pathname}${preservedParams}`;
}

```

The main loop skips any offer whose dedup key exists (lines 4250-4258), ensuring idempotent scans.

## Writing Results and Persistence

New offers append to [`data/pipeline.md`](https://github.com/santifer/career-ops/blob/main/data/pipeline.md) via `appendToPipeline` (lines 1931-1945):

```markdown

## ExampleCorp — Senior Platform Engineer

- **Location:** Berlin, Germany (Hybrid)
- **URL:** https://job-boards.greenhouse.io/example/jobs/12345
- **Posted:** 2024-01-15
- **Description:** [truncated]

```

Scan history appends to `data/scan-history.tsv` through `appendToScanHistory` (lines 1977-1990), creating an audit trail for future deduplication.

## CLI Usage and Options

Invoke the scanner directly with Node.js:

```bash

# Scan all companies defined in portals.yml

node scan.mjs

# Scan single company (useful for debugging provider logic)

node scan.mjs --company Cohere

# Only consider postings from last N days

node scan.mjs --since 7

# Validate without writing files

node scan.mjs --dry-run

```

CLI argument parsing appears at lines 23-37 of `scan.mjs`, supporting both interactive and automation workflows.

## Extending with Custom Providers

To add a new ATS (e.g., Workday, SmartRecruiters):

1. Create `providers/newats.mjs` exporting:
   - `id`: Provider identifier
   - `detect(entry)`: Return API URL or `null`
   - `fetch(entry, ctx)`: Return array of normalized job objects

2. The registry auto-loads your provider on next `scan.mjs` execution

Required output schema per job object:

```js
{
  title: string,      // Job title
  url: string,        // Direct application URL
  company: string,    // Normalized company name
  location: string,   // Human-readable location
  description?: string,  // Optional full text
  postedAt?: string     // ISO 8601 date
}

```

## Security and Sandboxing

Provider modules enforce hostname validation through assertion functions (`assertGreenhouseUrl`, `assertLeverUrl`, etc.). This prevents malicious [`portals.yml`](https://github.com/santifer/career-ops/blob/main/portals.yml) entries from exfiltrating data to untrusted endpoints. The `redirect: 'error'` fetch option adds defense-in-depth against open redirects.

## Summary

- **`scan.mjs`** performs zero-token job scanning by querying Greenhouse, Ashby, and Lever public APIs directly
- **Provider plugins** under `providers/*.mjs` handle ATS-specific URL detection, fetching, and normalization
- **Single HTTP requests** retrieve complete job listings—no per-job pagination or LLM processing
- **Local filter pipeline** in `scan.mjs` applies title, location, content, salary, and eligibility criteria without external dependencies
- **Deduplication system** tracks seen URLs and company-role pairs across runs via `data/scan-history.tsv`
- **Sandboxed HTTP** with whitelist validation guarantees safe, reproducible execution

## Frequently Asked Questions

### Does scan.mjs require API keys or authentication?

No. The Greenhouse, Ashby, and Lever APIs used by `scan.mjs` are public endpoints that do not require authentication. The scanner sends unauthenticated HTTP GET requests to standardized job board URLs, and these ATS vendors intentionally expose this data for indexing and aggregation purposes.

### How does the scanner avoid duplicate job postings across multiple runs?

`scan.mjs` implements two deduplication mechanisms. First, it normalizes and stores every scraped URL in `data/scan-history.tsv`. Second, it tracks composite "company::role" keys to catch reposted listings with different URLs. The `normalizeUrlForDedup` function carefully preserves essential identity parameters like `gh_jid` while removing cosmetic query strings that would otherwise fragment deduplication.

### Can I filter jobs by salary or visa sponsorship without using AI?

Yes. The filter pipeline in `scan.mjs` (lines 82-93 for salary, 120-123 for visa) operates on structured data fields returned by the ATS APIs. When an API includes salary ranges or visa sponsorship flags, these filters match against them directly. No natural language processing or LLM calls are required for these constraints.

### What happens if a careers page URL redirects to a different domain?

The scanner explicitly sets `redirect: 'error'` in all `ctx.fetchJson()` calls. This prevents silent redirection to unexpected domains, which could indicate a compromised link or tracking system. If a redirect occurs, the fetch throws an error and the provider skips that entry, maintaining the security boundary established by the hostname whitelists.