# How `scan-ats-full.mjs` Performs Checkpoint‑Based Reverse‑ATS Scanning Across Full Datasets in Career‑Ops

> Discover how scan-ats-full.mjs uses checkpoint-based reverse-ATS scanning to efficiently process full datasets in career-ops. Resumable pipeline survives interruptions without re-processing companies.

- Repository: [Santiago Fernández de Valderrama/career-ops](https://github.com/santifer/career-ops)
- Tags: internals
- Published: 2026-08-20

---

**`scan-ats-full.mjs` implements a resumable, checkpoint-driven pipeline that scans public job-board-aggregators for major ATS providers, filters matches by date and criteria, and survives interruptions without re-processing companies.**

This deep dive examines the architecture of **checkpoint-based reverse-ATS scanning** as implemented in the `santifer/career-ops` repository. The script orchestrates multi-hour sweeps across Greenhouse, Lever, Ashby, Workday, and iCIMS datasets while ensuring crash safety and data integrity through atomic checkpointing and dataset fingerprinting.

## Core Architecture: The Checkpoint-Driven Pipeline

The scanning engine follows a deterministic, resumable workflow designed for reliability at scale. Every significant state change is persisted, allowing users to interrupt and resume long-running jobs without loss of progress.

### CLI Configuration and Flag Validation

Entry-point flags are parsed and validated via `validateFlags` in `lib/cli-flags.mjs` ([lines 31-62](https://github.com/santifer/career-ops/blob/main/scan-ats-full.mjs#L31-L62)). Key options include:

- `--since <days>` — Date window for fresh postings (default: 3 days)
- `--limit <n>` — Cap companies per ATS source
- `--ats <list>` — Target specific providers (greenhouse, lever, ashby, workday, icims)
- `--resume` — Continue from last checkpoint
- `--shuffle` — Randomize company order (disables resume)
- `--include-undated` — Retain postings without parseable dates

```bash

# Resume an interrupted 10,000-company sweep

node scan-ats-full.mjs --resume

# Target only Greenhouse with custom date window

node scan-ats-full.mjs --ats greenhouse --since 7 --limit 500

```

### Dataset Loading and Cache Management

The script fetches cached JSON company lists from the public `job-board-aggregator` repository via `loadCompanyList` ([lines 31-40](https://github.com/santifer/career-ops/blob/main/scan-ats-full.mjs#L31-L40)). Cache entries expire after `CACHE_TTL_HOURS` (24 hours) to balance freshness with API rate limits.

Each ATS provider maps to a configuration in `SOURCES[name]` defining:
- `provider` — Fetch implementation module
- `concurrency` — Worker pool size for parallel requests

## Checkpoint Mechanics: Crash-Safe State Persistence

The **checkpoint system** is the foundation of resumable operation. Checkpoints are written to [`data/cache/ats-full-checkpoint.json`](https://github.com/santifer/career-ops/blob/main/data/cache/ats-full-checkpoint.json) and validated on resume to prevent silent data corruption.

### Checkpoint Compatibility Verification

When `--resume` is specified, `loadCheckpoint()` ([lines 96-118](https://github.com/santifer/career-ops/blob/main/scan-ats-full.mjs#L96-L118)) performs strict compatibility checks:

```javascript
// Pseudo-structure from source analysis
checkpointCompatible(current, checkpoint) {
  return (
    arraysEqual(current.ats, checkpoint.ats) &&
    current.limit === checkpoint.limit &&
    current.includeUndated === checkpoint.includeUndated &&
    hashDataset(current.dataset) === checkpoint.datasetFingerprint
  )
}

```

The **SHA-1 fingerprint** (`datasetFingerprint`) detects any drift in the underlying company list. If the dataset hash mismatches, the resume aborts—forcing a fresh scan rather than risking incomplete or inconsistent results.

### Atomic Checkpoint Writes

Checkpoints are written every `CHECKPOINT_EVERY` (500) companies via `writeCheckpoint()` ([lines 31-45](https://github.com/santifer/career-ops/blob/main/scan-ats-full.mjs#L31-L45)):

```javascript
// Atomic write pattern from source
const tmpPath = `${CHECKPOINT_PATH}.tmp`;
fs.writeFileSync(tmpPath, JSON.stringify(checkpoint));
fs.renameSync(tmpPath, CHECKPOINT_PATH);  // Atomic on POSIX

```

This **write-to-temporary-then-rename** pattern guarantees that a crash during checkpoint serialization cannot leave a corrupt partial file.

### Checkpoint Schema

```json
{
  "version": 1,
  "cutoffMs": 1724161921123,
  "ats": ["greenhouse", "lever", "ashby"],
  "limit": 200,
  "includeUndated": false,
  "completedSources": ["greenhouse"],
  "resumeAt": 1500,
  "offers": [...],
  "datasetFingerprint": "a1b2c3d4...",
  "savedAt": "2026-08-20T14:32:01.123Z"
}

```

## Parallel Fetching with Provider-Aware Concurrency

The script maximizes throughput while respecting host-specific rate limits through tiered concurrency control.

### Worker Pool Configuration

| Provider Type | Concurrency | Rationale |
|-------------|-------------|-----------|
| Single-host (Greenhouse, Lever, Ashby) | `SINGLE_HOST_CONCURRENCY = 6` | Avoid IP-based throttling |
| Distributed (Workday, iCIMS) | `CONCURRENCY = 20` | Higher fan-out across subdomains |

`parallelEach` spawns workers up to `source.concurrency` and routes each company through `source.provider.fetch(entry, ctx)`.

### Per-Company Timeouts and Failure Handling

Each fetch is wrapped with `withTimeout` enforcing `COMPANY_TIMEOUT_MS` (5 minutes). DNS resolver failures are tracked globally—after `RESOLVER_FAILURE_LIMIT` (50) consecutive failures, the sweep halts, writes a final checkpoint at the current offset, and exits cleanly ([lines 84-108](https://github.com/santifer/career-ops/blob/main/scan-ats-full.mjs#L84-L108)).

## Sampling, Shuffling, and Determinism

The `sampleCompanies` function ([lines 41-54](https://github.com/santifer/career-ops/blob/main/scan-ats-full.mjs#L41-L54)) handles list preparation:

- **Default mode**: Preserves source order for deterministic resume
- `--shuffle` mode: Randomizes order (mutually exclusive with `--resume`)

This design trade-off ensures reproducibility when resumability is needed, while allowing random exploration for one-off scans.

## Job-Level Filtering and Deduplication

Fetched jobs pass through a multi-stage pipeline before entering results:

1. **Date classification** (`classifyPostingDate`) — Filters by `sinceDays`, optionally retaining undated postings with `--include-undated`
2. **Title filter** (`buildTitleFilter`) — Matches against `title_filter` regexes from [`portals.yml`](https://github.com/santifer/career-ops/blob/main/portals.yml)
3. **Location filter** (`buildLocationFilter`) — Geographic constraints from configuration
4. **Content filter** (`buildContentFilter`) — Keyword requirements in description
5. **URL deduplication** (`normalizeUrlForDedup`) — Prevents duplicate postings across ATS migrations

Matching jobs append to `newOffers` for final processing ([lines 134-170](https://github.com/santifer/career-ops/blob/main/scan-ats-full.mjs#L134-L170)).

## Handling Truncated Boards: Sequential Retry

Some ATS implementations (Workday, iCIMS) return **truncated** job listings due to internal pagination limits or anti-scraping measures. These boards are flagged during the parallel phase and re-processed **sequentially** after the main sweep completes ([lines 282-304](https://github.com/santifer/career-ops/blob/main/scan-ats-full.mjs#L282-L304)).

This two-pass approach maintains high throughput for well-behaved providers while ensuring completeness for problematic ones.

## VC Seed Scanning Extension

With the `--seeds` flag, `runSeedScan` ([lines 76-118](https://github.com/santifer/career-ops/blob/main/scan-ats-full.mjs#L76-L118)) extends coverage to venture-capital portfolio lists defined in `seeds/vc-portfolios.mjs`:

- Y Combinator batches
- Andreessen Horowitz portfolio
- Other seed-stage company aggregators

Seed sources use the same provider detection and filtering pipeline, with results merged into the main `newOffers` collection.

## Post-Processing and Output Generation

After all ATS and seed scans complete, the script executes final quality gates ([lines 610-622](https://github.com/santifer/career-ops/blob/main/scan-ats-full.mjs#L610-L622)):

1. **Blacklist filtering** (`filterBlacklistedOffers`) — Removes companies in exclusion list
2. **Liveness verification** (`filterLive`) — Optional Playwright headless browser checks that job URLs still resolve
3. **Date sorting** — Most recent postings first
4. **Output writes**:
   - [`data/pipeline.md`](https://github.com/santifer/career-ops/blob/main/data/pipeline.md) — Append-only markdown of new opportunities
   - `data/scan-history.tsv` — Tabular log for analytics
   - `--md-out <path>` — Optional custom digest directory

## Complete Usage Examples

```bash

# Standard daily scan with resume capability

node scan-ats-full.mjs --since 1

# Full dataset sweep with verification and custom output

node scan-ats-full.mjs --since 7 --liveness --md-out reports/weekly

# Dry-run to preview matches without persistence

node scan-ats-full.mjs --dry-run --ats lever,ashby

# Resume after network interruption

node scan-ats-full.mjs --resume

# Include undated postings for aggressive coverage

node scan-ats-full.mjs --include-undated --limit 1000

```

## Summary

- **`scan-ats-full.mjs`** combines **checkpoint persistence**, **parallel fetching**, and **ATS-specific provider logic** to scan tens of thousands of public job boards reliably
- **Atomic checkpoint writes** every 500 companies ensure crash safety without filesystem corruption
- **SHA-1 dataset fingerprinting** prevents silent resume after upstream data changes
- **Tiered concurrency** (6 for single-host, 20 for distributed providers) optimizes throughput while respecting rate limits
- **Two-phase processing** (parallel sweep + sequential retry) handles truncated boards without sacrificing overall speed
- **Modular provider architecture** in `providers/*.mjs` enables clean extension to new ATS platforms

## Frequently Asked Questions

### How does the checkpoint system handle dataset changes from the job-board-aggregator source?

The checkpoint stores a SHA-1 hash (`datasetFingerprint`) of the company list at scan start. On resume, `checkpointCompatible` compares this hash against the current dataset. Mismatches force a fresh scan, preventing incomplete results from stale resumption.

### Why does `--shuffle` disable the resume capability?

Shuffling randomizes company order, which destroys the deterministic offset (`resumeAt`) used for checkpoint positioning. A resumed shuffled scan would skip unpredictable subsets or re-process already-seen companies. The script enforces this mutual exclusion to maintain data integrity.

### What triggers the DNS resolver failure abort?

After `RESOLVER_FAILURE_LIMIT` (50) consecutive DNS failures, the script assumes systemic network or configuration issues. It writes a final checkpoint at the current company offset and exits. This preserves progress before potential cascading failures corrupt state or trigger rate-limit penalties.

### How are truncated Workday and iCIMS boards handled differently?

These providers may paginate or throttle aggressively, returning partial results. The parallel phase flags such boards as truncated. After the main sweep, a sequential retry processes each flagged board with slower, more patient fetching—ensuring completeness without degrading overall scan performance.