How `scan-ats-full.mjs` Performs Checkpoint‑Based Reverse‑ATS Scanning Across Full Datasets in Career‑Ops

scan-ats-full.mjs implements a resumable, checkpoint-driven pipeline that scans public job-board-aggregators for major ATS providers, filters matches by date and criteria, and survives interruptions without re-processing companies.

This deep dive examines the architecture of checkpoint-based reverse-ATS scanning as implemented in the santifer/career-ops repository. The script orchestrates multi-hour sweeps across Greenhouse, Lever, Ashby, Workday, and iCIMS datasets while ensuring crash safety and data integrity through atomic checkpointing and dataset fingerprinting.

Core Architecture: The Checkpoint-Driven Pipeline

The scanning engine follows a deterministic, resumable workflow designed for reliability at scale. Every significant state change is persisted, allowing users to interrupt and resume long-running jobs without loss of progress.

CLI Configuration and Flag Validation

Entry-point flags are parsed and validated via validateFlags in lib/cli-flags.mjs (lines 31-62). Key options include:

  • --since <days> — Date window for fresh postings (default: 3 days)
  • --limit <n> — Cap companies per ATS source
  • --ats <list> — Target specific providers (greenhouse, lever, ashby, workday, icims)
  • --resume — Continue from last checkpoint
  • --shuffle — Randomize company order (disables resume)
  • --include-undated — Retain postings without parseable dates

# Resume an interrupted 10,000-company sweep

node scan-ats-full.mjs --resume

# Target only Greenhouse with custom date window

node scan-ats-full.mjs --ats greenhouse --since 7 --limit 500

Dataset Loading and Cache Management

The script fetches cached JSON company lists from the public job-board-aggregator repository via loadCompanyList (lines 31-40). Cache entries expire after CACHE_TTL_HOURS (24 hours) to balance freshness with API rate limits.

Each ATS provider maps to a configuration in SOURCES[name] defining:

  • provider — Fetch implementation module
  • concurrency — Worker pool size for parallel requests

Checkpoint Mechanics: Crash-Safe State Persistence

The checkpoint system is the foundation of resumable operation. Checkpoints are written to data/cache/ats-full-checkpoint.json and validated on resume to prevent silent data corruption.

Checkpoint Compatibility Verification

When --resume is specified, loadCheckpoint() (lines 96-118) performs strict compatibility checks:

// Pseudo-structure from source analysis
checkpointCompatible(current, checkpoint) {
  return (
    arraysEqual(current.ats, checkpoint.ats) &&
    current.limit === checkpoint.limit &&
    current.includeUndated === checkpoint.includeUndated &&
    hashDataset(current.dataset) === checkpoint.datasetFingerprint
  )
}

The SHA-1 fingerprint (datasetFingerprint) detects any drift in the underlying company list. If the dataset hash mismatches, the resume aborts—forcing a fresh scan rather than risking incomplete or inconsistent results.

Atomic Checkpoint Writes

Checkpoints are written every CHECKPOINT_EVERY (500) companies via writeCheckpoint() (lines 31-45):

// Atomic write pattern from source
const tmpPath = `${CHECKPOINT_PATH}.tmp`;
fs.writeFileSync(tmpPath, JSON.stringify(checkpoint));
fs.renameSync(tmpPath, CHECKPOINT_PATH);  // Atomic on POSIX

This write-to-temporary-then-rename pattern guarantees that a crash during checkpoint serialization cannot leave a corrupt partial file.

Checkpoint Schema

{
  "version": 1,
  "cutoffMs": 1724161921123,
  "ats": ["greenhouse", "lever", "ashby"],
  "limit": 200,
  "includeUndated": false,
  "completedSources": ["greenhouse"],
  "resumeAt": 1500,
  "offers": [...],
  "datasetFingerprint": "a1b2c3d4...",
  "savedAt": "2026-08-20T14:32:01.123Z"
}

Parallel Fetching with Provider-Aware Concurrency

The script maximizes throughput while respecting host-specific rate limits through tiered concurrency control.

Worker Pool Configuration

Provider Type Concurrency Rationale
Single-host (Greenhouse, Lever, Ashby) SINGLE_HOST_CONCURRENCY = 6 Avoid IP-based throttling
Distributed (Workday, iCIMS) CONCURRENCY = 20 Higher fan-out across subdomains

parallelEach spawns workers up to source.concurrency and routes each company through source.provider.fetch(entry, ctx).

Per-Company Timeouts and Failure Handling

Each fetch is wrapped with withTimeout enforcing COMPANY_TIMEOUT_MS (5 minutes). DNS resolver failures are tracked globally—after RESOLVER_FAILURE_LIMIT (50) consecutive failures, the sweep halts, writes a final checkpoint at the current offset, and exits cleanly (lines 84-108).

Sampling, Shuffling, and Determinism

The sampleCompanies function (lines 41-54) handles list preparation:

  • Default mode: Preserves source order for deterministic resume
  • --shuffle mode: Randomizes order (mutually exclusive with --resume)

This design trade-off ensures reproducibility when resumability is needed, while allowing random exploration for one-off scans.

Job-Level Filtering and Deduplication

Fetched jobs pass through a multi-stage pipeline before entering results:

  1. Date classification (classifyPostingDate) — Filters by sinceDays, optionally retaining undated postings with --include-undated
  2. Title filter (buildTitleFilter) — Matches against title_filter regexes from portals.yml
  3. Location filter (buildLocationFilter) — Geographic constraints from configuration
  4. Content filter (buildContentFilter) — Keyword requirements in description
  5. URL deduplication (normalizeUrlForDedup) — Prevents duplicate postings across ATS migrations

Matching jobs append to newOffers for final processing (lines 134-170).

Handling Truncated Boards: Sequential Retry

Some ATS implementations (Workday, iCIMS) return truncated job listings due to internal pagination limits or anti-scraping measures. These boards are flagged during the parallel phase and re-processed sequentially after the main sweep completes (lines 282-304).

This two-pass approach maintains high throughput for well-behaved providers while ensuring completeness for problematic ones.

VC Seed Scanning Extension

With the --seeds flag, runSeedScan (lines 76-118) extends coverage to venture-capital portfolio lists defined in seeds/vc-portfolios.mjs:

  • Y Combinator batches
  • Andreessen Horowitz portfolio
  • Other seed-stage company aggregators

Seed sources use the same provider detection and filtering pipeline, with results merged into the main newOffers collection.

Post-Processing and Output Generation

After all ATS and seed scans complete, the script executes final quality gates (lines 610-622):

  1. Blacklist filtering (filterBlacklistedOffers) — Removes companies in exclusion list
  2. Liveness verification (filterLive) — Optional Playwright headless browser checks that job URLs still resolve
  3. Date sorting — Most recent postings first
  4. Output writes:
    • data/pipeline.md — Append-only markdown of new opportunities
    • data/scan-history.tsv — Tabular log for analytics
    • --md-out <path> — Optional custom digest directory

Complete Usage Examples


# Standard daily scan with resume capability

node scan-ats-full.mjs --since 1

# Full dataset sweep with verification and custom output

node scan-ats-full.mjs --since 7 --liveness --md-out reports/weekly

# Dry-run to preview matches without persistence

node scan-ats-full.mjs --dry-run --ats lever,ashby

# Resume after network interruption

node scan-ats-full.mjs --resume

# Include undated postings for aggressive coverage

node scan-ats-full.mjs --include-undated --limit 1000

Summary

  • scan-ats-full.mjs combines checkpoint persistence, parallel fetching, and ATS-specific provider logic to scan tens of thousands of public job boards reliably
  • Atomic checkpoint writes every 500 companies ensure crash safety without filesystem corruption
  • SHA-1 dataset fingerprinting prevents silent resume after upstream data changes
  • Tiered concurrency (6 for single-host, 20 for distributed providers) optimizes throughput while respecting rate limits
  • Two-phase processing (parallel sweep + sequential retry) handles truncated boards without sacrificing overall speed
  • Modular provider architecture in providers/*.mjs enables clean extension to new ATS platforms

Frequently Asked Questions

How does the checkpoint system handle dataset changes from the job-board-aggregator source?

The checkpoint stores a SHA-1 hash (datasetFingerprint) of the company list at scan start. On resume, checkpointCompatible compares this hash against the current dataset. Mismatches force a fresh scan, preventing incomplete results from stale resumption.

Why does --shuffle disable the resume capability?

Shuffling randomizes company order, which destroys the deterministic offset (resumeAt) used for checkpoint positioning. A resumed shuffled scan would skip unpredictable subsets or re-process already-seen companies. The script enforces this mutual exclusion to maintain data integrity.

What triggers the DNS resolver failure abort?

After RESOLVER_FAILURE_LIMIT (50) consecutive DNS failures, the script assumes systemic network or configuration issues. It writes a final checkpoint at the current company offset and exits. This preserves progress before potential cascading failures corrupt state or trigger rate-limit penalties.

How are truncated Workday and iCIMS boards handled differently?

These providers may paginate or throttle aggressively, returning partial results. The parallel phase flags such boards as truncated. After the main sweep, a sequential retry processes each flagged board with slower, more patient fetching—ensuring completeness without degrading overall scan performance.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →