How `scan-ats-full.mjs` Performs Checkpoint‑Based Reverse‑ATS Scanning Across Full Datasets in Career‑Ops
scan-ats-full.mjs implements a resumable, checkpoint-driven pipeline that scans public job-board-aggregators for major ATS providers, filters matches by date and criteria, and survives interruptions without re-processing companies.
This deep dive examines the architecture of checkpoint-based reverse-ATS scanning as implemented in the santifer/career-ops repository. The script orchestrates multi-hour sweeps across Greenhouse, Lever, Ashby, Workday, and iCIMS datasets while ensuring crash safety and data integrity through atomic checkpointing and dataset fingerprinting.
Core Architecture: The Checkpoint-Driven Pipeline
The scanning engine follows a deterministic, resumable workflow designed for reliability at scale. Every significant state change is persisted, allowing users to interrupt and resume long-running jobs without loss of progress.
CLI Configuration and Flag Validation
Entry-point flags are parsed and validated via validateFlags in lib/cli-flags.mjs (lines 31-62). Key options include:
--since <days>— Date window for fresh postings (default: 3 days)--limit <n>— Cap companies per ATS source--ats <list>— Target specific providers (greenhouse, lever, ashby, workday, icims)--resume— Continue from last checkpoint--shuffle— Randomize company order (disables resume)--include-undated— Retain postings without parseable dates
# Resume an interrupted 10,000-company sweep
node scan-ats-full.mjs --resume
# Target only Greenhouse with custom date window
node scan-ats-full.mjs --ats greenhouse --since 7 --limit 500
Dataset Loading and Cache Management
The script fetches cached JSON company lists from the public job-board-aggregator repository via loadCompanyList (lines 31-40). Cache entries expire after CACHE_TTL_HOURS (24 hours) to balance freshness with API rate limits.
Each ATS provider maps to a configuration in SOURCES[name] defining:
provider— Fetch implementation moduleconcurrency— Worker pool size for parallel requests
Checkpoint Mechanics: Crash-Safe State Persistence
The checkpoint system is the foundation of resumable operation. Checkpoints are written to data/cache/ats-full-checkpoint.json and validated on resume to prevent silent data corruption.
Checkpoint Compatibility Verification
When --resume is specified, loadCheckpoint() (lines 96-118) performs strict compatibility checks:
// Pseudo-structure from source analysis
checkpointCompatible(current, checkpoint) {
return (
arraysEqual(current.ats, checkpoint.ats) &&
current.limit === checkpoint.limit &&
current.includeUndated === checkpoint.includeUndated &&
hashDataset(current.dataset) === checkpoint.datasetFingerprint
)
}
The SHA-1 fingerprint (datasetFingerprint) detects any drift in the underlying company list. If the dataset hash mismatches, the resume aborts—forcing a fresh scan rather than risking incomplete or inconsistent results.
Atomic Checkpoint Writes
Checkpoints are written every CHECKPOINT_EVERY (500) companies via writeCheckpoint() (lines 31-45):
// Atomic write pattern from source
const tmpPath = `${CHECKPOINT_PATH}.tmp`;
fs.writeFileSync(tmpPath, JSON.stringify(checkpoint));
fs.renameSync(tmpPath, CHECKPOINT_PATH); // Atomic on POSIX
This write-to-temporary-then-rename pattern guarantees that a crash during checkpoint serialization cannot leave a corrupt partial file.
Checkpoint Schema
{
"version": 1,
"cutoffMs": 1724161921123,
"ats": ["greenhouse", "lever", "ashby"],
"limit": 200,
"includeUndated": false,
"completedSources": ["greenhouse"],
"resumeAt": 1500,
"offers": [...],
"datasetFingerprint": "a1b2c3d4...",
"savedAt": "2026-08-20T14:32:01.123Z"
}
Parallel Fetching with Provider-Aware Concurrency
The script maximizes throughput while respecting host-specific rate limits through tiered concurrency control.
Worker Pool Configuration
| Provider Type | Concurrency | Rationale |
|---|---|---|
| Single-host (Greenhouse, Lever, Ashby) | SINGLE_HOST_CONCURRENCY = 6 |
Avoid IP-based throttling |
| Distributed (Workday, iCIMS) | CONCURRENCY = 20 |
Higher fan-out across subdomains |
parallelEach spawns workers up to source.concurrency and routes each company through source.provider.fetch(entry, ctx).
Per-Company Timeouts and Failure Handling
Each fetch is wrapped with withTimeout enforcing COMPANY_TIMEOUT_MS (5 minutes). DNS resolver failures are tracked globally—after RESOLVER_FAILURE_LIMIT (50) consecutive failures, the sweep halts, writes a final checkpoint at the current offset, and exits cleanly (lines 84-108).
Sampling, Shuffling, and Determinism
The sampleCompanies function (lines 41-54) handles list preparation:
- Default mode: Preserves source order for deterministic resume
--shufflemode: Randomizes order (mutually exclusive with--resume)
This design trade-off ensures reproducibility when resumability is needed, while allowing random exploration for one-off scans.
Job-Level Filtering and Deduplication
Fetched jobs pass through a multi-stage pipeline before entering results:
- Date classification (
classifyPostingDate) — Filters bysinceDays, optionally retaining undated postings with--include-undated - Title filter (
buildTitleFilter) — Matches againsttitle_filterregexes fromportals.yml - Location filter (
buildLocationFilter) — Geographic constraints from configuration - Content filter (
buildContentFilter) — Keyword requirements in description - URL deduplication (
normalizeUrlForDedup) — Prevents duplicate postings across ATS migrations
Matching jobs append to newOffers for final processing (lines 134-170).
Handling Truncated Boards: Sequential Retry
Some ATS implementations (Workday, iCIMS) return truncated job listings due to internal pagination limits or anti-scraping measures. These boards are flagged during the parallel phase and re-processed sequentially after the main sweep completes (lines 282-304).
This two-pass approach maintains high throughput for well-behaved providers while ensuring completeness for problematic ones.
VC Seed Scanning Extension
With the --seeds flag, runSeedScan (lines 76-118) extends coverage to venture-capital portfolio lists defined in seeds/vc-portfolios.mjs:
- Y Combinator batches
- Andreessen Horowitz portfolio
- Other seed-stage company aggregators
Seed sources use the same provider detection and filtering pipeline, with results merged into the main newOffers collection.
Post-Processing and Output Generation
After all ATS and seed scans complete, the script executes final quality gates (lines 610-622):
- Blacklist filtering (
filterBlacklistedOffers) — Removes companies in exclusion list - Liveness verification (
filterLive) — Optional Playwright headless browser checks that job URLs still resolve - Date sorting — Most recent postings first
- Output writes:
data/pipeline.md— Append-only markdown of new opportunitiesdata/scan-history.tsv— Tabular log for analytics--md-out <path>— Optional custom digest directory
Complete Usage Examples
# Standard daily scan with resume capability
node scan-ats-full.mjs --since 1
# Full dataset sweep with verification and custom output
node scan-ats-full.mjs --since 7 --liveness --md-out reports/weekly
# Dry-run to preview matches without persistence
node scan-ats-full.mjs --dry-run --ats lever,ashby
# Resume after network interruption
node scan-ats-full.mjs --resume
# Include undated postings for aggressive coverage
node scan-ats-full.mjs --include-undated --limit 1000
Summary
scan-ats-full.mjscombines checkpoint persistence, parallel fetching, and ATS-specific provider logic to scan tens of thousands of public job boards reliably- Atomic checkpoint writes every 500 companies ensure crash safety without filesystem corruption
- SHA-1 dataset fingerprinting prevents silent resume after upstream data changes
- Tiered concurrency (6 for single-host, 20 for distributed providers) optimizes throughput while respecting rate limits
- Two-phase processing (parallel sweep + sequential retry) handles truncated boards without sacrificing overall speed
- Modular provider architecture in
providers/*.mjsenables clean extension to new ATS platforms
Frequently Asked Questions
How does the checkpoint system handle dataset changes from the job-board-aggregator source?
The checkpoint stores a SHA-1 hash (datasetFingerprint) of the company list at scan start. On resume, checkpointCompatible compares this hash against the current dataset. Mismatches force a fresh scan, preventing incomplete results from stale resumption.
Why does --shuffle disable the resume capability?
Shuffling randomizes company order, which destroys the deterministic offset (resumeAt) used for checkpoint positioning. A resumed shuffled scan would skip unpredictable subsets or re-process already-seen companies. The script enforces this mutual exclusion to maintain data integrity.
What triggers the DNS resolver failure abort?
After RESOLVER_FAILURE_LIMIT (50) consecutive DNS failures, the script assumes systemic network or configuration issues. It writes a final checkpoint at the current company offset and exits. This preserves progress before potential cascading failures corrupt state or trigger rate-limit penalties.
How are truncated Workday and iCIMS boards handled differently?
These providers may paginate or throttle aggressively, returning partial results. The parallel phase flags such boards as truncated. After the main sweep, a sequential retry processes each flagged board with slower, more patient fetching—ensuring completeness without degrading overall scan performance.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →