How `report-number reservation` Prevents Race Conditions in Parallel Batch Evaluation in santifer/career-ops
The career-ops project prevents race conditions during parallel batch evaluation by using a shared atomic report-number allocator in reserve-report-num.mjs that combines exclusive filesystem sentinel files, a global tracker lock, and ownership tokens with automatic retry logic.
When running multiple evaluator processes simultaneously—such as the pipeline command that evaluates dozens of job postings at once—each process must obtain a unique report identifier that persists consistently across the tracker file and the filesystem. The santifer/career-ops repository solves this coordination problem through a carefully designed allocation mechanism that guarantees monotonic, collision-free report numbering even under heavy concurrency.
Three Coordinated Mechanisms for Atomic Allocation
The reserveReportNumbers() function in reserve-report-num.mjs implements three complementary safeguards that work together to eliminate race conditions:
1. Exclusive Filesystem Sentinel Files with Atomic Creation
Each candidate report number is claimed by writing a sentinel file named NNN-RESERVED.md. The allocator uses writeFileSync() with the O_CREAT|O_EXCL flag ({ flag: 'wx' }), which guarantees atomic failure if another process has already created that file. This filesystem-level primitive ensures that only one process can ever own a given number, regardless of timing or process count.
// From reserve-report-num.mjs – the core atomic claim operation
await fs.writeFile(
path.join(reservationsDir, `${num}-RESERVED.md`),
`Reserved by ${token} at ${new Date().toISOString()}\n`,
{ flag: 'wx' } // Fails atomically if file exists
);
2. Global Tracker Lock for Serializable Reads
Before computing which numbers are available, the allocator acquires an exclusive lock via acquireTrackerLock() (implemented in tracker-utils.mjs). This lock serializes access to both the tracker file and the reservation directory, preventing the classic "time-of-check to time-of-use" race where two processes simultaneously read the same "next free" number.
3. Ownership Tokens and Automatic Retry with Rollback
Every reservation carries a UUID token. If a process cannot claim its full requested contiguous block, it immediately releases any partial claims, refreshes the occupied set by re-reading existing reports and the tracker, and retries with a new base number. The token ensures that only the owning process can release its sentinels, protecting against accidental or malicious deletions.
// Reserving a block of report numbers for parallel batch processing
const block = await reserveReportNumbers(8, {
rootDir: __dirname,
reportsDir: path.join(__dirname, 'reports')
});
// Returns: [42, 43, 44, 45, 46, 47, 48, 49] with hidden RESERVATION_TOKEN
Complete Reservation Lifecycle
The allocation process in reserve-report-num.mjs follows a precise sequence to maintain correctness:
- Collect occupied numbers – Scan
reports/directory and parsedata/applications.mdtracker to build aSetof used IDs - Acquire tracker lock – Call
acquireTrackerLock()for exclusive access - Select base candidate – Start from
max(occupied) + 1 - Attempt atomic claims – Loop through desired range, calling
claimSlot()withflag: 'wx' - Rollback and retry on collision – If any slot is taken, release partial claims, refresh occupied set, and retry (up to
MAX_RETRIES) - Return verified array – On success, return numbers with hidden
RESERVATION_TOKENsymbol for authenticated release - Clean release – After writing reports,
releaseReportNumbers()removes sentinels after verifying ownership
Practical Usage in the Evaluation Pipeline
The openrouter-runner.mjs file demonstrates the reservation pattern in production use. Before generating any report, the runner reserves its number; after successful write or on failure, it releases the reservation to prevent leakage.
// openrouter-runner.mjs – typical reservation pattern
const reservedNumbers = await reserveReportNumbers(1, {
rootDir: __dirname,
reportsDir: path.join(__dirname, 'reports')
});
const reportId = reservedNumbers[0];
try {
// ... generate and write report to reports/${reportId}.md ...
} finally {
await releaseReportNumbers(reservedNumbers, {
reportsDir: path.join(__dirname, 'reports')
});
}
For parallel batches, each worker reserves a distinct block:
// Worker pool pattern: reserve contiguous block, process in parallel
const myBlock = await reserveReportNumbers(BATCH_SIZE, { rootDir, reportsDir });
await Promise.all(myBlock.map(id => evaluateAndWrite(id)));
await releaseReportNumbers(myBlock, { reportsDir });
Automatic Cleanup of Stale Reservations
Crashed or hung processes could leave orphaned sentinel files. The gcStaleReportReservations() function addresses this by removing sentinel files older than four hours if the owning process is no longer alive. This garbage collection can run periodically or on startup to reclaim leaked numbers.
// Startup safety: clean old reservations before beginning batch
await gcStaleReportReservations(); // Removes stale *-RESERVED.md files
Summary
- Atomic filesystem operations (
flag: 'wx') guarantee exclusive ownership of individual report numbers - Global tracker lock (
acquireTrackerLock) serializes the "find next free" computation across all processes - UUID ownership tokens enable authenticated release and prevent unauthorized deletion
- Automatic rollback and retry with refreshed state ensures forward progress under contention
- Garbage collection reclaims numbers from crashed workers after a timeout
- Contiguous block reservation supports efficient parallel batch evaluation without ID fragmentation
Frequently Asked Questions
What happens if two processes try to reserve the same report number simultaneously?
Only one succeeds. The writeFileSync with flag: 'wx' fails atomically for the second process, which then releases any partial claims, refreshes its view of occupied numbers, and retries with a new candidate. The global tracker lock ensures both processes cannot be computing "next free" at the same time.
Why use filesystem sentinels instead of a database or in-memory store?
Filesystem sentinels provide durability across process restarts, require no external dependencies, and leverage kernel-level atomicity guarantees. This matches career-ops's design philosophy of using git-trackable Markdown files for all persistent state.
How does the garbage collector know if a reservation is truly stale?
gcStaleReportReservations() checks both the file modification time (four-hour threshold) and whether the process that created the sentinel is still running. This dual check prevents deleting reservations from legitimate long-running batches.
Can I reserve non-contiguous report numbers for better packing?
The current allocator in reserve-report-num.mjs optimizes for contiguous blocks to support parallel batch evaluation efficiently. While the underlying mechanism could support scattered allocation, the public API emphasizes ranges that map cleanly to worker pool distribution.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →