# How `report-number reservation` Prevents Race Conditions in Parallel Batch Evaluation in santifer/career-ops

> Discover how santifer/career-ops uses report-number reservation with atomic allocation, sentinel files, and locks to prevent race conditions in parallel batch evaluation. Learn about this robust solution.

- Repository: [Santiago Fernández de Valderrama/career-ops](https://github.com/santifer/career-ops)
- Tags: internals
- Published: 2026-08-20

---

**The `career-ops` project prevents race conditions during parallel batch evaluation by using a shared atomic report-number allocator in `reserve-report-num.mjs` that combines exclusive filesystem sentinel files, a global tracker lock, and ownership tokens with automatic retry logic.**

When running multiple evaluator processes simultaneously—such as the `pipeline` command that evaluates dozens of job postings at once—each process must obtain a unique report identifier that persists consistently across the tracker file and the filesystem. The `santifer/career-ops` repository solves this coordination problem through a carefully designed allocation mechanism that guarantees monotonic, collision-free report numbering even under heavy concurrency.

## Three Coordinated Mechanisms for Atomic Allocation

The `reserveReportNumbers()` function in `reserve-report-num.mjs` implements three complementary safeguards that work together to eliminate race conditions:

### 1. Exclusive Filesystem Sentinel Files with Atomic Creation

Each candidate report number is claimed by writing a sentinel file named [`NNN-RESERVED.md`](https://github.com/santifer/career-ops/blob/main/NNN-RESERVED.md). The allocator uses `writeFileSync()` with the `O_CREAT|O_EXCL` flag (`{ flag: 'wx' }`), which guarantees atomic failure if another process has already created that file. This filesystem-level primitive ensures that only one process can ever own a given number, regardless of timing or process count.

```javascript
// From reserve-report-num.mjs – the core atomic claim operation
await fs.writeFile(
  path.join(reservationsDir, `${num}-RESERVED.md`),
  `Reserved by ${token} at ${new Date().toISOString()}\n`,
  { flag: 'wx' }  // Fails atomically if file exists
);

```

### 2. Global Tracker Lock for Serializable Reads

Before computing which numbers are available, the allocator acquires an exclusive lock via `acquireTrackerLock()` (implemented in `tracker-utils.mjs`). This lock serializes access to both the tracker file and the reservation directory, preventing the classic "time-of-check to time-of-use" race where two processes simultaneously read the same "next free" number.

### 3. Ownership Tokens and Automatic Retry with Rollback

Every reservation carries a UUID token. If a process cannot claim its full requested contiguous block, it immediately releases any partial claims, refreshes the occupied set by re-reading existing reports and the tracker, and retries with a new base number. The token ensures that only the owning process can release its sentinels, protecting against accidental or malicious deletions.

```javascript
// Reserving a block of report numbers for parallel batch processing
const block = await reserveReportNumbers(8, {
  rootDir: __dirname,
  reportsDir: path.join(__dirname, 'reports')
});
// Returns: [42, 43, 44, 45, 46, 47, 48, 49] with hidden RESERVATION_TOKEN

```

## Complete Reservation Lifecycle

The allocation process in `reserve-report-num.mjs` follows a precise sequence to maintain correctness:

1. **Collect occupied numbers** – Scan `reports/` directory and parse [`data/applications.md`](https://github.com/santifer/career-ops/blob/main/data/applications.md) tracker to build a `Set` of used IDs
2. **Acquire tracker lock** – Call `acquireTrackerLock()` for exclusive access
3. **Select base candidate** – Start from `max(occupied) + 1`
4. **Attempt atomic claims** – Loop through desired range, calling `claimSlot()` with `flag: 'wx'`
5. **Rollback and retry on collision** – If any slot is taken, release partial claims, refresh occupied set, and retry (up to `MAX_RETRIES`)
6. **Return verified array** – On success, return numbers with hidden `RESERVATION_TOKEN` symbol for authenticated release
7. **Clean release** – After writing reports, `releaseReportNumbers()` removes sentinels after verifying ownership

## Practical Usage in the Evaluation Pipeline

The `openrouter-runner.mjs` file demonstrates the reservation pattern in production use. Before generating any report, the runner reserves its number; after successful write or on failure, it releases the reservation to prevent leakage.

```javascript
// openrouter-runner.mjs – typical reservation pattern
const reservedNumbers = await reserveReportNumbers(1, {
  rootDir: __dirname,
  reportsDir: path.join(__dirname, 'reports')
});
const reportId = reservedNumbers[0];

try {
  // ... generate and write report to reports/${reportId}.md ...
} finally {
  await releaseReportNumbers(reservedNumbers, {
    reportsDir: path.join(__dirname, 'reports')
  });
}

```

For parallel batches, each worker reserves a distinct block:

```javascript
// Worker pool pattern: reserve contiguous block, process in parallel
const myBlock = await reserveReportNumbers(BATCH_SIZE, { rootDir, reportsDir });
await Promise.all(myBlock.map(id => evaluateAndWrite(id)));
await releaseReportNumbers(myBlock, { reportsDir });

```

## Automatic Cleanup of Stale Reservations

Crashed or hung processes could leave orphaned sentinel files. The `gcStaleReportReservations()` function addresses this by removing sentinel files older than four hours if the owning process is no longer alive. This garbage collection can run periodically or on startup to reclaim leaked numbers.

```javascript
// Startup safety: clean old reservations before beginning batch
await gcStaleReportReservations();  // Removes stale *-RESERVED.md files

```

## Summary

- **Atomic filesystem operations** (`flag: 'wx'`) guarantee exclusive ownership of individual report numbers
- **Global tracker lock** (`acquireTrackerLock`) serializes the "find next free" computation across all processes
- **UUID ownership tokens** enable authenticated release and prevent unauthorized deletion
- **Automatic rollback and retry** with refreshed state ensures forward progress under contention
- **Garbage collection** reclaims numbers from crashed workers after a timeout
- **Contiguous block reservation** supports efficient parallel batch evaluation without ID fragmentation

## Frequently Asked Questions

### What happens if two processes try to reserve the same report number simultaneously?

Only one succeeds. The `writeFileSync` with `flag: 'wx'` fails atomically for the second process, which then releases any partial claims, refreshes its view of occupied numbers, and retries with a new candidate. The global tracker lock ensures both processes cannot be computing "next free" at the same time.

### Why use filesystem sentinels instead of a database or in-memory store?

Filesystem sentinels provide durability across process restarts, require no external dependencies, and leverage kernel-level atomicity guarantees. This matches `career-ops`'s design philosophy of using git-trackable Markdown files for all persistent state.

### How does the garbage collector know if a reservation is truly stale?

`gcStaleReportReservations()` checks both the file modification time (four-hour threshold) and whether the process that created the sentinel is still running. This dual check prevents deleting reservations from legitimate long-running batches.

### Can I reserve non-contiguous report numbers for better packing?

The current allocator in `reserve-report-num.mjs` optimizes for contiguous blocks to support parallel batch evaluation efficiently. While the underlying mechanism could support scattered allocation, the public API emphasizes ranges that map cleanly to worker pool distribution.