# How the Career-Ops Batch Runner Orchestrates Parallel Headless Evaluations

> Discover how the Career-Ops batch runner orchestrates parallel headless evaluations using a shell-based system. Learn about atomic reservations, file locking, and recovery for resumability.

- Repository: [Santiago Fernández de Valderrama/career-ops](https://github.com/santifer/career-ops)
- Tags: internals
- Published: 2026-08-19

---

**The batch runner uses a shell-based orchestrator that spawns multiple headless CLAUDE-Code workers in parallel**, coordinating them through atomic report-number reservations, file-based locking on a shared state file, and recovery records for resumability.

The `career-ops` repository automates the evaluation of job postings using a sophisticated batch processing system. When processing thousands of job descriptions simultaneously, the **[`batch/batch-runner.sh`](https://github.com/santifer/career-ops/blob/main/batch/batch-runner.sh)** script manages parallel execution by orchestrating independent headless workers that share state safely without race conditions.

## Parsing the Parallelism Configuration

The orchestration begins with argument parsing that determines the worker pool size. The script accepts a **`--parallel N`** flag at line 63 of [`batch/batch-runner.sh`](https://github.com/santifer/career-ops/blob/main/batch/batch-runner.sh), which controls how many simultaneous workers run against the input file. The default value is `1`, ensuring safe sequential execution unless explicitly overridden.

When `N` is greater than one, the runner activates parallel mode, which triggers specific optimizations including shared browser instances and mandatory state-file locking.

## Atomic Report Number Reservation

Before spawning any workers, the system must guarantee unique identifiers for each job evaluation. The script calls **`reserve-report-num.mjs`** to atomically claim a contiguous block of report numbers from a central counter.

The reservation logic implements two functions in [`batch/batch-runner.sh`](https://github.com/santifer/career-ops/blob/main/batch/batch-runner.sh):
- **`reserve_report_num_unlocked`** (line 668): The core reservation logic that reads the current counter and increments it
- **`reserve_report_num`** (line 699-704): A public wrapper that handles retry logic and synchronization

Each worker receives a unique report number from this reserved block, ensuring that parallel processes never generate colliding identifiers when writing to shared output files.

## Spawning Headless Workers

With report numbers allocated, the runner enters the main processing loop around lines 725-731, labeled *"Process a single offer"* in the source comments. For each URL in `batch-input.tsv`, the script launches a background worker using the **`claude -p`** command:

```bash
claude -p --model "$MODEL" --append-system-prompt-file batch-prompt.md \
        --dangerously-skip-permissions "$url" "$report_num" &

```

The worker executes in **headless mode** (no UI), receiving its unique `$report_num` and the job description file path as arguments. When `--parallel` exceeds 1, workers share a single browser instance (lines 800-802), significantly reducing memory overhead compared to launching separate browser processes per worker.

## State Coordination via File Locking

All workers synchronize through **`batch-state.tsv`**, a persistent tab-separated file tracking each job's status, score, and completion state. To prevent race conditions during parallel writes, the runner implements a **lock directory mechanism** (`batch-runner.pid`).

The lock acquisition logic at line 143 creates an exclusive lock before any read-modify-write cycle. If the lock remains unavailable beyond a timeout threshold, the system triggers lock-timeout handling (lines 204-208) to prevent indefinite blocking while maintaining data integrity.

Workers write progress updates to the shared state file only while holding this lock, ensuring that status transitions remain atomic even with multiple concurrent writers.

## Recovery Records for Lock Contention

When a worker cannot acquire the lock—typically because another process holds it—the runner implements a **recovery record** mechanism (lines 474-488) to prevent state loss. Instead of blocking indefinitely or dropping the update, the worker:

1. Creates a temporary file using `mktemp` with `O_CREAT|O_EXCL` flags to ensure uniqueness
2. Writes the pending state transition to this recovery file in a designated recovery directory
3. Exits cleanly without corrupting the shared state

At the start of the next run, the system automatically merges these recovery records back into `batch-state.tsv` before processing new jobs. This design guarantees **resumability** even if the batch process terminates unexpectedly or encounters high lock contention.

## Post-Run Aggregation and Validation

Once all background workers complete, the runner executes a deterministic aggregation pipeline (lines 981-994). The **`merge-tracker.mjs`** script consolidates TSV rows written by individual workers into the canonical tracker at [`data/applications.md`](https://github.com/santifer/career-ops/blob/main/data/applications.md).

Following the merge, the runner executes sanity-check scripts to ensure data consistency:
- **`verify-pipeline.mjs`**: Validates the integrity of merged records
- **`normalize-statuses.mjs`**: Standardizes status fields across entries
- **`dedup-tracker.mjs`**: Removes duplicate entries that may occur from retried operations

## Summary

The `career-ops` batch runner orchestrates parallel headless evaluations through:

- **Configuration-driven parallelism** via the `--parallel` flag that controls worker pool size
- **Atomic reservation** of unique report numbers through `reserve-report-num.mjs` to prevent identifier collisions
- **Headless worker spawning** using `claude -p` commands that share browser instances when running concurrently
- **File-based locking** on `batch-state.tsv` to serialize state updates and prevent race conditions
- **Recovery record system** that writes pending updates to temporary files during lock contention, merged back on subsequent runs
- **Validated aggregation** through `merge-tracker.mjs` and verification scripts that consolidate outputs into the master tracker

## Frequently Asked Questions

### How does the batch runner prevent duplicate report numbers when running in parallel?

The runner calls `reserve-report-num.mjs` before spawning any workers, atomically reserving a contiguous block of report numbers. Each worker receives a unique number from this block through the `reserve_report_num` function (lines 699-704), ensuring no two workers ever claim the same identifier even when launched simultaneously.

### What happens if a worker cannot acquire the lock on the state file?

If the lock on `batch-state.tsv` is unavailable, the worker writes a **recovery record** to a uniquely named temporary file in the recovery directory (lines 474-488). These records capture the pending state transition and are merged back into the main state file at the beginning of the next run, guaranteeing no data loss occurs during high-contention scenarios.

### Why do parallel workers share a single browser instance?

When `--parallel` is greater than 1, the runner configures workers to share a single browser instance (lines 800-802) to minimize memory consumption. This optimization prevents the resource exhaustion that would occur if each headless worker launched its own separate browser process while evaluating thousands of job postings.

### Can the batch process resume if interrupted mid-run?

Yes, the system is fully resumable. The `batch-state.tsv` file persists the status of every job, and recovery records capture any updates that couldn't be written due to lock contention or crashes. When restarted, the runner reads this state file and skips jobs already marked complete, processing only pending or failed entries from `batch-input.tsv`.