How the Career-Ops Batch Runner Orchestrates Parallel Headless Evaluations

The batch runner uses a shell-based orchestrator that spawns multiple headless CLAUDE-Code workers in parallel, coordinating them through atomic report-number reservations, file-based locking on a shared state file, and recovery records for resumability.

The career-ops repository automates the evaluation of job postings using a sophisticated batch processing system. When processing thousands of job descriptions simultaneously, the batch/batch-runner.sh script manages parallel execution by orchestrating independent headless workers that share state safely without race conditions.

Parsing the Parallelism Configuration

The orchestration begins with argument parsing that determines the worker pool size. The script accepts a --parallel N flag at line 63 of batch/batch-runner.sh, which controls how many simultaneous workers run against the input file. The default value is 1, ensuring safe sequential execution unless explicitly overridden.

When N is greater than one, the runner activates parallel mode, which triggers specific optimizations including shared browser instances and mandatory state-file locking.

Atomic Report Number Reservation

Before spawning any workers, the system must guarantee unique identifiers for each job evaluation. The script calls reserve-report-num.mjs to atomically claim a contiguous block of report numbers from a central counter.

The reservation logic implements two functions in batch/batch-runner.sh:

  • reserve_report_num_unlocked (line 668): The core reservation logic that reads the current counter and increments it
  • reserve_report_num (line 699-704): A public wrapper that handles retry logic and synchronization

Each worker receives a unique report number from this reserved block, ensuring that parallel processes never generate colliding identifiers when writing to shared output files.

Spawning Headless Workers

With report numbers allocated, the runner enters the main processing loop around lines 725-731, labeled "Process a single offer" in the source comments. For each URL in batch-input.tsv, the script launches a background worker using the claude -p command:

claude -p --model "$MODEL" --append-system-prompt-file batch-prompt.md \
        --dangerously-skip-permissions "$url" "$report_num" &

The worker executes in headless mode (no UI), receiving its unique $report_num and the job description file path as arguments. When --parallel exceeds 1, workers share a single browser instance (lines 800-802), significantly reducing memory overhead compared to launching separate browser processes per worker.

State Coordination via File Locking

All workers synchronize through batch-state.tsv, a persistent tab-separated file tracking each job's status, score, and completion state. To prevent race conditions during parallel writes, the runner implements a lock directory mechanism (batch-runner.pid).

The lock acquisition logic at line 143 creates an exclusive lock before any read-modify-write cycle. If the lock remains unavailable beyond a timeout threshold, the system triggers lock-timeout handling (lines 204-208) to prevent indefinite blocking while maintaining data integrity.

Workers write progress updates to the shared state file only while holding this lock, ensuring that status transitions remain atomic even with multiple concurrent writers.

Recovery Records for Lock Contention

When a worker cannot acquire the lock—typically because another process holds it—the runner implements a recovery record mechanism (lines 474-488) to prevent state loss. Instead of blocking indefinitely or dropping the update, the worker:

  1. Creates a temporary file using mktemp with O_CREAT|O_EXCL flags to ensure uniqueness
  2. Writes the pending state transition to this recovery file in a designated recovery directory
  3. Exits cleanly without corrupting the shared state

At the start of the next run, the system automatically merges these recovery records back into batch-state.tsv before processing new jobs. This design guarantees resumability even if the batch process terminates unexpectedly or encounters high lock contention.

Post-Run Aggregation and Validation

Once all background workers complete, the runner executes a deterministic aggregation pipeline (lines 981-994). The merge-tracker.mjs script consolidates TSV rows written by individual workers into the canonical tracker at data/applications.md.

Following the merge, the runner executes sanity-check scripts to ensure data consistency:

  • verify-pipeline.mjs: Validates the integrity of merged records
  • normalize-statuses.mjs: Standardizes status fields across entries
  • dedup-tracker.mjs: Removes duplicate entries that may occur from retried operations

Summary

The career-ops batch runner orchestrates parallel headless evaluations through:

  • Configuration-driven parallelism via the --parallel flag that controls worker pool size
  • Atomic reservation of unique report numbers through reserve-report-num.mjs to prevent identifier collisions
  • Headless worker spawning using claude -p commands that share browser instances when running concurrently
  • File-based locking on batch-state.tsv to serialize state updates and prevent race conditions
  • Recovery record system that writes pending updates to temporary files during lock contention, merged back on subsequent runs
  • Validated aggregation through merge-tracker.mjs and verification scripts that consolidate outputs into the master tracker

Frequently Asked Questions

How does the batch runner prevent duplicate report numbers when running in parallel?

The runner calls reserve-report-num.mjs before spawning any workers, atomically reserving a contiguous block of report numbers. Each worker receives a unique number from this block through the reserve_report_num function (lines 699-704), ensuring no two workers ever claim the same identifier even when launched simultaneously.

What happens if a worker cannot acquire the lock on the state file?

If the lock on batch-state.tsv is unavailable, the worker writes a recovery record to a uniquely named temporary file in the recovery directory (lines 474-488). These records capture the pending state transition and are merged back into the main state file at the beginning of the next run, guaranteeing no data loss occurs during high-contention scenarios.

Why do parallel workers share a single browser instance?

When --parallel is greater than 1, the runner configures workers to share a single browser instance (lines 800-802) to minimize memory consumption. This optimization prevents the resource exhaustion that would occur if each headless worker launched its own separate browser process while evaluating thousands of job postings.

Can the batch process resume if interrupted mid-run?

Yes, the system is fully resumable. The batch-state.tsv file persists the status of every job, and recovery records capture any updates that couldn't be written due to lock contention or crashes. When restarted, the runner reads this state file and skips jobs already marked complete, processing only pending or failed entries from batch-input.tsv.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →