How Batch Processing with Parallel AI Agents Works in Career‑Ops

Career‑Ops evaluates multiple job postings simultaneously by atomically reserving unique report numbers, spawning independent Node.js workers that each run the full auto-pipeline in isolation, and aggregating the results back into the central tracker after all parallel processes complete.

The santifer/career-ops repository implements an embarrassingly parallel architecture for high‑throughput job evaluation. By eliminating shared mutable state and reserving identifiers upfront, the system safely scales batch processing with parallel AI agents across many job descriptions without race conditions or duplicate reports.

Atomic Report Number Reservation

Before any workers start, the system must guarantee that every parallel agent writes to a distinct report ID. The script reserve-report-num.mjs handles this by atomically claiming a contiguous block of report numbers—such as 042-049—and persisting the reservation to a lock file.

This reservation step is the integrity layer of the pipeline. Because each worker receives a pre‑allocated range, simultaneous processes can never claim the same identifier. The lock file ensures that even under heavy concurrency, the allocated block remains exclusive to the current batch job.

Parallel Worker Execution with batch-tailor.mjs

Once the report range is reserved, batch-tailor.mjs distributes the IDs to independent workers. Each process is launched via a simple Node.js command and operates on its own reserved slice of the range.

node batch-tailor.mjs --start 042 --count 8

In this example, the script spawns workers beginning at report 042 for a total of eight parallel agents. Each worker invokes the standard evaluation mode, referred to in the codebase as auto-pipeline, which parses the job description, scores the fit, and generates a PDF. Because every agent writes to its own files—such as reports/042-*.md and output/042-*.pdf—there is no shared mutable state between processes.

Result Aggregation and Tracker Consolidation

After all parallel agents finish, the pipeline enters a single‑threaded aggregation phase. The script merge-tracker.mjs collects the generated TSV additions produced by each worker and integrates them into the canonical data/applications.md tracker. This step performs consistency checks including deduplication and status validation, ensuring the tracker remains the single source of truth.

Token usage across all workers is summed by batch/aggregate-tokens.mjs, which produces a concise usage report. This accounting script runs once the parallel jobs are complete, giving an accurate total of API consumption for the entire batch without needing cross‑process communication during execution.

Summary

  • reserve-report-num.mjs atomically reserves a contiguous range of report numbers to prevent race conditions before workers start.
  • batch-tailor.mjs launches independent Node.js workers via command‑line flags like --start and --count, keeping execution embarrassingly parallel.
  • Each agent runs the full auto-pipeline in isolation, writing to unique file paths such as reports/042-*.md and output/042-*.pdf.
  • merge-tracker.mjs consolidates worker outputs into data/applications.md with deduplication and validation.
  • batch/aggregate-tokens.mjs tallies total token usage after all workers finish, providing a unified cost report.

Frequently Asked Questions

How does Career‑Ops prevent duplicate report numbers during parallel execution?

The system calls reserve-report-num.mjs before spawning any agents. This script atomically claims a contiguous block of report numbers and writes the reservation to a lock file, guaranteeing that each parallel worker receives a unique identifier and eliminating race conditions.

What command launches the parallel AI agents in a batch job?

You start the workers by running node batch-tailor.mjs --start <number> --count <number>. For example, node batch-tailor.mjs --start 042 --count 8 distributes reserved IDs to eight independent processes that each execute the full evaluation pipeline.

How are results merged after all parallel agents finish?

The script merge-tracker.mjs gathers the TSV additions generated by each worker and merges them into data/applications.md. It performs deduplication and status validation to maintain a consistent, canonical tracker.

Where is token usage aggregated across all batch workers?

Token consumption is totaled by batch/aggregate-tokens.mjs after all parallel jobs complete. This produces a single usage report without requiring inter‑process communication during the parallel execution phase.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →