How Career-Ops Batch Processing Mode Utilizes Headless Workers for Parallel Job Evaluations
Career-Ops batch processing mode orchestrates parallel job evaluations by spawning headless CLAUDE-Code workers through a Bash-based job-control loop, using file-based locking to prevent race conditions when writing to shared state files.
The batch processing mode in the santifer/career-ops repository enables automated, high-throughput evaluation of job postings without UI interaction. By leveraging headless workers and a sophisticated state management system, the batch/batch-runner.sh orchestrator can process dozens of offers concurrently while maintaining data integrity. This architecture treats each job offer as a discrete unit of work, distributing evaluations across parallel processes that safely share report numbers and state through atomic file operations.
Loading the Job Queue and State Management
The batch processing mode begins by reading the job queue from batch/batch-input.tsv, a tab-separated file containing columns for id, url, source, and notes.
Simultaneously, the runner initializes batch/batch-state.tsv to track every offer’s lifecycle. This state file records critical metadata including status, timestamps, report numbers, scores, and retry counts. The runner maintains a file-based lock mechanism using .batch-state.lock to ensure that concurrent headless workers never corrupt the shared state during parallel executions.
Report Number Reservation with Locking
Before spawning workers, the system must allocate a unique report number for each offer. The reserve-report-num.mjs script handles this sequentially while holding the state lock.
This reservation prevents race conditions where parallel workers might claim the same identifier. The lock file guarantees that even when running with --parallel N where N > 1, report numbers remain deterministic and non-overlapping across the entire batch.
Worker Spawning and Prompt Resolution
The process_offer function (lines 406-487 in batch/batch-runner.sh) constructs the headless worker command that evaluates each job posting.
Each worker receives a resolved system prompt stored as .resolved-prompt-${id}.md. The runner generates this file by substituting placeholders in batch/batch-prompt.md:
{{URL}}– The job posting URL{{JD_FILE}}– Path to the job description file{{REPORT_NUM}}– The reserved report number{{DATE}}– Current timestamp{{ID}}– The offer identifier
The runner appends personalization files (modes/_profile.md and config/profile.yml) to ensure batch evaluations mirror the interactive pipeline’s behavior. The actual command uses claude -p to launch a headless CLAUDE-Code instance for each offer.
Parallel Execution Control
The batch processing mode supports both sequential and parallel execution through the --parallel flag.
When --parallel 1 is specified, the script calls process_offer directly in the main process. For parallel execution with --parallel N where N > 1, the runner implements a Bash job-control loop (lines 730-815) that:
- Forks background jobs up to the specified limit
- Tracks running PIDs in an array
- Waits for slot availability before spawning new workers
- Maintains the state lock during report number reservation and status updates
This lightweight concurrency model avoids complex threading while maximizing throughput for I/O-bound evaluation tasks.
Rate Limiting and Session Management
The batch processing mode includes sophisticated handling for API rate limits and session restrictions. If a worker’s output contains rate-limit keywords, the runner writes a paused_rate_limit entry to the state file and creates batch-runner.paused to halt new worker scheduling.
You can resume a paused batch using the --resume-paused flag, which clears the paused state and continues processing remaining offers. The --rate-limit-sleep parameter configures backoff delays between retries, while --max-retries controls resilience for transient failures.
Result Aggregation and Pipeline Reconciliation
After workers complete, each headless process writes a tracker line to batch/tracker-additions/ as a TSV file. The merge-tracker.mjs script consolidates these individual entries into data/applications.md, maintaining the central job application tracker.
Finally, reconcile-pipeline.mjs moves processed offers from the active pipeline to the "Processed" section, while verify-pipeline.mjs runs sanity checks to ensure consistency between the pipeline and tracker files. This three-phase approach (evaluation → aggregation → verification) ensures the UI-facing pipeline.md remains synchronized with batch operations.
Practical Usage Examples
Create the input file with your target job postings:
cat > batch/batch-input.tsv <<'EOF'
id url source notes
1 https://example.com/job/123 greenhouse Initial batch run
2 https://another.com/role lever High-priority
EOF
Execute the batch with three parallel workers, retry logic, and rate-limit protection:
./batch/batch-runner.sh \
--parallel 3 \
--retry-failed \
--rate-limit-sleep 60 \
--max-retries 4
Check the current status without processing new offers:
./batch/batch-runner.sh --status
Resume a run that was paused due to rate limiting:
./batch/batch-runner.sh --resume-paused
Summary
- batch/batch-runner.sh serves as the central orchestrator, managing the full lifecycle from queue ingestion to result aggregation.
- Headless workers are spawned via
claude -pcommands within theprocess_offerfunction, each receiving a resolved prompt with substituted variables. - Parallel execution is controlled through a Bash job-control loop that respects the
--parallellimit while tracking background PIDs. - File-based locking via
.batch-state.lockandreserve-report-num.mjsguarantees atomic report number allocation and prevents write conflicts inbatch-state.tsv. - State persistence in
batch/batch-state.tsvenables pause/resume functionality and retry logic for failed evaluations. - Results consolidation flows through
merge-tracker.mjsintodata/applications.md, followed by pipeline reconciliation and verification scripts.
Frequently Asked Questions
How does the batch processing mode prevent race conditions when multiple headless workers write to the same state file?
The system uses a file-based lock mechanism centered on .batch-state.lock. When reserving report numbers or updating batch/batch-state.tsv, the reserve-report-num.mjs script and the runner itself hold this lock exclusively. This ensures that even with --parallel set to values greater than 1, only one worker at a time can modify shared state, preventing overlapping writes or duplicate report numbers.
What happens if a headless worker encounters a rate limit during parallel job evaluation?
If a worker detects rate-limiting keywords in its output, the runner immediately writes a paused_rate_limit status to batch/batch-state.tsv and creates a batch-runner.paused file. The orchestrator stops scheduling new offers but preserves the state of in-progress workers. You can resume the batch processing mode later using ./batch/batch-runner.sh --resume-paused, which clears the paused state and continues with remaining offers.
Can I customize the prompt that headless workers use for job evaluations?
Yes. The batch processing mode uses batch/batch-prompt.md as a template containing placeholders like {{URL}}, {{JD_FILE}}, and {{REPORT_NUM}}. The runner generates per-offer resolved prompts (.resolved-prompt-${id}.md) by substituting these variables. Additionally, the system automatically appends modes/_profile.md and config/profile.yml to each worker’s prompt, ensuring batch evaluations match your personalized interactive configuration.
How does the system handle failed offers during parallel execution?
The batch processing mode supports retry logic through the --retry-failed flag and --max-retries parameter. Failed offers are recorded in batch/batch-state.tsv with their retry count incremented. When running with --retry-failed, the process_offer function re-evaluates these offers up to the maximum retry limit, allowing transient failures to recover without manual intervention while maintaining the parallel execution workflow.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →