How the Career-Ops Pipeline Integrity System Works: Inside `merge-tracker.mjs`
Career-Ops ensures data consistency through a pipeline integrity system that uses file locking, atomic writes, and intelligent deduplication in merge-tracker.mjs to safely merge batch evaluation results into a canonical markdown tracker.
Career-Ops is an open-source job application tracking system that processes evaluation data through a multi-stage pipeline. The pipeline integrity system prevents data corruption when concurrent batch workers write to the central tracker, ensuring that applications.md remains the single source of truth for all job offers. This architecture guarantees that downstream analytical tools like followup-cadence.mjs and analyze-patterns.mjs always operate on clean, consistent data.
What Is the Career-Ops Pipeline Integrity System?
The pipeline consists of three integrated stages: batch workers generate TSV files in batch/tracker-additions/, the tracker (applications.md or data/applications.md) stores the canonical state as a markdown table, and the merge step incorporates new data while preventing conflicts and handling duplicates. This design ensures that simultaneous evaluations never overwrite each other, that partial writes cannot corrupt the tracker, and that duplicate offers are resolved intelligently based on evaluation scores.
How merge-tracker.mjs Enforces Pipeline Integrity
The core script (merge-tracker.mjs) implements eight specific mechanisms to maintain consistency throughout the pipeline.
Exclusive Access via UUID-Based Locking
To prevent race conditions between concurrent workers, the script creates a lock directory using either the CAREER_OPS_TRACKER_LOCK environment variable or a deterministic temp directory. The acquireTrackerLock function (lines 31-84) guards this directory with a UUID token, implementing retry logic with exponential backoff and automatic recovery from stale locks.
Atomic File Writes
The helper writeFileAtomic (lines 96-106) ensures the tracker is never left in a partially written state. It writes to a temporary file in the same directory and uses renameSync to swap the files atomically, guaranteeing that other processes always see a complete file or the previous valid state.
Column-Aware Parsing for Custom Layouts
The script imports detectColumns from tracker-parse.mjs (lines 21-23) to support custom tracker layouts—such as adding a Location column—without breaking existing parsing logic. This modular approach allows the pipeline to evolve while maintaining backward compatibility with existing tracker formats.
Deduplication and Fuzzy Matching
When merging batch results, the script normalizes company names using normalizeCompany and extracts report numbers via extractReportNum. For role titles, it employs roleFuzzyMatch (lines 90-112) to identify duplicates even when wording varies slightly, preventing duplicate rows for the same offer while accommodating minor description differences.
Score-Based Updates
The system only upgrades existing entries when new evaluations score higher. The parseScore function extracts numerical ratings, and the comparison logic (lines 115-131) replaces the old row only if the new score is greater, preserving the highest-quality evaluation data in the tracker.
Canonical Status Enforcement
The validateStatus function (lines 38-73) maps language-specific or legacy status strings to the canonical set defined in templates/states.yml (e.g., Evaluated, Applied). Non-canonical values are logged for review, ensuring downstream tools receive predictable input regardless of how batch workers label states.
Migration and Cleanup Support
When invoked with --migrate, the script triggers legacy report-link normalization (lines 99-114) to fix paths relative to the tracker location. After successful merges, it moves processed TSV files to batch/tracker-additions/merged/ (lines 174-177) to prevent reprocessing.
Optional Pipeline Verification
The --verify flag invokes verify-pipeline.mjs (lines 186-192) after merging to validate the entire dataset integrity, catching any inconsistencies that might have been introduced during the batch generation phase.
Practical Usage Examples
Dry-Run Mode
Preview changes without modifying files:
node merge-tracker.mjs --dry-run
This displays how many entries would be added, updated, or skipped without touching the tracker.
Manual Lock Acquisition
Reuse the locking logic in other scripts:
import { acquireTrackerLock } from './merge-tracker.mjs';
const lock = await acquireTrackerLock('/tmp/career-ops-lock');
try {
// Critical section operations
} finally {
lock.release();
}
This demonstrates the same UUID-based locking that protects the tracker during standard merges.
Batch Input Format
Create TSV files in batch/tracker-additions/ following this structure:
001 2024-06-30 Acme Corp Senior Engineer 4.3/5 Applied ✅ [001](reports/001-acme-senior-engineer-2024-06-30.md) First evaluation
When merge-tracker.mjs runs, it inserts this row after the header separator in applications.md (lines 54-66), applying all deduplication and normalization rules.
Key Files in the Integrity System
Several components support this architecture:
tracker-parse.mjs: Shared column-detection logic used by multiple pipeline stagesrole-matcher.mjs: Fuzzy matching algorithms for duplicate detectiontracker-links.mjs: Path normalization for report linkstemplates/states.yml: Canonical status definitions enforced byvalidateStatusverify-pipeline.mjs: Optional post-merge validation suite
Summary
- The pipeline integrity system in Career-Ops combines file locking, atomic writes, and intelligent merging to prevent data corruption during concurrent batch operations.
merge-tracker.mjscoordinates exclusive access via UUID-based locks and atomicrenameSyncoperations inwriteFileAtomic.- Deduplication uses fuzzy matching on company names and role titles while upgrading entries based on evaluation scores via
parseScore. - Status normalization enforces canonical values from
templates/states.ymlfor downstream compatibility. - Cleanup and verification steps archive processed files to
batch/tracker-additions/merged/and optionally validate the entire pipeline.
Frequently Asked Questions
What prevents two batch workers from corrupting the tracker simultaneously?
The acquireTrackerLock function in merge-tracker.mjs creates a dedicated lock directory guarded by a UUID token. It implements retry logic with exponential backoff and can recover from stale locks, ensuring only one process writes to the tracker at any time according to the Career-Ops source code.
How does the system handle duplicate job offers across different evaluations?
The script normalizes company names using normalizeCompany and uses roleFuzzyMatch to compare role titles for similarity. When duplicates are detected, it compares scores using parseScore and only retains the higher-rated entry, preventing redundant rows while preserving the best evaluation data.
Can I customize the tracker table columns without breaking the merge process?
Yes. The script imports detectColumns from tracker-parse.mjs to dynamically identify column positions rather than using hardcoded indices. This allows you to add columns like Location without modifying the merge logic, maintaining compatibility with the existing pipeline.
What happens if the merge process is interrupted mid-write?
Atomic write operations via writeFileAtomic ensure the original tracker remains intact. The function writes to a temporary file first, then uses renameSync to replace the original only after the write completes successfully, eliminating the risk of partial writes that could corrupt applications.md.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →