How the Career-Ops Pipeline Integrity System Handles Merging and Deduplication

The Career-Ops pipeline integrity system uses two specialized scripts—merge-tracker.mjs and dedup-tracker.mjs—to atomically merge batch TSV additions into the tracker while preventing race conditions, and to remove fuzzy duplicates while preserving advanced application statuses.

The santifer/career-ops repository maintains a job-application tracker (applications.md) that requires strict consistency guarantees when ingesting batch evaluations. The pipeline integrity system ensures concurrent updates never corrupt data through filesystem locking and atomic writes, while intelligent deduplication algorithms preserve the highest-scoring entries and most advanced workflow statuses.

How New Entries Are Merged into the Tracker

The merge-tracker.mjs script functions as the entry point for all new job evaluations. It processes TSV files generated by batch workers and integrates them into the central tracker with crash-safe guarantees.

Locating the Tracker and Acquiring Exclusive Locks

Before reading any data, the script canonicalizes the tracker path to avoid lock collisions. According to the source in merge-tracker.mjs lines 29-35, the logic checks for data/applications.md first, falling back to applications.md in the root directory.

The acquireTrackerLock function (lines 42-66) creates an atomic filesystem lock using mkdirSync. The lock directory stores the process PID and a random token. If another process holds the lock, the script implements retry logic with exponential back-off and can recover stale locks from crashed processes.

Parsing and Normalizing TSV Input

The script reads all *.tsv files from batch/tracker-additions, sorting them numerically before processing. The parseTsvContent function (lines 96-108) handles both 8-column and 9-column variants, including pipe-delimited rows within fields.

Report links undergo normalization through normalizeReportLink (lines 62-66 in tracker-links.mjs), converting root-relative paths like reports/… to be relative to the tracker file's directory. This ensures links remain valid regardless of where the tracker is accessed from.

Duplicate Detection and Conflict Resolution

The merging algorithm detects existing entries through three matching strategies defined in lines 83-110:

  • Exact report number match using extractReportNum
  • Same tracker number combined with normalized company name
  • Same normalized company with fuzzy role matching via roleFuzzyMatch from role-matcher.mjs

When duplicates are found, the script parses scores using parseScore (lines 119-124). If the new entry carries a higher score, the existing row updates in-place; otherwise, the addition is skipped entirely. This guarantees the tracker always contains the best available evaluation for any given opportunity.

Atomic Write and Cleanup

Non-duplicate rows append after the header separator using the next available tracker number (maxNum). The writeFileAtomic helper (lines 97-105) writes to a temporary file before calling renameSync, ensuring the update is atomic and crash-resistant.

After successful merging, processed TSV files move to batch/tracker-additions/merged/ and the filesystem lock releases (lines 173-188).

How Deduplication Cleans the Tracker

While merging prevents exact duplicates, fuzzy matches—similar roles at the same company—require periodic cleanup via dedup-tracker.mjs.

Grouping by Normalized Company Names

The script loads the tracker using the same path resolution logic as the merge step (lines 21-26). The normalizeCompany function (lines 57-64) strips punctuation, parentheses, and extra whitespace to create stable grouping keys. A Map stores all entries for each normalized company (lines 73-79).

Fuzzy Role Matching with Status Protection

Within each company group, roleFuzzyMatch compares role titles. However, the algorithm protects entries with "advanced" statuses—defined in templates/states.yml as Applied, Interview, Offer, or later stages. When either entry in a comparison pair holds an advanced status, the duplicate pair is preserved regardless of fuzzy similarity (lines 86-102).

Selecting Keepers and Promoting Statuses

For duplicate clusters that pass the status filter, rows sort by numeric score via parseScore. The highest-scoring row becomes the keeper (lines 104-108). Before removing duplicates, the script checks if any discarded row holds a more advanced status than the keeper; if so, that status promotes onto the surviving row (lines 121-128).

Removal happens in reverse index order to preserve array integrity (lines 144-148). Unless running in dry-run mode, the script creates a backup (applications.md.bak) and writes the cleaned tracker atomically (lines 152-157).

Running the Pipeline Scripts

Execute merges safely using dry-run mode to preview changes without modifying applications.md:

node career-ops/merge-tracker.mjs --dry-run

Verify data integrity after merging:

node career-ops/merge-tracker.mjs --verify

Preview deduplication without committing changes:

node career-ops/dedup-tracker.mjs --dry-run

Perform actual deduplication with backup creation:

node career-ops/dedup-tracker.mjs

Summary

  • Atomic merging via merge-tracker.mjs uses filesystem locks and temporary file renaming to prevent data corruption during concurrent writes.
  • Intelligent duplicate detection matches entries by report number, tracker ID, or fuzzy role similarity, always retaining the highest-scoring version.
  • Safe deduplication through dedup-tracker.mjs groups entries by normalized company names and protects advanced application statuses from removal.
  • Status promotion ensures that even when removing duplicate rows, the most advanced workflow state transfers to the surviving entry.
  • Crash-resistant writes utilize writeFileAtomic patterns with renameSync to guarantee tracker consistency even if the process terminates unexpectedly.

Frequently Asked Questions

What prevents race conditions when multiple batch workers finish simultaneously?

The acquireTrackerLock function in merge-tracker.mjs creates an exclusive lock using atomic directory creation (mkdirSync). The lock stores the process PID and a random token, with stale-lock detection allowing recovery if a previous process crashed. This ensures only one merge operation writes to applications.md at any time.

How does the system decide which duplicate entry to keep?

During merging, the script compares scores using parseScore and retains the higher-scoring entry. During deduplication, entries within the same company group sort by score, keeping the highest. If any removed duplicate holds an advanced status (Applied, Interview, Offer), that status promotes to the keeper, ensuring no workflow progress is lost.

Can I preview changes before modifying the tracker?

Both scripts support --dry-run flags. Running merge-tracker.mjs --dry-run displays how many rows would be added, updated, or skipped. Similarly, dedup-tracker.mjs --dry-run shows which duplicate rows would be removed and which statuses would be promoted, without writing changes to disk.

What file formats does the merge script accept?

The parseTsvContent function accepts TSV files with either 8 or 9 columns, tolerating pipe-delimited content within fields. The script processes all *.tsv files in batch/tracker-additions, handling root-relative report links by normalizing them to be relative to the tracker file's location using tracker-links.mjs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →