# How the Career-Ops Pipeline Integrity System Handles Merging and Deduplication

> Learn how Career Ops pipeline integrity uses merge-tracker and dedup-tracker scripts to atomically merge batch TSV additions and remove fuzzy duplicates, preventing race conditions.

- Repository: [Santiago Fernández de Valderrama/career-ops](https://github.com/santifer/career-ops)
- Tags: internals
- Published: 2026-06-13

---

**The Career-Ops pipeline integrity system uses two specialized scripts—`merge-tracker.mjs` and `dedup-tracker.mjs`—to atomically merge batch TSV additions into the tracker while preventing race conditions, and to remove fuzzy duplicates while preserving advanced application statuses.**

The santifer/career-ops repository maintains a job-application tracker ([`applications.md`](https://github.com/santifer/career-ops/blob/main/applications.md)) that requires strict consistency guarantees when ingesting batch evaluations. The pipeline integrity system ensures concurrent updates never corrupt data through filesystem locking and atomic writes, while intelligent deduplication algorithms preserve the highest-scoring entries and most advanced workflow statuses.

## How New Entries Are Merged into the Tracker

The `merge-tracker.mjs` script functions as the entry point for all new job evaluations. It processes TSV files generated by batch workers and integrates them into the central tracker with crash-safe guarantees.

### Locating the Tracker and Acquiring Exclusive Locks

Before reading any data, the script canonicalizes the tracker path to avoid lock collisions. According to the source in `merge-tracker.mjs` lines 29-35, the logic checks for [`data/applications.md`](https://github.com/santifer/career-ops/blob/main/data/applications.md) first, falling back to [`applications.md`](https://github.com/santifer/career-ops/blob/main/applications.md) in the root directory.

The `acquireTrackerLock` function (lines 42-66) creates an atomic filesystem lock using `mkdirSync`. The lock directory stores the process PID and a random token. If another process holds the lock, the script implements retry logic with exponential back-off and can recover stale locks from crashed processes.

### Parsing and Normalizing TSV Input

The script reads all `*.tsv` files from `batch/tracker-additions`, sorting them numerically before processing. The `parseTsvContent` function (lines 96-108) handles both 8-column and 9-column variants, including pipe-delimited rows within fields.

Report links undergo normalization through `normalizeReportLink` (lines 62-66 in `tracker-links.mjs`), converting root-relative paths like `reports/…` to be relative to the tracker file's directory. This ensures links remain valid regardless of where the tracker is accessed from.

### Duplicate Detection and Conflict Resolution

The merging algorithm detects existing entries through three matching strategies defined in lines 83-110:

- **Exact report number match** using `extractReportNum`
- **Same tracker number** combined with normalized company name
- **Same normalized company** with fuzzy role matching via `roleFuzzyMatch` from `role-matcher.mjs`

When duplicates are found, the script parses scores using `parseScore` (lines 119-124). If the new entry carries a higher score, the existing row updates in-place; otherwise, the addition is skipped entirely. This guarantees the tracker always contains the best available evaluation for any given opportunity.

### Atomic Write and Cleanup

Non-duplicate rows append after the header separator using the next available tracker number (`maxNum`). The `writeFileAtomic` helper (lines 97-105) writes to a temporary file before calling `renameSync`, ensuring the update is atomic and crash-resistant.

After successful merging, processed TSV files move to `batch/tracker-additions/merged/` and the filesystem lock releases (lines 173-188).

## How Deduplication Cleans the Tracker

While merging prevents exact duplicates, fuzzy matches—similar roles at the same company—require periodic cleanup via `dedup-tracker.mjs`.

### Grouping by Normalized Company Names

The script loads the tracker using the same path resolution logic as the merge step (lines 21-26). The `normalizeCompany` function (lines 57-64) strips punctuation, parentheses, and extra whitespace to create stable grouping keys. A `Map` stores all entries for each normalized company (lines 73-79).

### Fuzzy Role Matching with Status Protection

Within each company group, `roleFuzzyMatch` compares role titles. However, the algorithm protects entries with "advanced" statuses—defined in [`templates/states.yml`](https://github.com/santifer/career-ops/blob/main/templates/states.yml) as Applied, Interview, Offer, or later stages. When either entry in a comparison pair holds an advanced status, the duplicate pair is preserved regardless of fuzzy similarity (lines 86-102).

### Selecting Keepers and Promoting Statuses

For duplicate clusters that pass the status filter, rows sort by numeric score via `parseScore`. The highest-scoring row becomes the keeper (lines 104-108). Before removing duplicates, the script checks if any discarded row holds a more advanced status than the keeper; if so, that status promotes onto the surviving row (lines 121-128).

Removal happens in reverse index order to preserve array integrity (lines 144-148). Unless running in dry-run mode, the script creates a backup (`applications.md.bak`) and writes the cleaned tracker atomically (lines 152-157).

## Running the Pipeline Scripts

Execute merges safely using dry-run mode to preview changes without modifying [`applications.md`](https://github.com/santifer/career-ops/blob/main/applications.md):

```bash
node career-ops/merge-tracker.mjs --dry-run

```

Verify data integrity after merging:

```bash
node career-ops/merge-tracker.mjs --verify

```

Preview deduplication without committing changes:

```bash
node career-ops/dedup-tracker.mjs --dry-run

```

Perform actual deduplication with backup creation:

```bash
node career-ops/dedup-tracker.mjs

```

## Summary

- **Atomic merging** via `merge-tracker.mjs` uses filesystem locks and temporary file renaming to prevent data corruption during concurrent writes.
- **Intelligent duplicate detection** matches entries by report number, tracker ID, or fuzzy role similarity, always retaining the highest-scoring version.
- **Safe deduplication** through `dedup-tracker.mjs` groups entries by normalized company names and protects advanced application statuses from removal.
- **Status promotion** ensures that even when removing duplicate rows, the most advanced workflow state transfers to the surviving entry.
- **Crash-resistant writes** utilize `writeFileAtomic` patterns with `renameSync` to guarantee tracker consistency even if the process terminates unexpectedly.

## Frequently Asked Questions

### What prevents race conditions when multiple batch workers finish simultaneously?

The `acquireTrackerLock` function in `merge-tracker.mjs` creates an exclusive lock using atomic directory creation (`mkdirSync`). The lock stores the process PID and a random token, with stale-lock detection allowing recovery if a previous process crashed. This ensures only one merge operation writes to [`applications.md`](https://github.com/santifer/career-ops/blob/main/applications.md) at any time.

### How does the system decide which duplicate entry to keep?

During merging, the script compares scores using `parseScore` and retains the higher-scoring entry. During deduplication, entries within the same company group sort by score, keeping the highest. If any removed duplicate holds an advanced status (Applied, Interview, Offer), that status promotes to the keeper, ensuring no workflow progress is lost.

### Can I preview changes before modifying the tracker?

Both scripts support `--dry-run` flags. Running `merge-tracker.mjs --dry-run` displays how many rows would be added, updated, or skipped. Similarly, `dedup-tracker.mjs --dry-run` shows which duplicate rows would be removed and which statuses would be promoted, without writing changes to disk.

### What file formats does the merge script accept?

The `parseTsvContent` function accepts TSV files with either 8 or 9 columns, tolerating pipe-delimited content within fields. The script processes all `*.tsv` files in `batch/tracker-additions`, handling root-relative report links by normalizing them to be relative to the tracker file's location using `tracker-links.mjs`.