# How the Career-Ops Pipeline Integrity System Works: Inside `merge-tracker.mjs`

> Discover Career-Ops pipeline integrity. Learn how merge-tracker.mjs uses file locking and atomic writes to safely merge batch evaluation results into a canonical markdown tracker.

- Repository: [Santiago Fernández de Valderrama/career-ops](https://github.com/santifer/career-ops)
- Tags: internals
- Published: 2026-07-03

---

**Career-Ops ensures data consistency through a pipeline integrity system that uses file locking, atomic writes, and intelligent deduplication in `merge-tracker.mjs` to safely merge batch evaluation results into a canonical markdown tracker.**

Career-Ops is an open-source job application tracking system that processes evaluation data through a multi-stage pipeline. The **pipeline integrity system** prevents data corruption when concurrent batch workers write to the central tracker, ensuring that [`applications.md`](https://github.com/santifer/career-ops/blob/main/applications.md) remains the single source of truth for all job offers. This architecture guarantees that downstream analytical tools like `followup-cadence.mjs` and `analyze-patterns.mjs` always operate on clean, consistent data.

## What Is the Career-Ops Pipeline Integrity System?

The pipeline consists of three integrated stages: **batch workers** generate TSV files in `batch/tracker-additions/`, the **tracker** ([`applications.md`](https://github.com/santifer/career-ops/blob/main/applications.md) or [`data/applications.md`](https://github.com/santifer/career-ops/blob/main/data/applications.md)) stores the canonical state as a markdown table, and the **merge step** incorporates new data while preventing conflicts and handling duplicates. This design ensures that simultaneous evaluations never overwrite each other, that partial writes cannot corrupt the tracker, and that duplicate offers are resolved intelligently based on evaluation scores.

## How `merge-tracker.mjs` Enforces Pipeline Integrity

The core script (`merge-tracker.mjs`) implements eight specific mechanisms to maintain consistency throughout the pipeline.

### Exclusive Access via UUID-Based Locking

To prevent race conditions between concurrent workers, the script creates a lock directory using either the `CAREER_OPS_TRACKER_LOCK` environment variable or a deterministic temp directory. The `acquireTrackerLock` function (lines 31-84) guards this directory with a UUID token, implementing retry logic with exponential backoff and automatic recovery from stale locks.

### Atomic File Writes

The helper `writeFileAtomic` (lines 96-106) ensures the tracker is never left in a partially written state. It writes to a temporary file in the same directory and uses `renameSync` to swap the files atomically, guaranteeing that other processes always see a complete file or the previous valid state.

### Column-Aware Parsing for Custom Layouts

The script imports `detectColumns` from `tracker-parse.mjs` (lines 21-23) to support custom tracker layouts—such as adding a *Location* column—without breaking existing parsing logic. This modular approach allows the pipeline to evolve while maintaining backward compatibility with existing tracker formats.

### Deduplication and Fuzzy Matching

When merging batch results, the script normalizes company names using `normalizeCompany` and extracts report numbers via `extractReportNum`. For role titles, it employs `roleFuzzyMatch` (lines 90-112) to identify duplicates even when wording varies slightly, preventing duplicate rows for the same offer while accommodating minor description differences.

### Score-Based Updates

The system only upgrades existing entries when new evaluations score higher. The `parseScore` function extracts numerical ratings, and the comparison logic (lines 115-131) replaces the old row only if the new score is greater, preserving the highest-quality evaluation data in the tracker.

### Canonical Status Enforcement

The `validateStatus` function (lines 38-73) maps language-specific or legacy status strings to the canonical set defined in [`templates/states.yml`](https://github.com/santifer/career-ops/blob/main/templates/states.yml) (e.g., `Evaluated`, `Applied`). Non-canonical values are logged for review, ensuring downstream tools receive predictable input regardless of how batch workers label states.

### Migration and Cleanup Support

When invoked with `--migrate`, the script triggers legacy report-link normalization (lines 99-114) to fix paths relative to the tracker location. After successful merges, it moves processed TSV files to `batch/tracker-additions/merged/` (lines 174-177) to prevent reprocessing.

### Optional Pipeline Verification

The `--verify` flag invokes `verify-pipeline.mjs` (lines 186-192) after merging to validate the entire dataset integrity, catching any inconsistencies that might have been introduced during the batch generation phase.

## Practical Usage Examples

### Dry-Run Mode

Preview changes without modifying files:

```bash
node merge-tracker.mjs --dry-run

```

This displays how many entries would be added, updated, or skipped without touching the tracker.

### Manual Lock Acquisition

Reuse the locking logic in other scripts:

```javascript
import { acquireTrackerLock } from './merge-tracker.mjs';

const lock = await acquireTrackerLock('/tmp/career-ops-lock');
try {
  // Critical section operations
} finally {
  lock.release();
}

```

This demonstrates the same UUID-based locking that protects the tracker during standard merges.

### Batch Input Format

Create TSV files in `batch/tracker-additions/` following this structure:

```text
001	2024-06-30	Acme Corp	Senior Engineer	4.3/5	Applied	✅	[001](reports/001-acme-senior-engineer-2024-06-30.md)	First evaluation

```

When `merge-tracker.mjs` runs, it inserts this row after the header separator in [`applications.md`](https://github.com/santifer/career-ops/blob/main/applications.md) (lines 54-66), applying all deduplication and normalization rules.

## Key Files in the Integrity System

Several components support this architecture:

- **`tracker-parse.mjs`**: Shared column-detection logic used by multiple pipeline stages
- **`role-matcher.mjs`**: Fuzzy matching algorithms for duplicate detection
- **`tracker-links.mjs`**: Path normalization for report links
- **[`templates/states.yml`](https://github.com/santifer/career-ops/blob/main/templates/states.yml)**: Canonical status definitions enforced by `validateStatus`
- **`verify-pipeline.mjs`**: Optional post-merge validation suite

## Summary

- **The pipeline integrity system** in Career-Ops combines file locking, atomic writes, and intelligent merging to prevent data corruption during concurrent batch operations.
- **`merge-tracker.mjs`** coordinates exclusive access via UUID-based locks and atomic `renameSync` operations in `writeFileAtomic`.
- **Deduplication** uses fuzzy matching on company names and role titles while upgrading entries based on evaluation scores via `parseScore`.
- **Status normalization** enforces canonical values from [`templates/states.yml`](https://github.com/santifer/career-ops/blob/main/templates/states.yml) for downstream compatibility.
- **Cleanup and verification** steps archive processed files to `batch/tracker-additions/merged/` and optionally validate the entire pipeline.

## Frequently Asked Questions

### What prevents two batch workers from corrupting the tracker simultaneously?

The `acquireTrackerLock` function in `merge-tracker.mjs` creates a dedicated lock directory guarded by a UUID token. It implements retry logic with exponential backoff and can recover from stale locks, ensuring only one process writes to the tracker at any time according to the Career-Ops source code.

### How does the system handle duplicate job offers across different evaluations?

The script normalizes company names using `normalizeCompany` and uses `roleFuzzyMatch` to compare role titles for similarity. When duplicates are detected, it compares scores using `parseScore` and only retains the higher-rated entry, preventing redundant rows while preserving the best evaluation data.

### Can I customize the tracker table columns without breaking the merge process?

Yes. The script imports `detectColumns` from `tracker-parse.mjs` to dynamically identify column positions rather than using hardcoded indices. This allows you to add columns like *Location* without modifying the merge logic, maintaining compatibility with the existing pipeline.

### What happens if the merge process is interrupted mid-write?

Atomic write operations via `writeFileAtomic` ensure the original tracker remains intact. The function writes to a temporary file first, then uses `renameSync` to replace the original only after the write completes successfully, eliminating the risk of partial writes that could corrupt [`applications.md`](https://github.com/santifer/career-ops/blob/main/applications.md).