# How to Manage Application Pipeline Tracking with Integrity Checks in Career-Ops

> Learn how to manage application pipeline tracking with integrity checks in Career-Ops. Ensure consistency and deduplication using atomic merge scripts and file-based locking.

- Repository: [Santiago Fernández de Valderrama/career-ops](https://github.com/santifer/career-ops)
- Tags: how-to-guide
- Published: 2026-08-19

---

**Career-Ops manages application pipeline tracking with integrity checks by storing every job application in a single [`data/applications.md`](https://github.com/santifer/career-ops/blob/main/data/applications.md) tracker and enforcing consistency through atomic merge scripts, canonical state normalization, file-based locking, and score-based deduplication.**

The `santifer/career-ops` repository treats your job search as a formal data pipeline. Instead of manually editing spreadsheets, it centralizes application pipeline tracking with integrity checks by routing all updates through validation scripts that prevent duplicates, reject failed batch reports, and normalize every status to a canonical vocabulary.

## How the Tracker Pipeline Works

Every job application lives in one master file: [`data/applications.md`](https://github.com/santifer/career-ops/blob/main/data/applications.md). Batch workers never touch this file directly. Instead, they write new rows as **TSV files** into `batch/tracker-additions/`. A chain of integrity scripts then ingests, validates, and cleans those additions before they become permanent.

### Batch Additions and the Merge Stage

Workers create TSV files such as `batch/tracker-additions/023-Acme-Senior-AI-Engineer.tsv`. Each file contains a single tab-separated row with columns for report number, date, company, role, status, score, flag, report link, and notes.

The `merge-tracker.mjs` script processes these files atomically:

```bash
node merge-tracker.mjs --dry-run   # preview changes

node merge-tracker.mjs             # apply changes

```

During a merge, the script performs three critical checks. It normalizes statuses using `validateStatus`, deduplicates by **company plus role plus report number**, and rejects any row whose report number appears in `batch/batch-state.tsv` with a `failed` status.

### Deduplication by Score and Pipeline Status

After a merge, `dedup-tracker.mjs` removes duplicate rows that describe the same opening:

```bash
node dedup-tracker.mjs --dry-run   # preview removals

node dedup-tracker.mjs             # perform deduplication

```

The script groups rows using `normalizeCompany` and exact-normalized role strings from `normalizeRole`. When two rows refer to the same report number, the script keeps the entry with the higher score. If scores are equal, it uses `statusRank` to preserve the more advanced pipeline status:

```js
function statusRank(status) {
  const map = { skip: 0, discarded: 0, rejected: 1, evaluated: 2,
                applied: 3, responded: 4, interview: 5, offer: 6, hired: 7 };
  return map[normalizeStatus(status)] || 0;
}

```

## Normalizing Canonical Pipeline States

Both scripts enforce a single source of truth for status labels. The `validateStatus` function in `merge-tracker.mjs` and the `normalizeStatus` function in `dedup-tracker.mjs` map every spelling variant to the **canonical states** defined in [`templates/states.yml`](https://github.com/santifer/career-ops/blob/main/templates/states.yml).

The `validateStatus` helper strips markdown decorations and trailing dates, then resolves aliases:

```js
function validateStatus(status) {
  const clean = status.replace(/\*\*/g, '').replace(/\s+\d{4}-\d{2}-\d{2}.*$/, '').trim();
  const lower = clean.toLowerCase();
  const aliases = { 'aplicado': 'Applied', 'entrevista': 'Interview' };
  return CANONICAL_STATES.find(s => s.toLowerCase() === lower) ||
         aliases[lower] || 'Evaluated';
}

```

Any unrecognized status falls back to `Evaluated` and triggers a warning, preventing malformed labels from polluting the tracker.

## Enforcing Concurrency Safety

Concurrent writes are blocked by a **file-based lock** managed in `tracker-utils.mjs`. Before any script modifies [`data/applications.md`](https://github.com/santifer/career-ops/blob/main/data/applications.md), it acquires the lock in the `tracker-lock/` directory. This guarantees that only one process can mutate the tracker at a time, eliminating race conditions between batch workers and manual CLI operations.

After a successful merge, the source TSV is moved to `batch/tracker-additions/merged/` so it cannot be processed again. For targeted single-row updates, use `set-status.mjs`, which applies the same lock and canonical-state logic.

## Running the Complete Integrity Workflow

You can add a new application programmatically and run the full pipeline from the shell:

```js
const fs = require('fs');
const tsv = [
  '024',
  '2026-08-20',
  'Beta Labs',
  'Machine Learning Engineer',
  'Applied',
  '4.8/5',
  '✅',
  '[024](reports/024-beta-ml-engineer-2026-08-20.md)',
  'Sent application via LinkedIn'
].join('\t');
fs.writeFileSync('batch/tracker-additions/024-Beta-ML-Engineer.tsv', tsv);

require('child_process').execSync('node merge-tracker.mjs');

```

Then run the standard maintenance sequence:

```bash
node merge-tracker.mjs
node dedup-tracker.mjs
node verify-pipeline.mjs

```

The `verify-pipeline.mjs` script performs a final sanity check for unknown states, missing columns, and formatting drift after the merge and dedup steps complete.

## Summary

- **Store updates as TSV additions** in `batch/tracker-additions/` instead of editing [`data/applications.md`](https://github.com/santifer/career-ops/blob/main/data/applications.md) by hand.
- **Run `merge-tracker.mjs`** to atomically ingest new rows, normalize statuses, and exclude failed reports listed in `batch/batch-state.tsv`.
- **Run `dedup-tracker.mjs`** to collapse duplicate openings, keeping the highest score or most advanced pipeline rank.
- **Respect the file lock** in `tracker-lock/`; the scripts acquire it automatically, so never edit the tracker during an active merge.
- **Define custom states** in [`templates/states.yml`](https://github.com/santifer/career-ops/blob/main/templates/states.yml); otherwise `validateStatus` falls back to `Evaluated`.
- **Schedule `verify-pipeline.mjs`** periodically to catch stray formatting issues before they accumulate.

## Frequently Asked Questions

### What is the master tracker file in Career-Ops?

The master tracker is [`data/applications.md`](https://github.com/santifer/career-ops/blob/main/data/applications.md). It contains a single markdown table that holds every tracked job application. All writes flow through `merge-tracker.mjs` or `set-status.mjs` to preserve table integrity.

### How does Career-Ops prevent duplicate job applications?

The system prevents duplicates in two stages. First, `merge-tracker.mjs` deduplicates by company, role, and report number during ingestion. Second, `dedup-tracker.mjs` groups rows with `normalizeCompany` and `normalizeRole`, then uses `statusRank` to keep the most favorable entry when collisions remain.

### What happens if a batch worker reports a failed status?

If a report number appears in `batch/batch-state.tsv` with status `failed`, `merge-tracker.mjs` rejects the corresponding TSV row entirely. This safety check prevents fabricated or erroneous batch results from entering the tracker.

### Can I edit [`applications.md`](https://github.com/santifer/career-ops/blob/main/applications.md) directly?

No. Direct manual edits bypass the integrity checks, status normalization, and file-based locking system. Always add TSV files to `batch/tracker-additions/` and run `merge-tracker.mjs`, or use `set-status.mjs` for targeted updates.