# How `detect-reposts.mjs` Flags Roles Re-Listed 2+ Times in 90 Days in santifer/career-ops

> Discover how detect-reposts.mjs flags roles reposted over twice in 90 days within santifer/career-ops. Learn the script's logic for identifying duplicate job listings.

- Repository: [Santiago Fernández de Valderrama/career-ops](https://github.com/santifer/career-ops)
- Tags: how-to-guide
- Published: 2026-08-20

---

**The `detect-reposts.mjs` script identifies reposted job roles by clustering scan history entries where the same company and matching title appear with different URLs within a 90-day sliding window, requiring at least two distinct URLs to trigger a flag.**

The `detect-reposts.mjs` script in the [santifer/career-ops](https://github.com/santifer/career-ops) repository provides a self-contained pipeline for detecting when employers re-list the same role multiple times. This detection helps users avoid stale postings and track employer posting patterns. The implementation processes the `data/scan-history.tsv` file through normalization, fuzzy matching, and sliding-window clustering to surface repost clusters.

## Loading and Filtering Scan History

The script begins by reading the append-only scan log at `data/scan-history.tsv` via `parseScanHistory()` in `detect-reposts.mjs` (lines 70-106). Each TSV line converts to a structured row object with fields including `date`, `company`, `title`, `url`, and `status`.

Only rows with `status === 'added'` are retained. Rows marked as `skipped_expired` or other statuses are discarded early (lines 12-14). This filtering ensures that detection operates on actual new postings rather than metadata about skipped entries.

## Normalizing Company Identity

Consistent company grouping requires handling name variations like "Acme Inc." versus "Acme". The `companyKey()` function (lines 15-19) implements a two-tier strategy:

- **Primary**: Use the pre-normalized company column if present in the TSV
- **Fallback**: Call `normalizeCompanyName()` imported from `invite-match.mjs`

This normalization ensures that minor naming differences do not fragment repost detection across artificial company boundaries.

## Grouping by Title with Fuzzy Matching

Within each company bucket, `groupRowsByTitle()` (lines 74-146) performs efficient title clustering through a two-phase approach:

1. **Exact matching**: Case-insensitive title comparison buckets rows in O(N) time
2. **Fuzzy matching**: For non-exact titles, a token-based inverted index pre-filters candidates before invoking `roleFuzzyMatch()` from `role-matcher.mjs`

The `roleFuzzyMatch()` function implements sophisticated token-set comparison to identify titles like "Senior Software Engineer" and "Sr. Software Engineer - Backend" as representing the same role. The inverted index dramatically reduces expensive fuzzy comparisons by only testing titles that share at least one significant token.

## Sliding-Window Clustering Algorithm

The core repost detection logic uses a date-sorted sliding window. For each title group sorted by `date` (lines 86-100):

- Initialize an empty `cluster` with the earliest row
- Append subsequent rows while `currentDate - firstDate <= windowDays` (default 90)
- When the span exceeds the window, seal the current cluster and start fresh

This algorithm guarantees that every emitted cluster respects the temporal bound. The window is configurable via the `--window` CLI flag.

## Deduplication and Validation Criteria

Before finalizing a cluster, `buildRepostCluster()` (lines 69-85) applies two critical filters:

- **URL deduplication**: Rows sharing identical URLs collapse to their earliest sighting, preventing the same re-scrape from counting multiple times
- **Minimum distinct URLs**: The cluster must contain ≥ 2 unique URLs
- **Window enforcement**: `lastSeen - firstSeen` must not exceed `windowDays`

A repost flag triggers only when all conditions satisfy: same normalized company, fuzzy-matched title, multiple distinct URLs, and 90-day containment.

## Output Formats and CLI Usage

The script emits results sorted by most-recent `lastSeen` (lines 71-72). Two output modes are available (lines 16-18, 400-430):

```bash

# Full JSON with complete cluster details

node detect-reposts.mjs

# Human-readable table view

node detect-reposts.mjs --summary

# Custom window length (days)

node detect-reposts.mjs --window 60 --summary

# Run built-in test suite

node detect-reposts.mjs --self-test

```

## Programmatic Integration

Core functions are importable for custom workflows:

```javascript
import { parseScanHistory, detectReposts } from './detect-reposts.mjs';
import { readFile } from 'fs/promises';

const tsv = await readFile('data/scan-history.tsv', 'utf-8');
const rows = parseScanHistory(tsv);
const clusters = detectReposts(rows, 90); // 90-day window

console.log(`Found ${clusters.length} repost clusters`);

```

## Key Implementation Files

| File | Responsibility |
|------|--------------|
| `detect-reposts.mjs` | Main orchestration: parsing, grouping, window clustering, output |
| `role-matcher.mjs` | `roleFuzzyMatch()` and tokenization for title similarity |
| `invite-match.mjs` | `normalizeCompanyName()` for consistent company keys |
| `data/scan-history.tsv` | Source data: append-only log of all scanned postings |
| `lib/cli-flags.mjs` | CLI argument parsing (`--window`, `--summary`, `--self-test`) |

## Summary

- **Repost detection** requires matching company (normalized), fuzzy title match, ≥ 2 distinct URLs, and 90-day window containment
- **Performance optimization** uses exact bucketing first, then token-indexed fuzzy matching to minimize expensive comparisons
- **URL deduplication** prevents re-scans of identical postings from inflating cluster counts
- **Sliding-window algorithm** guarantees temporal bounds are respected across arbitrarily long posting histories
- **Configurable window** supports custom timeframes via `--window` flag without code changes

## Frequently Asked Questions

### What defines a "repost" in this system?

A repost is a job posting where the same normalized company and matching role title appears with at least two different URLs within the configured time window. The URL requirement distinguishes genuine re-listings from simple re-scrapes of identical postings.

### How does the script handle title variations like "Senior" vs "Sr."?

The `roleFuzzyMatch()` function in `role-matcher.mjs` tokenizes titles and performs set-based similarity comparison. This catches common abbreviations, ordering differences, and minor wording changes without requiring manual synonym lists.

### Can I adjust the 90-day window for different detection sensitivity?

Yes. Pass `--window <days>` to override the default. Shorter windows catch aggressive reposters; longer windows reveal seasonal re-listing patterns. The window applies to the span between first and last sighting in each cluster.

### Why does the script ignore rows with `skipped_expired` status?

These rows represent postings that were filtered out during the scan (expired, already seen, or otherwise excluded). Including them would contaminate repost detection with entries that never entered the active pipeline.