How `detect-reposts.mjs` Flags Roles Re-Listed 2+ Times in 90 Days in santifer/career-ops

The detect-reposts.mjs script identifies reposted job roles by clustering scan history entries where the same company and matching title appear with different URLs within a 90-day sliding window, requiring at least two distinct URLs to trigger a flag.

The detect-reposts.mjs script in the santifer/career-ops repository provides a self-contained pipeline for detecting when employers re-list the same role multiple times. This detection helps users avoid stale postings and track employer posting patterns. The implementation processes the data/scan-history.tsv file through normalization, fuzzy matching, and sliding-window clustering to surface repost clusters.

Loading and Filtering Scan History

The script begins by reading the append-only scan log at data/scan-history.tsv via parseScanHistory() in detect-reposts.mjs (lines 70-106). Each TSV line converts to a structured row object with fields including date, company, title, url, and status.

Only rows with status === 'added' are retained. Rows marked as skipped_expired or other statuses are discarded early (lines 12-14). This filtering ensures that detection operates on actual new postings rather than metadata about skipped entries.

Normalizing Company Identity

Consistent company grouping requires handling name variations like "Acme Inc." versus "Acme". The companyKey() function (lines 15-19) implements a two-tier strategy:

  • Primary: Use the pre-normalized company column if present in the TSV
  • Fallback: Call normalizeCompanyName() imported from invite-match.mjs

This normalization ensures that minor naming differences do not fragment repost detection across artificial company boundaries.

Grouping by Title with Fuzzy Matching

Within each company bucket, groupRowsByTitle() (lines 74-146) performs efficient title clustering through a two-phase approach:

  1. Exact matching: Case-insensitive title comparison buckets rows in O(N) time
  2. Fuzzy matching: For non-exact titles, a token-based inverted index pre-filters candidates before invoking roleFuzzyMatch() from role-matcher.mjs

The roleFuzzyMatch() function implements sophisticated token-set comparison to identify titles like "Senior Software Engineer" and "Sr. Software Engineer - Backend" as representing the same role. The inverted index dramatically reduces expensive fuzzy comparisons by only testing titles that share at least one significant token.

Sliding-Window Clustering Algorithm

The core repost detection logic uses a date-sorted sliding window. For each title group sorted by date (lines 86-100):

  • Initialize an empty cluster with the earliest row
  • Append subsequent rows while currentDate - firstDate <= windowDays (default 90)
  • When the span exceeds the window, seal the current cluster and start fresh

This algorithm guarantees that every emitted cluster respects the temporal bound. The window is configurable via the --window CLI flag.

Deduplication and Validation Criteria

Before finalizing a cluster, buildRepostCluster() (lines 69-85) applies two critical filters:

  • URL deduplication: Rows sharing identical URLs collapse to their earliest sighting, preventing the same re-scrape from counting multiple times
  • Minimum distinct URLs: The cluster must contain ≥ 2 unique URLs
  • Window enforcement: lastSeen - firstSeen must not exceed windowDays

A repost flag triggers only when all conditions satisfy: same normalized company, fuzzy-matched title, multiple distinct URLs, and 90-day containment.

Output Formats and CLI Usage

The script emits results sorted by most-recent lastSeen (lines 71-72). Two output modes are available (lines 16-18, 400-430):


# Full JSON with complete cluster details

node detect-reposts.mjs

# Human-readable table view

node detect-reposts.mjs --summary

# Custom window length (days)

node detect-reposts.mjs --window 60 --summary

# Run built-in test suite

node detect-reposts.mjs --self-test

Programmatic Integration

Core functions are importable for custom workflows:

import { parseScanHistory, detectReposts } from './detect-reposts.mjs';
import { readFile } from 'fs/promises';

const tsv = await readFile('data/scan-history.tsv', 'utf-8');
const rows = parseScanHistory(tsv);
const clusters = detectReposts(rows, 90); // 90-day window

console.log(`Found ${clusters.length} repost clusters`);

Key Implementation Files

File Responsibility
detect-reposts.mjs Main orchestration: parsing, grouping, window clustering, output
role-matcher.mjs roleFuzzyMatch() and tokenization for title similarity
invite-match.mjs normalizeCompanyName() for consistent company keys
data/scan-history.tsv Source data: append-only log of all scanned postings
lib/cli-flags.mjs CLI argument parsing (--window, --summary, --self-test)

Summary

  • Repost detection requires matching company (normalized), fuzzy title match, ≥ 2 distinct URLs, and 90-day window containment
  • Performance optimization uses exact bucketing first, then token-indexed fuzzy matching to minimize expensive comparisons
  • URL deduplication prevents re-scans of identical postings from inflating cluster counts
  • Sliding-window algorithm guarantees temporal bounds are respected across arbitrarily long posting histories
  • Configurable window supports custom timeframes via --window flag without code changes

Frequently Asked Questions

What defines a "repost" in this system?

A repost is a job posting where the same normalized company and matching role title appears with at least two different URLs within the configured time window. The URL requirement distinguishes genuine re-listings from simple re-scrapes of identical postings.

How does the script handle title variations like "Senior" vs "Sr."?

The roleFuzzyMatch() function in role-matcher.mjs tokenizes titles and performs set-based similarity comparison. This catches common abbreviations, ordering differences, and minor wording changes without requiring manual synonym lists.

Can I adjust the 90-day window for different detection sensitivity?

Yes. Pass --window <days> to override the default. Shorter windows catch aggressive reposters; longer windows reveal seasonal re-listing patterns. The window applies to the span between first and last sighting in each cluster.

Why does the script ignore rows with skipped_expired status?

These rows represent postings that were filtered out during the scan (expired, already seen, or otherwise excluded). Including them would contaminate repost detection with entries that never entered the active pipeline.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →