How `jd-capture.mjs` Resolves Archived JDs by Report Number with Prefix Matching in santifer/career-ops

The jd-capture.mjs module locates archived job descriptions in the jds/ folder by matching the leading numeric prefix of filenames against a target report number, with optional company slug disambiguation and recency-based tie-breaking.

In the Career-Ops repository, the jd-capture.mjs module provides a robust lookup mechanism for retrieving captured job descriptions (JDs) that belong to specific tracker entries. The system supports both legacy and modern filename formats, ensuring backward compatibility while preventing false matches. This article explains the algorithm implemented in santifer/career-ops for resolving archived JDs by report number.

Normalizing Report Numbers for Consistent Matching

The resolution process begins with reportPrefix(), a helper function that standardizes numeric IDs. Located at lines 21-24 in jd-capture.mjs, this function pads report numbers to three digits:

console.log(reportPrefix(7));   // → "007"
console.log(reportPrefix(64));  // → "064"

This normalization ensures consistency with the reports/ folder naming convention, though the matching algorithm itself handles both padded and unpadded formats in archived filenames.

Scanning the Capture Directory with Prefix Matching

The core lookup logic resides in findCaptureForReport(). This function first reads the entire jds/ directory, returning null immediately if the folder does not exist.

The prefix matching implementation at lines 77-82 uses a regex to extract leading digits:

const matches = entries.filter(name => {
  const m = /^(\d+)-/.exec(name);
  return m ? parseInt(m[1], 10) === target : false;
});

Key design properties of this matcher:

  • The ^(\d+)- pattern requires a trailing hyphen, preventing 64 from matching 640-company.txt
  • Numeric comparison via parseInt() allows 064-acme.txt and 64-acme.txt to both match report 64
  • Non-matching files are excluded silently

Ordering Candidates by Recency

When multiple captures exist for the same report number, the system applies deterministic selection at lines 85-94:

const sorted = matches
  .map(name => {
    const st = statSync(join(dirPath, name));
    return { name, mtimeMs: st.mtimeMs };
  })
  .sort((a, b) => {
    if (b.mtimeMs !== a.mtimeMs) return b.mtimeMs - a.mtimeMs;
    return a.name.localeCompare(b.name);
  });

The sort prioritizes modification timestamp (newest first) and falls back to alphabetical order for deterministic tie-breaking.

Optional Company Slug Disambiguation

Callers may supply a companySlug option to filter among multiple captures. The companyMatches() helper (lines 26-50) implements strict slug validation:

  1. Removes the numeric prefix and optional date from the filename
  2. Verifies the remaining string starts with the supplied slug
  3. Requires a delimiter (-, _, .) or end-of-string after the slug
const captureForAcme = findCaptureForReport(resolve('jds'), 70, {
  companySlug: 'iota-old'
});

If a slug is provided but no candidate matches, the function returns null rather than risk attaching the wrong JD. The selection step at lines 95-108 implements this guard logic.

Complete API and Return Value

The function returns a structured result object at lines 111-116 containing:

{
  path,        // Absolute file path
  filename,    // File name with extension
  ext,         // File extension
  candidates   // Full ordered list of matching files
}

When no match exists, the function returns null.

Practical Usage Examples

import { findCaptureForReport, reportPrefix } from './jd-capture.mjs';
import { resolve } from 'path';

// Simple lookup by report number
const capture = findCaptureForReport(resolve('jds'), 64);
if (capture) {
  console.log('Found JD:', capture.filename);   // → "064-acme.txt"
}

// With company disambiguation
const specific = findCaptureForReport(resolve('jds'), 70, {
  companySlug: 'iota-old'
});

// Generate standardized prefix
console.log(reportPrefix(7));   // → "007"

These patterns are exercised by outcome.mjs (lines 34-35) when recording application outcomes.

Summary

  • reportPrefix() creates three-digit normalized report identifiers
  • findCaptureForReport() implements regex-based prefix matching with parseInt() normalization for cross-format compatibility
  • Recency sorting by mtimeMs with alphabetical tie-breaking provides deterministic selection
  • companyMatches() enforces strict slug boundaries to prevent false matches
  • The API returns full metadata including the complete candidate list for debugging
  • Null-safe handling throughout prevents crashes on missing directories or unmatched queries

Frequently Asked Questions

How does the prefix matcher avoid confusing report 64 with 640?

The regex ^(\d+)- requires a hyphen immediately after the leading digits. Since 640-company.txt has no boundary between 64 and 0, the entire prefix 640 is captured and parsed as the integer 640, which does not equal 64. Only files with a clear delimiter after the target number will match.

Does the system prefer padded or unpadded filenames?

The matching logic treats both formats equally through parseInt() normalization. However, the recency-based sort typically favors newer captures, which in practice use the modern three-digit padded format from reportPrefix().

What happens if two JDs exist for the same company and report number?

The modification timestamp (mtimeMs) determines priority. If timestamps are identical, alphabetical filename order serves as the tie-breaker. The full candidates array in the return value allows callers to implement alternative selection logic if needed.

Can findCaptureForReport() be used without a company slug?

Yes. The companySlug parameter is optional. When omitted, the function returns the most recent capture matching the report number regardless of company name, which is useful for single-company reports or quick lookups.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →