How `jd-capture.mjs` Resolves Archived JDs by Report Number with Prefix Matching in santifer/career-ops
The jd-capture.mjs module locates archived job descriptions in the jds/ folder by matching the leading numeric prefix of filenames against a target report number, with optional company slug disambiguation and recency-based tie-breaking.
In the Career-Ops repository, the jd-capture.mjs module provides a robust lookup mechanism for retrieving captured job descriptions (JDs) that belong to specific tracker entries. The system supports both legacy and modern filename formats, ensuring backward compatibility while preventing false matches. This article explains the algorithm implemented in santifer/career-ops for resolving archived JDs by report number.
Normalizing Report Numbers for Consistent Matching
The resolution process begins with reportPrefix(), a helper function that standardizes numeric IDs. Located at lines 21-24 in jd-capture.mjs, this function pads report numbers to three digits:
console.log(reportPrefix(7)); // → "007"
console.log(reportPrefix(64)); // → "064"
This normalization ensures consistency with the reports/ folder naming convention, though the matching algorithm itself handles both padded and unpadded formats in archived filenames.
Scanning the Capture Directory with Prefix Matching
The core lookup logic resides in findCaptureForReport(). This function first reads the entire jds/ directory, returning null immediately if the folder does not exist.
The prefix matching implementation at lines 77-82 uses a regex to extract leading digits:
const matches = entries.filter(name => {
const m = /^(\d+)-/.exec(name);
return m ? parseInt(m[1], 10) === target : false;
});
Key design properties of this matcher:
- The
^(\d+)-pattern requires a trailing hyphen, preventing64from matching640-company.txt - Numeric comparison via
parseInt()allows064-acme.txtand64-acme.txtto both match report64 - Non-matching files are excluded silently
Ordering Candidates by Recency
When multiple captures exist for the same report number, the system applies deterministic selection at lines 85-94:
const sorted = matches
.map(name => {
const st = statSync(join(dirPath, name));
return { name, mtimeMs: st.mtimeMs };
})
.sort((a, b) => {
if (b.mtimeMs !== a.mtimeMs) return b.mtimeMs - a.mtimeMs;
return a.name.localeCompare(b.name);
});
The sort prioritizes modification timestamp (newest first) and falls back to alphabetical order for deterministic tie-breaking.
Optional Company Slug Disambiguation
Callers may supply a companySlug option to filter among multiple captures. The companyMatches() helper (lines 26-50) implements strict slug validation:
- Removes the numeric prefix and optional date from the filename
- Verifies the remaining string starts with the supplied slug
- Requires a delimiter (
-,_,.) or end-of-string after the slug
const captureForAcme = findCaptureForReport(resolve('jds'), 70, {
companySlug: 'iota-old'
});
If a slug is provided but no candidate matches, the function returns null rather than risk attaching the wrong JD. The selection step at lines 95-108 implements this guard logic.
Complete API and Return Value
The function returns a structured result object at lines 111-116 containing:
{
path, // Absolute file path
filename, // File name with extension
ext, // File extension
candidates // Full ordered list of matching files
}
When no match exists, the function returns null.
Practical Usage Examples
import { findCaptureForReport, reportPrefix } from './jd-capture.mjs';
import { resolve } from 'path';
// Simple lookup by report number
const capture = findCaptureForReport(resolve('jds'), 64);
if (capture) {
console.log('Found JD:', capture.filename); // → "064-acme.txt"
}
// With company disambiguation
const specific = findCaptureForReport(resolve('jds'), 70, {
companySlug: 'iota-old'
});
// Generate standardized prefix
console.log(reportPrefix(7)); // → "007"
These patterns are exercised by outcome.mjs (lines 34-35) when recording application outcomes.
Summary
reportPrefix()creates three-digit normalized report identifiersfindCaptureForReport()implements regex-based prefix matching withparseInt()normalization for cross-format compatibility- Recency sorting by
mtimeMswith alphabetical tie-breaking provides deterministic selection companyMatches()enforces strict slug boundaries to prevent false matches- The API returns full metadata including the complete candidate list for debugging
- Null-safe handling throughout prevents crashes on missing directories or unmatched queries
Frequently Asked Questions
How does the prefix matcher avoid confusing report 64 with 640?
The regex ^(\d+)- requires a hyphen immediately after the leading digits. Since 640-company.txt has no boundary between 64 and 0, the entire prefix 640 is captured and parsed as the integer 640, which does not equal 64. Only files with a clear delimiter after the target number will match.
Does the system prefer padded or unpadded filenames?
The matching logic treats both formats equally through parseInt() normalization. However, the recency-based sort typically favors newer captures, which in practice use the modern three-digit padded format from reportPrefix().
What happens if two JDs exist for the same company and report number?
The modification timestamp (mtimeMs) determines priority. If timestamps are identical, alphabetical filename order serves as the tie-breaker. The full candidates array in the return value allows callers to implement alternative selection logic if needed.
Can findCaptureForReport() be used without a company slug?
Yes. The companySlug parameter is optional. When omitted, the function returns the most recent capture matching the report number regardless of company name, which is useful for single-company reports or quick lookups.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →