How `jd-skill-gap.mjs` Classifies JD Skills Against a CV Without LLM Calls in `santifer/career-ops`

jd-skill-gap.mjs uses deterministic regexes, token sets, and word-boundary searches to compare job description skills against a CV—no LLM or external AI service is ever invoked.

The santifer/career-ops repository provides a fully rule-based pipeline for identifying skill gaps between a job description (JD) and a candidate's resume. Unlike modern AI-powered resume analyzers, this tool achieves fast, offline, and fully traceable results using only pattern matching and canonicalization tables.


JD Skill Extraction via Regex-Only Parsing

The classification begins with extractJdSkills(), which scans the JD line-by-line using purely deterministic pattern matching.

The script identifies the requirements section using REQUIREMENT_HEADER_RE and stops parsing when NON_REQUIREMENT_HEADER_RE matches a subsequent heading. Within this block, BULLET_LINE_RE extracts individual requirement lines, and SKILL_TOKEN_RE tokenizes each bullet—capturing capital-letter-starting tokens and symbol-edge tokens like C++ or C#. A stop-word filter prevents generic nouns from polluting the skill set.

This regex-only approach is implemented in scanJd() at lines 45-92 of jd-skill-gap.mjs:

// From jd-skill-gap.mjs
const REQUIREMENT_HEADER_RE = /^#+\s*(requirements?|qualifications?|skills?|experience needed)/i;
const NON_REQUIREMENT_HEADER_RE = /^#+\s*(benefits?|compensation|about us|apply)/i;
const BULLET_LINE_RE = /^\s*[-*]\s+(.+)$/;
const SKILL_TOKEN_RE = /\b[A-Z][a-z]*\b|\b\w+#?\b/g;

Canonicalization: Unifying Skill Spellings

Before comparison, all skill tokens pass through skill-extract.mjs for canonicalization. This shared module defines:

  • SKILL_TOKENS: The master vocabulary of recognized technologies
  • CANONICAL: A mapping of aliases to standard forms
  • canonicalize(token): Lower-cases input, resolves aliases, and returns consistent display names

For example, k8s → Kubernetes, golang → Go, and reactjs → React. This ensures that variant spellings across JD and CV are treated as identical skills without any semantic inference.

The canonicalization logic resides at lines 95-103 and 190-197 of skill-extract.mjs:

// From skill-extract.mjs
export function canonicalize(token) {
  const lower = token.toLowerCase();
  const canonicalKey = CANONICAL[lower] || lower;
  return DISPLAY[canonicalKey] || canonicalKey;
}

CV Segmentation: Skills Section vs. Prose

The tool treats explicit skill listings differently from narrative mentions. splitSkillsSection(cvText) (lines 81-99 of jd-skill-gap.mjs) locates any markdown heading matching # Skills (at any level) and splits the CV into:

Segment Purpose
namedSkillsText Explicitly listed skills under a "Skills" heading
proseText Remaining resume content—experience, projects, education

This separation enables nuanced classification: a skill appearing in the dedicated list carries higher confidence than one mentioned only in passing.


The Classification Algorithm

The core classifySkillGaps(jdSkills, cvText) function (lines 124-164 of jd-skill-gap.mjs) applies a tiered matching strategy:

  1. Canonical matching against named skills — JD skill canonicalized, checked against namedCanon set → classified as existing
  2. Canonical matching against prose — if not in named skills, checked against proseCanon set → classified as supportedByResume
  3. Fallback word-boundary search — skillMentionedInText() performs exact substring matches with word boundaries for tokens outside the canonical vocabulary
  4. Remaining skills → classified as gap
// Simplified classification logic from jd-skill-gap.mjs
for (const skill of jdSkills) {
  const canon = canonicalize(skill);
  
  if (namedCanon.has(canon)) {
    result.existing.push(canon);
  } else if (proseCanon.has(canon)) {
    result.supportedByResume.push(canon);
  } else if (skillMentionedInText(canon, namedText) || skillMentionedInText(canon, proseText)) {
    result.supportedByResume.push(canon); // fallback hit
  } else {
    result.gap.push(canon);
  }
}

This deterministic cascade guarantees reproducible results—the same JD and CV always produce identical classifications.


Diagnostic Validation and Confidence Flags

The pipeline includes diagnoseExtraction() (lines 31-48 of jd-skill-gap.mjs) to prevent false negatives. It distinguishes three diagnostic states:

  • Conclusive: At least one JD skill was processed → returns null
  • No requirements section: JD lacked a recognizable requirements block
  • No skill candidates: Requirements section found but yielded zero tokens

In --summary mode, these states emit warnings so users know whether empty gap results indicate genuine alignment or parsing failure.


Running the Tool

Execute the classifier from command line or import its functions:


# Human-readable summary

node jd-skill-gap.mjs jds/example.md --summary

# JSON output for downstream processing

node jd-skill-gap.mjs jds/example.md --json

Programmatic usage:

import { extractJdSkills, classifySkillGaps } from './jd-skill-gap.mjs';
import { readFileSync } from 'fs';

const jd = readFileSync('jds/backend-engineer.md', 'utf8');
const cv = readFileSync('cv.md', 'utf8');

const jdSkills = extractJdSkills(jd);
const analysis = classifySkillGaps(jdSkills, cv);

console.log('Explicitly listed:', analysis.existing);
console.log('Mentioned in prose:', analysis.supportedByResume);
console.log('True gaps:', analysis.gap);

Summary

  • jd-skill-gap.mjs implements a zero-LLM pipeline for JD-to-CV skill gap analysis in santifer/career-ops
  • Regex-based extraction identifies requirements sections and tokenizes skills without semantic parsing
  • Canonicalization via skill-extract.mjs unifies spelling variants using alias tables
  • CV segmentation separates explicit skill lists from narrative content for tiered classification
  • Deterministic matching cascade produces reproducible, fully traceable results
  • Diagnostic validation prevents silent failures when JD structure is unexpected

Frequently Asked Questions

How does the tool handle technologies with multiple common spellings?

The canonicalization system in skill-extract.mjs maintains explicit alias mappings. Variants like k8s, kubernetes, and Kubernetes all resolve to the canonical form Kubernetes through the CANONICAL lookup table, ensuring consistent matching regardless of which spelling appears in the JD or CV.

What happens if a JD lacks a clear requirements section?

diagnoseExtraction() detects this condition and returns a diagnostic flag. In --summary mode, the tool emits a low-confidence warning explaining that no requirements block was found, distinguishing this parsing failure from a genuine "no gaps" result.

Why split the CV into named skills and prose sections?

This separation enables more accurate confidence assessment. A skill listed explicitly under a "Skills" heading represents deliberate self-identification, while the same term buried in project descriptions may indicate incidental exposure. The classifier surfaces this distinction in its existing versus supportedByResume categories.

Can the tool recognize skills not in its canonical vocabulary?

Yes, through fallback word-boundary matching. If a JD skill token fails canonical lookup, skillMentionedInText() performs exact substring searches with word boundaries against both the named skills section and prose. This catches exact matches for emerging technologies not yet added to SKILL_TOKENS.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →