# How `jd-skill-gap.mjs` Classifies JD Skills Against a CV Without LLM Calls in `santifer/career-ops`

> Learn how santifer/career-ops' jd-skill-gap.mjs classifies JD skills against a CV using deterministic regexes and token sets without LLM calls. Enhance your career ops analysis.

- Repository: [Santiago Fernández de Valderrama/career-ops](https://github.com/santifer/career-ops)
- Tags: deep-dive
- Published: 2026-08-20

---

**`jd-skill-gap.mjs` uses deterministic regexes, token sets, and word-boundary searches to compare job description skills against a CV—no LLM or external AI service is ever invoked.**

The `santifer/career-ops` repository provides a fully **rule-based pipeline** for identifying skill gaps between a job description (JD) and a candidate's resume. Unlike modern AI-powered resume analyzers, this tool achieves fast, offline, and fully traceable results using only pattern matching and canonicalization tables.

---

## JD Skill Extraction via Regex-Only Parsing

The classification begins with `extractJdSkills()`, which scans the JD line-by-line using purely deterministic pattern matching.

The script identifies the **requirements section** using `REQUIREMENT_HEADER_RE` and stops parsing when `NON_REQUIREMENT_HEADER_RE` matches a subsequent heading. Within this block, `BULLET_LINE_RE` extracts individual requirement lines, and `SKILL_TOKEN_RE` tokenizes each bullet—capturing capital-letter-starting tokens and symbol-edge tokens like `C++` or `C#`. A stop-word filter prevents generic nouns from polluting the skill set.

This regex-only approach is implemented in `scanJd()` at lines 45-92 of `jd-skill-gap.mjs`:

```javascript
// From jd-skill-gap.mjs
const REQUIREMENT_HEADER_RE = /^#+\s*(requirements?|qualifications?|skills?|experience needed)/i;
const NON_REQUIREMENT_HEADER_RE = /^#+\s*(benefits?|compensation|about us|apply)/i;
const BULLET_LINE_RE = /^\s*[-*]\s+(.+)$/;
const SKILL_TOKEN_RE = /\b[A-Z][a-z]*\b|\b\w+#?\b/g;

```

---

## Canonicalization: Unifying Skill Spellings

Before comparison, all skill tokens pass through `skill-extract.mjs` for **canonicalization**. This shared module defines:

- `SKILL_TOKENS`: The master vocabulary of recognized technologies
- `CANONICAL`: A mapping of aliases to standard forms
- `canonicalize(token)`: Lower-cases input, resolves aliases, and returns consistent display names

For example, `k8s` → `Kubernetes`, `golang` → `Go`, and `reactjs` → `React`. This ensures that variant spellings across JD and CV are treated as identical skills without any semantic inference.

The canonicalization logic resides at lines 95-103 and 190-197 of `skill-extract.mjs`:

```javascript
// From skill-extract.mjs
export function canonicalize(token) {
  const lower = token.toLowerCase();
  const canonicalKey = CANONICAL[lower] || lower;
  return DISPLAY[canonicalKey] || canonicalKey;
}

```

---

## CV Segmentation: Skills Section vs. Prose

The tool treats explicit skill listings differently from narrative mentions. `splitSkillsSection(cvText)` (lines 81-99 of `jd-skill-gap.mjs`) locates any markdown heading matching `# Skills` (at any level) and splits the CV into:

| Segment | Purpose |
|---------|---------|
| `namedSkillsText` | Explicitly listed skills under a "Skills" heading |
| `proseText` | Remaining resume content—experience, projects, education |

This separation enables nuanced classification: a skill appearing in the dedicated list carries higher confidence than one mentioned only in passing.

---

## The Classification Algorithm

The core `classifySkillGaps(jdSkills, cvText)` function (lines 124-164 of `jd-skill-gap.mjs`) applies a tiered matching strategy:

1. **Canonical matching against named skills** — JD skill canonicalized, checked against `namedCanon` set → classified as **existing**
2. **Canonical matching against prose** — if not in named skills, checked against `proseCanon` set → classified as **supportedByResume**
3. **Fallback word-boundary search** — `skillMentionedInText()` performs exact substring matches with word boundaries for tokens outside the canonical vocabulary
4. **Remaining skills** → classified as **gap**

```javascript
// Simplified classification logic from jd-skill-gap.mjs
for (const skill of jdSkills) {
  const canon = canonicalize(skill);
  
  if (namedCanon.has(canon)) {
    result.existing.push(canon);
  } else if (proseCanon.has(canon)) {
    result.supportedByResume.push(canon);
  } else if (skillMentionedInText(canon, namedText) || skillMentionedInText(canon, proseText)) {
    result.supportedByResume.push(canon); // fallback hit
  } else {
    result.gap.push(canon);
  }
}

```

This deterministic cascade guarantees reproducible results—the same JD and CV always produce identical classifications.

---

## Diagnostic Validation and Confidence Flags

The pipeline includes `diagnoseExtraction()` (lines 31-48 of `jd-skill-gap.mjs`) to prevent false negatives. It distinguishes three diagnostic states:

- **Conclusive**: At least one JD skill was processed → returns `null`
- **No requirements section**: JD lacked a recognizable requirements block
- **No skill candidates**: Requirements section found but yielded zero tokens

In `--summary` mode, these states emit warnings so users know whether empty gap results indicate genuine alignment or parsing failure.

---

## Running the Tool

Execute the classifier from command line or import its functions:

```bash

# Human-readable summary

node jd-skill-gap.mjs jds/example.md --summary

# JSON output for downstream processing

node jd-skill-gap.mjs jds/example.md --json

```

Programmatic usage:

```javascript
import { extractJdSkills, classifySkillGaps } from './jd-skill-gap.mjs';
import { readFileSync } from 'fs';

const jd = readFileSync('jds/backend-engineer.md', 'utf8');
const cv = readFileSync('cv.md', 'utf8');

const jdSkills = extractJdSkills(jd);
const analysis = classifySkillGaps(jdSkills, cv);

console.log('Explicitly listed:', analysis.existing);
console.log('Mentioned in prose:', analysis.supportedByResume);
console.log('True gaps:', analysis.gap);

```

---

## Summary

- **`jd-skill-gap.mjs`** implements a zero-LLM pipeline for JD-to-CV skill gap analysis in `santifer/career-ops`
- **Regex-based extraction** identifies requirements sections and tokenizes skills without semantic parsing
- **Canonicalization via `skill-extract.mjs`** unifies spelling variants using alias tables
- **CV segmentation** separates explicit skill lists from narrative content for tiered classification
- **Deterministic matching cascade** produces reproducible, fully traceable results
- **Diagnostic validation** prevents silent failures when JD structure is unexpected

---

## Frequently Asked Questions

### How does the tool handle technologies with multiple common spellings?

The canonicalization system in `skill-extract.mjs` maintains explicit alias mappings. Variants like `k8s`, `kubernetes`, and `Kubernetes` all resolve to the canonical form `Kubernetes` through the `CANONICAL` lookup table, ensuring consistent matching regardless of which spelling appears in the JD or CV.

### What happens if a JD lacks a clear requirements section?

`diagnoseExtraction()` detects this condition and returns a diagnostic flag. In `--summary` mode, the tool emits a low-confidence warning explaining that no requirements block was found, distinguishing this parsing failure from a genuine "no gaps" result.

### Why split the CV into named skills and prose sections?

This separation enables more accurate confidence assessment. A skill listed explicitly under a "Skills" heading represents deliberate self-identification, while the same term buried in project descriptions may indicate incidental exposure. The classifier surfaces this distinction in its `existing` versus `supportedByResume` categories.

### Can the tool recognize skills not in its canonical vocabulary?

Yes, through fallback word-boundary matching. If a JD skill token fails canonical lookup, `skillMentionedInText()` performs exact substring searches with word boundaries against both the named skills section and prose. This catches exact matches for emerging technologies not yet added to `SKILL_TOKENS`.