How CareerOps Performs Archetype Detection on Job Listings: A Technical Deep Dive
CareerOps classifies every job posting into one of six deterministic archetypes (or a hybrid match) by comparing normalized job description text against a curated keyword table in modes/_shared.md.
The CareerOps open-source project (available at santifer/career-ops) implements a transparent, rule-based classification system that runs as the first step of its evaluation pipeline. Unlike black-box AI classifiers, this archetype detection algorithm is fully deterministic, source-controlled, and auditable, using keyword presence scoring to categorize roles and drive downstream CV tailoring.
The Six Archetypes and Signal Keywords
The foundation of the detection system lives in modes/_shared.md, which contains a static table mapping each archetype to its hallmark terminology. This file serves as the single source of truth for what constitutes each role type.
Each archetype definition includes specific key signals—terms that typically appear in matching job descriptions. For example, words like "observability," "agent," or "roadmap" serve as strong indicators for specific archetypes. When the CareerOps engine starts, it parses this table once and converts each keyword list into lower-cased regular expressions for boundary-safe matching.
The Detection Algorithm Step-by-Step
The archetype detection process follows a deterministic five-step pipeline implemented in match-star.mjs.
Loading Archetype Definitions
At startup, the engine reads and parses modes/_shared.md using a helper function. Each archetype is transformed into an object containing the archetype name and its associated keyword regexes.
import { parseArchetypeTable } from './modes/_shared.mjs';
const archetypes = parseArchetypeTable();
// Returns: [{name: 'Systems', keywords: [/observability/, /infrastructure/, ...]}, ...]
Tokenizing and Normalizing JD Text
The raw job description—whether fetched via Playwright or the built-in scanner—undergoes preprocessing before matching occurs. The engine strips HTML tags, markdown formatting, and punctuation, then converts the entire text to lowercase to ensure case-insensitive matching against the keyword regexes.
Ranking and Hybrid Detection
The core scoring logic counts keyword presence (not frequency) for each archetype:
function detectArchetype(jdText) {
const lower = jdText.toLowerCase();
const scores = archetypes.map(a => ({
name: a.name,
count: a.keywords.reduce((c, kw) => c + (lower.includes(kw) ? 1 : 0), 0)
}));
scores.sort((x, y) => y.count - x.count);
const primary = scores[0];
const secondary = scores[1] && (scores[0].count - scores[1].count <= 1)
? scores[1]
: null;
return { primary: primary.name, secondary: secondary?.name };
}
If the second-ranked archetype's score falls within the default tolerance of ≤1 point of the top score, CareerOps reports it as a secondary (hybrid) match. This dual-classification allows the system to handle job descriptions that blend responsibilities across archetypes.
Implementation in match-star.mjs
The match-star.mjs module contains the canonical implementation of this logic. It exports the detection function used universally across the codebase. The real implementation uses RegExp objects with word-boundary checks and maintains a small stop-word list to prevent false positives, though the scoring principle remains identical to the simplified version above.
This module returns a structured object containing the primary archetype and optional secondary classification, which populates the archetype field in downstream processing.
Integration with Evaluation Modes
Every evaluation driver in CareerOps invokes the archetype detector before generating reports:
openai-eval.mjs– Calls the detector to frame prompts for OpenAI modelsollama-eval.mjs– Uses classifications for local LLM evaluationgemini-eval.mjs– Applies archetype context to Google's Gemini prompts
The detected archetype drives two critical downstream features:
- North-Star Alignment Scoring – The classification determines how closely the job matches the user's target profile (referenced in
web/src/lib/profile-keywords.mjs) - Adaptive Framing – The resulting archetype appears in the report header (e.g.,
**Archetype:** Systemsinmodes/oferta.md) and tailors the CV and cover-letter language to match the role's expected terminology
Summary
- CareerOps performs archetype detection through deterministic keyword matching against definitions stored in
modes/_shared.md - The algorithm counts keyword presence (not frequency) across six predefined archetypes, with optional hybrid detection when scores differ by ≤1 point
- Core logic resides in
match-star.mjs, invoked by all evaluation modes includingopenai-eval.mjs,ollama-eval.mjs, andgemini-eval.mjs - Results populate the
archetypefield used for North-Star alignment scoring and adaptive CV framing - The system is fully transparent, with no machine learning—just source-controlled keyword matching
Frequently Asked Questions
What are the six archetypes used in CareerOps?
The six archetypes are defined in the modes/_shared.md file within the santifer/career-ops repository. While the specific names aren't enumerated in the core algorithm files, the system categorizes roles based on signal words typical to infrastructure, product management, agent-based systems, observability, and other specialized software engineering domains. You can view the complete table and keyword mappings directly in the repository's shared modes documentation.
How does CareerOps handle job descriptions that fit multiple archetypes?
When the detection algorithm runs in match-star.mjs, it calculates scores for all six archetypes. If the second-highest scoring archetype falls within one point of the top score (default tolerance), CareerOps returns both as primary and secondary matches. This hybrid detection allows the system to handle blended roles—such as a position requiring both systems architecture and product management skills—by acknowledging both archetypes in the generated report and adaptive framing.
Is the archetype detection powered by AI or machine learning?
No, the archetype detection is entirely deterministic and rule-based. It uses regular expression matching against a static keyword table in modes/_shared.md. There are no neural networks or probabilistic models involved in the classification step. The AI components (OpenAI, Ollama, or Gemini) are used only after detection occurs, to generate tailored content based on the pre-determined archetype classification.
Where does CareerOps store the archetype definitions and keywords?
All archetype definitions and their associated keyword signals live in modes/_shared.md at the repository root. This markdown file contains a table that maps each archetype to its indicative terms. The parseArchetypeTable() function in the codebase reads this file at startup, making the classification system fully auditable and modifiable without changing any JavaScript code.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →