How CareerOps Performs Archetype Detection on Job Listings: A Technical Deep Dive

CareerOps classifies every job posting into one of six deterministic archetypes (or a hybrid match) by comparing normalized job description text against a curated keyword table in modes/_shared.md.

The CareerOps open-source project (available at santifer/career-ops) implements a transparent, rule-based classification system that runs as the first step of its evaluation pipeline. Unlike black-box AI classifiers, this archetype detection algorithm is fully deterministic, source-controlled, and auditable, using keyword presence scoring to categorize roles and drive downstream CV tailoring.

The Six Archetypes and Signal Keywords

The foundation of the detection system lives in modes/_shared.md, which contains a static table mapping each archetype to its hallmark terminology. This file serves as the single source of truth for what constitutes each role type.

Each archetype definition includes specific key signals—terms that typically appear in matching job descriptions. For example, words like "observability," "agent," or "roadmap" serve as strong indicators for specific archetypes. When the CareerOps engine starts, it parses this table once and converts each keyword list into lower-cased regular expressions for boundary-safe matching.

The Detection Algorithm Step-by-Step

The archetype detection process follows a deterministic five-step pipeline implemented in match-star.mjs.

Loading Archetype Definitions

At startup, the engine reads and parses modes/_shared.md using a helper function. Each archetype is transformed into an object containing the archetype name and its associated keyword regexes.

import { parseArchetypeTable } from './modes/_shared.mjs';

const archetypes = parseArchetypeTable(); 
// Returns: [{name: 'Systems', keywords: [/observability/, /infrastructure/, ...]}, ...]

Tokenizing and Normalizing JD Text

The raw job description—whether fetched via Playwright or the built-in scanner—undergoes preprocessing before matching occurs. The engine strips HTML tags, markdown formatting, and punctuation, then converts the entire text to lowercase to ensure case-insensitive matching against the keyword regexes.

Ranking and Hybrid Detection

The core scoring logic counts keyword presence (not frequency) for each archetype:

function detectArchetype(jdText) {
  const lower = jdText.toLowerCase();
  
  const scores = archetypes.map(a => ({
    name: a.name,
    count: a.keywords.reduce((c, kw) => c + (lower.includes(kw) ? 1 : 0), 0)
  }));
  
  scores.sort((x, y) => y.count - x.count);
  
  const primary = scores[0];
  const secondary = scores[1] && (scores[0].count - scores[1].count <= 1)
    ? scores[1] 
    : null;
    
  return { primary: primary.name, secondary: secondary?.name };
}

If the second-ranked archetype's score falls within the default tolerance of ≤1 point of the top score, CareerOps reports it as a secondary (hybrid) match. This dual-classification allows the system to handle job descriptions that blend responsibilities across archetypes.

Implementation in match-star.mjs

The match-star.mjs module contains the canonical implementation of this logic. It exports the detection function used universally across the codebase. The real implementation uses RegExp objects with word-boundary checks and maintains a small stop-word list to prevent false positives, though the scoring principle remains identical to the simplified version above.

This module returns a structured object containing the primary archetype and optional secondary classification, which populates the archetype field in downstream processing.

Integration with Evaluation Modes

Every evaluation driver in CareerOps invokes the archetype detector before generating reports:

  • openai-eval.mjs – Calls the detector to frame prompts for OpenAI models
  • ollama-eval.mjs – Uses classifications for local LLM evaluation
  • gemini-eval.mjs – Applies archetype context to Google's Gemini prompts

The detected archetype drives two critical downstream features:

  1. North-Star Alignment Scoring – The classification determines how closely the job matches the user's target profile (referenced in web/src/lib/profile-keywords.mjs)
  2. Adaptive Framing – The resulting archetype appears in the report header (e.g., **Archetype:** Systems in modes/oferta.md) and tailors the CV and cover-letter language to match the role's expected terminology

Summary

  • CareerOps performs archetype detection through deterministic keyword matching against definitions stored in modes/_shared.md
  • The algorithm counts keyword presence (not frequency) across six predefined archetypes, with optional hybrid detection when scores differ by ≤1 point
  • Core logic resides in match-star.mjs, invoked by all evaluation modes including openai-eval.mjs, ollama-eval.mjs, and gemini-eval.mjs
  • Results populate the archetype field used for North-Star alignment scoring and adaptive CV framing
  • The system is fully transparent, with no machine learning—just source-controlled keyword matching

Frequently Asked Questions

What are the six archetypes used in CareerOps?

The six archetypes are defined in the modes/_shared.md file within the santifer/career-ops repository. While the specific names aren't enumerated in the core algorithm files, the system categorizes roles based on signal words typical to infrastructure, product management, agent-based systems, observability, and other specialized software engineering domains. You can view the complete table and keyword mappings directly in the repository's shared modes documentation.

How does CareerOps handle job descriptions that fit multiple archetypes?

When the detection algorithm runs in match-star.mjs, it calculates scores for all six archetypes. If the second-highest scoring archetype falls within one point of the top score (default tolerance), CareerOps returns both as primary and secondary matches. This hybrid detection allows the system to handle blended roles—such as a position requiring both systems architecture and product management skills—by acknowledging both archetypes in the generated report and adaptive framing.

Is the archetype detection powered by AI or machine learning?

No, the archetype detection is entirely deterministic and rule-based. It uses regular expression matching against a static keyword table in modes/_shared.md. There are no neural networks or probabilistic models involved in the classification step. The AI components (OpenAI, Ollama, or Gemini) are used only after detection occurs, to generate tailored content based on the pre-determined archetype classification.

Where does CareerOps store the archetype definitions and keywords?

All archetype definitions and their associated keyword signals live in modes/_shared.md at the repository root. This markdown file contains a table that maps each archetype to its indicative terms. The parseArchetypeTable() function in the codebase reads this file at startup, making the classification system fully auditable and modifiable without changing any JavaScript code.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →