What Data Powers the A-H Evaluation Blocks in CareerOps: Complete Data Flow Guide
The A-H Evaluation blocks in CareerOps consume job description text, Playwright browser snapshots, candidate CVs, YAML configuration files, and auxiliary reference templates to generate comprehensive role analysis, match scoring, and risk assessment reports.
The open-source CareerOps toolchain (available at santifer/career-ops) processes job opportunities through a structured A-H evaluation pipeline. Each block transforms specific input data sources—from live URL snapshots to local profile configurations—into structured intelligence. Understanding these data dependencies is critical for customizing evaluations, debugging outputs, or extending the framework.
The A-G+H Pipeline Structure
CareerOps implements an eight-stage evaluation workflow defined in modes/oferta.md. The pipeline runs sequentially from Block A through Block G, with Block H serving as a final aggregation layer that collates risk signals generated throughout the process. Each stage maintains strict data contracts, consuming only the specific artifacts required for its analysis domain.
The architecture separates concerns into distinct functional units: role classification (A), candidate matching (B), career strategy (C), market analysis (D), document optimization (E), interview preparation (F), legitimacy verification (G), and risk summarization (H).
Data Sources by Evaluation Block
Block A – Role Summary
Block A ingests the raw job posting and determines the foundational characteristics of the opportunity. The primary data sources include:
- JD text or live URL snapshot: Either direct text input or a Playwright-captured snapshot of the posting page
- Archetype detection logic: Scoring algorithms defined in
modes/_shared.mdthat classify roles into one of six archetypes - Candidate profile:
config/profile.ymlcontaining work authorization status, authorized countries, and spend tier preferences - Optional blacklist:
data/blacklist.mdfor case-insensitive company name matching that can abort the pipeline before processing continues
This block outputs a structured table including archetype, domain, function, seniority, remote setup, team size, culture-screen outcome, and a one-sentence TL;DR summary.
Block B – CV Match
Block B performs bi-directional requirement mapping between the job posting and candidate background:
- JD requirements: Structured requirements extracted during Block A processing
- Candidate CV:
cv.mdcontaining the candidate's experience and skills
The block maps every JD requirement to exact line(s) in the CV that satisfy it, flags capability gaps, and suggests mitigation strategies for missing qualifications.
Block C – Level & Strategy
This block determines position alignment with career trajectory using:
- Inferred level: Seniority classification derived from JD analysis in Block A
- Candidate seniority: Archetype-specific seniority data from
config/profile.yml
Block C determines if the role represents a level-up or level-down move, proposes a "sell senior without lying" narrative framework, and generates fallback negotiation tactics.
Block D – Compensation & Demand
Compensation analysis relies on bounded external research:
- Salary data: Explicit ranges present in the JD
- Market research: Maximum of five web-search queries (enforced by the "Bounded Research Budget" constraint)
- Company classification: Public corporation, startup, or other entity type classification
The block generates advertised-range analysis, reliability tier classification, and breakdowns of guaranteed base, variable components, and non-cash benefits.
Block E – Customisation Plan
Document optimization requires:
- CV source:
cv.mdfor gap analysis - LinkedIn profile: Optional LinkedIn data for social presence alignment
Block E outputs the top-five prioritized CV edits and LinkedIn tweaks necessary to improve match scores against the specific JD.
Block F – Interview Plan
Interview preparation consumes:
- JD requirements: From Block A output
- Story bank:
interview-prep/story-bank.mdcontaining STAR+R formatted experience narratives
The block aligns 6-10 stories to specific JD items, adding reflection columns that demonstrate seniority-level thinking.
Block G – Posting Legitimacy
The most data-intensive block, G performs multi-signal fraud and quality detection using:
- Playwright snapshot: HTML capture from the liveness gate validation
- JD text: Processed description content
- Scan history:
data/scan-history.tsvfor reposting detection and temporal analysis - Jurisdiction tables:
templates/agency-licensing.yml,templates/immigration-status-requirements.yml, andtemplates/jurisdiction-prohibited-content.ymlfor compliance verification
Block G analyzes freshness signals, description quality metrics, hiring pattern anomalies, employment-classification risks, AI-buzzword mismatches, benefits-terminology inconsistencies, platform-location tag mismatches, agency licensing validity, immigration-status overreach, prohibited content, pay-range width anomalies, minimum-wage compliance, and AI-screening disclosure requirements.
Block H – Risk Summary
The final block aggregates observational data:
- All flags from Block G: Culture screen results, sponsorship issues, AI-buzzword mismatches, and other risk signals
Block H compiles these into a compact ordered list (## Risk Summary) allowing immediate identification of critical concerns without modifying underlying scores.
Critical Data Files and Their Roles
| File Path | Pipeline Function |
|---|---|
modes/oferta.md |
Master schema defining Blocks A-H data requirements and execution order |
modes/_shared.md |
Shared archetype detection logic and scoring rubrics |
cv.md |
Source of candidate experience for matching (Block B) and story selection (Block F) |
config/profile.yml |
Candidate metadata including work authorization, geographic constraints, and compensation tiers |
data/blacklist.md |
Optional pre-flight gate for company exclusion |
data/scan-history.tsv |
Historical posting data powering reposting detection in Block G |
interview-prep/story-bank.md |
Repository of STAR+R narratives for interview alignment |
templates/agency-licensing.yml |
Jurisdiction-specific staffing agency regulatory data |
templates/immigration-status-requirements.yml |
Work authorization requirement validation rules |
templates/jurisdiction-prohibited-content.yml |
Regional legal constraints on job posting content |
Data Flow and Processing Gates
Input Gate Validation
When a URL is supplied, the liveness gate executes browser_navigate and snapshot operations via Playwright to fetch the current page state. The system validates posting freshness before any evaluation blocks execute, preventing analysis of expired or removed listings.
Blacklist Pre-Check
If data/blacklist.md exists, CareerOps performs case-insensitive string matching against the company name immediately after the liveness gate. Matches trigger immediate pipeline abort before Block A processing begins.
Archetype Detection
The system uses the shared scoring system in modes/_shared.md to classify roles into six distinct archetypes. This classification drives content generation in Blocks B through F, ensuring evaluation criteria match role types.
Bounded Research Constraints
Block D enforces a strict maximum of five web-search queries for market data. This prevents excessive external API usage while maintaining sufficient data for compensation benchmarking.
Observational Risk Architecture
All Block G signals operate in read-only mode. These legitimacy indicators never modify numerical scores but append risk rows to the data structure consumed by Block H for final risk aggregation.
Running the A-H Evaluation
Execute a complete evaluation against a live posting URL:
node oferta.mjs https://example.com/jobs/1234
Process raw JD text without URL fetching:
node oferta.mjs --jd "Full-stack Engineer – Remote …"
Extract the generated risk summary from a specific report:
grep -A5 "## Risk Summary" reports/042-example-company-2024-08-22.md
Invoke the culture-screen component (Block A sub-process) independently:
node culture-screen.mjs --jd-file jd.txt --profile config/profile.yml
These commands demonstrate how CareerOps wires URL snapshots, text inputs, CV data, and profile configurations through the sequential A-H evaluation pipeline.
Summary
- Eight sequential blocks (A-H) process job postings through distinct analytical lenses, each with specific data requirements defined in
modes/oferta.md. - Primary inputs include Playwright browser snapshots,
cv.md,config/profile.yml, and jurisdiction-specific templates for compliance checking. - Block G performs the most comprehensive data synthesis, consuming scan history, agency licensing tables, and immigration requirements to generate 15+ distinct legitimacy signals.
- Block H operates as a pure aggregation layer, collating risk flags without mutating underlying analysis scores.
- Bounded research limits (5 queries maximum in Block D) and blacklist pre-checks prevent resource exhaustion and wasted computation on undesirable companies.
Frequently Asked Questions
What triggers the A-H Evaluation block sequence in CareerOps?
The pipeline initiates when you invoke node oferta.mjs with either a URL or raw JD text. The liveness gate first validates URL accessibility via Playwright, then checks data/blacklist.md for company exclusions before Block A begins processing archetype detection.
How does CareerOps handle live job posting URLs?
The system uses Playwright's browser_navigate and snapshot functions to capture the current DOM state and visible text. This snapshot serves as the authoritative JD source for Blocks A and G, ensuring analysis reflects the live posting state rather than cached or stale data.
What is the bounded research budget mentioned in Block D?
CareerOps enforces a strict limit of five web-search queries during compensation research in Block D. This constraint prevents excessive search API consumption while gathering sufficient market data to classify salary reliability tiers and benchmark components.
Can I customize the data sources for Block G legitimacy checks?
Yes. Block G consumes external reference data from the templates/ directory. You can modify templates/agency-licensing.yml, templates/immigration-status-requirements.yml, and templates/jurisdiction-prohibited-content.yml to add jurisdiction-specific rules, licensing requirements, or prohibited content categories that match your regional compliance needs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →