How Career-Ops Handles Quantified Claims from Derived Sources Like Story-Bank

Career-ops uses a read-only provenance checker to classify numeric claims from story-bank.md into four distinct buckets, preventing unverified interview anecdotes from silently becoming resume facts.

The santifer/career-ops repository implements a strict source-of-truth hierarchy to manage factual integrity across career documents. When users draft behavioral stories in interview-prep/story-bank.md, they often include approximate numbers—attendance figures, budget estimates, or efficiency percentages—that may not yet appear in the primary cv.md. The system treats these quantified claims from derived sources with specialized skepticism, ensuring that only verified or explicitly marked data propagates to generated CVs and cover letters.

The Two-Tier Trust Model

Career-ops maintains a hard boundary between primary and derived sources:

  • Primary tier (cv.md): User-authored content with full trust. Numbers here are treated as ground truth.
  • Derived tier (interview-prep/story-bank.md): Accumulated STAR+R stories with narrative trust only. Quantified claims here require provenance verification before promotion to the primary tier.

According to the AGENTS.md architecture document, any quantified claim, scale figure, or scope-of-responsibility claim originating in a derived file must trace to a primary file or carry an explicit provenance marker. Absent such markers, the system classifies the claim as derived-unverified, blocking its use in authoritative outputs.

The Four Classification Buckets

The provenance logic in story-provenance-check.mjs sorts each numeric pattern into one of four buckets:

  1. existing: The exact number or a verified provenance marker is found in cv.md, or an explicit marker like source: cv.md or user-stated YYYY-MM-DD forces this classification.
  2. supportedByResume: The claim's surrounding context words overlap with cv.md, making the fact plausible but not precisely verified.
  3. derived-unverified: The number exists only in the story-bank with no provenance marker and no contextual overlap with the primary CV.
  4. user-cannot-confirm: An explicit **Provenance:** user-cannot-confirm marker overrides all heuristics, guaranteeing the claim is never treated as verified and remains narrative-only.

Provenance Checking Implementation

The core classification logic lives in story-provenance-check.mjs. This script parses each story block using numeric pattern extraction—identifying percentages, "+ noun" constructs, hour ranges, scale-hyphen patterns, and scale-noun phrases. It then cross-references both the extracted numbers and their contextual sentences against cv.md to determine the appropriate bucket.

The checker is read-only by design. It reports classifications (such as derived-unverified) but never modifies source files. A separate UX flow handles user prompts to confirm, correct, or discard claims, writing the appropriate **Provenance:** line back into story-bank.md only after explicit user action. This architecture guarantees that a claim marked user-cannot-confirm never drifts back into a verified state through automated processes.

Working with the Provenance Checker

You can invoke the provenance system via command line or programmatically.

Run the summary report:

node story-provenance-check.mjs --summary

This outputs a human-readable report showing each claim, its detected pattern, the assigned bucket, and the provenance reason.

Programmatic integration:

import { classifyStoryBank } from './story-provenance-check.mjs';
import { readFileSync } from 'fs';

const storyBank = readFileSync('interview-prep/story-bank.md', 'utf-8');
const cv = readFileSync('cv.md', 'utf-8');

const result = classifyStoryBank(storyBank, cv);
console.log('Derived‑unverified claims:', result.derivedUnverified);

The returned result object contains four arrays (existing, supportedByResume, derivedUnverified, userCannotConfirm) that downstream generators consume when building CVs, cover letters, or interview scripts.

Preventing Factual Drift with Provenance Markers

Explicit markers in story-bank.md override heuristic classification:

Verified claim with date marker:


### [Scale] Conference Talk

**Provenance:** user‑stated 2026‑01‑15
**Situation:** Presented at a regional conference.
**Task:** Design a session for a large mixed audience.
**Action:** Delivered a workshop attended by roughly 300 students.
**Result:** Session feedback scores were among the highest of the day.

Despite cv.md not containing "300 students," the user-stated marker forces classification into existing, allowing the number to appear in generated documents.

Explicitly unverifiable claim:


### [Budget] Vendor Negotiation

**Provenance:** user‑cannot‑confirm
**Result:** Estimated 40% savings on licensing costs.

The user-cannot-confirm marker forces this into its own bucket, ensuring downstream generators treat it as unsupported narrative only.

Summary

  • Career-ops enforces a strict boundary between primary (cv.md) and derived (story-bank.md) sources to prevent unverified numbers from becoming facts.
  • The story-provenance-check.mjs script classifies quantified claims into four buckets: existing, supportedByResume, derived-unverified, and user-cannot-confirm.
  • Provenance markers (user-stated YYYY-MM-DD, source: cv.md, user-cannot-confirm) manually override automatic classification.
  • The checking logic is read-only, ensuring that verification requires explicit user action through a separate UX flow.
  • Numeric pattern extraction covers percentages, scale references, hour ranges, and other common resume quantification formats.

Frequently Asked Questions

How does career-ops prevent made-up numbers from appearing in generated CVs?

The system treats interview-prep/story-bank.md as a derived source with narrative trust only. Any quantified claim found only in this file is classified as derived-unverified by the provenance checker in story-provenance-check.mjs. Downstream generators ignore these claims unless the user explicitly confirms them and adds a provenance marker like user-stated YYYY-MM-DD.

What is the difference between supportedByResume and derived-unverified buckets?

supportedByResume indicates that while the exact number isn't in cv.md, the surrounding context words overlap with the primary CV, suggesting the claim is plausible. derived-unverified means the number appears exclusively in the story-bank with no textual overlap or provenance marker, indicating it may be an estimate invented during interview preparation.

Can the provenance checker automatically update my story-bank.md file?

No. The provenance check implemented in story-provenance-check.mjs is strictly read-only. It reports classifications but never writes to source files. A separate UX workflow handles user confirmations and writes **Provenance:** markers back to story-bank.md, ensuring that user-cannot-confirm claims cannot be automatically reclassified.

Why does the system use four buckets instead of three?

While the companion script jd-skill-gap.mjs uses three buckets for skill mentions, numeric claims require an additional user-cannot-confirm tier. This fourth bucket explicitly handles cases where users acknowledge they cannot verify a number, preventing accidental "laundering" of guesses into facts through subsequent heuristic checks.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →