# How Career-Ops Handles Quantified Claims from Derived Sources Like Story-Bank

> Learn how career-ops uses a provenance checker to handle quantified claims from sources like story-bank, ensuring interview anecdotes don't become unverified resume facts.

- Repository: [Santiago Fernández de Valderrama/career-ops](https://github.com/santifer/career-ops)
- Tags: internals
- Published: 2026-08-28

---

**Career-ops uses a read-only provenance checker to classify numeric claims from [`story-bank.md`](https://github.com/santifer/career-ops/blob/main/story-bank.md) into four distinct buckets, preventing unverified interview anecdotes from silently becoming resume facts.**

The `santifer/career-ops` repository implements a strict source-of-truth hierarchy to manage factual integrity across career documents. When users draft behavioral stories in [`interview-prep/story-bank.md`](https://github.com/santifer/career-ops/blob/main/interview-prep/story-bank.md), they often include approximate numbers—attendance figures, budget estimates, or efficiency percentages—that may not yet appear in the primary [`cv.md`](https://github.com/santifer/career-ops/blob/main/cv.md). The system treats these **quantified claims from derived sources** with specialized skepticism, ensuring that only verified or explicitly marked data propagates to generated CVs and cover letters.

## The Two-Tier Trust Model

Career-ops maintains a hard boundary between primary and derived sources:

- **Primary tier ([`cv.md`](https://github.com/santifer/career-ops/blob/main/cv.md))**: User-authored content with **full trust**. Numbers here are treated as ground truth.
- **Derived tier ([`interview-prep/story-bank.md`](https://github.com/santifer/career-ops/blob/main/interview-prep/story-bank.md))**: Accumulated STAR+R stories with **narrative trust only**. Quantified claims here require provenance verification before promotion to the primary tier.

According to the [`AGENTS.md`](https://github.com/santifer/career-ops/blob/main/AGENTS.md) architecture document, any quantified claim, scale figure, or scope-of-responsibility claim originating in a derived file must trace to a primary file or carry an explicit provenance marker. Absent such markers, the system classifies the claim as `derived-unverified`, blocking its use in authoritative outputs.

## The Four Classification Buckets

The provenance logic in `story-provenance-check.mjs` sorts each numeric pattern into one of four buckets:

1. **existing**: The exact number or a verified provenance marker is found in [`cv.md`](https://github.com/santifer/career-ops/blob/main/cv.md), or an explicit marker like `source: cv.md` or `user-stated YYYY-MM-DD` forces this classification.
2. **supportedByResume**: The claim's surrounding context words overlap with [`cv.md`](https://github.com/santifer/career-ops/blob/main/cv.md), making the fact plausible but not precisely verified.
3. **derived-unverified**: The number exists only in the story-bank with no provenance marker and no contextual overlap with the primary CV.
4. **user-cannot-confirm**: An explicit `**Provenance:** user-cannot-confirm` marker overrides all heuristics, guaranteeing the claim is never treated as verified and remains narrative-only.

## Provenance Checking Implementation

The core classification logic lives in **`story-provenance-check.mjs`**. This script parses each story block using numeric pattern extraction—identifying percentages, "+ noun" constructs, hour ranges, scale-hyphen patterns, and scale-noun phrases. It then cross-references both the extracted numbers and their contextual sentences against [`cv.md`](https://github.com/santifer/career-ops/blob/main/cv.md) to determine the appropriate bucket.

The checker is **read-only** by design. It reports classifications (such as `derived-unverified`) but never modifies source files. A separate UX flow handles user prompts to confirm, correct, or discard claims, writing the appropriate `**Provenance:**` line back into [`story-bank.md`](https://github.com/santifer/career-ops/blob/main/story-bank.md) only after explicit user action. This architecture guarantees that a claim marked `user-cannot-confirm` never drifts back into a verified state through automated processes.

## Working with the Provenance Checker

You can invoke the provenance system via command line or programmatically.

**Run the summary report:**

```bash
node story-provenance-check.mjs --summary

```

This outputs a human-readable report showing each claim, its detected pattern, the assigned bucket, and the provenance reason.

**Programmatic integration:**

```javascript
import { classifyStoryBank } from './story-provenance-check.mjs';
import { readFileSync } from 'fs';

const storyBank = readFileSync('interview-prep/story-bank.md', 'utf-8');
const cv = readFileSync('cv.md', 'utf-8');

const result = classifyStoryBank(storyBank, cv);
console.log('Derived‑unverified claims:', result.derivedUnverified);

```

The returned `result` object contains four arrays (`existing`, `supportedByResume`, `derivedUnverified`, `userCannotConfirm`) that downstream generators consume when building CVs, cover letters, or interview scripts.

## Preventing Factual Drift with Provenance Markers

Explicit markers in [`story-bank.md`](https://github.com/santifer/career-ops/blob/main/story-bank.md) override heuristic classification:

**Verified claim with date marker:**

```markdown

### [Scale] Conference Talk

**Provenance:** user‑stated 2026‑01‑15
**Situation:** Presented at a regional conference.
**Task:** Design a session for a large mixed audience.
**Action:** Delivered a workshop attended by roughly 300 students.
**Result:** Session feedback scores were among the highest of the day.

```

Despite [`cv.md`](https://github.com/santifer/career-ops/blob/main/cv.md) not containing "300 students," the `user-stated` marker forces classification into **existing**, allowing the number to appear in generated documents.

**Explicitly unverifiable claim:**

```markdown

### [Budget] Vendor Negotiation

**Provenance:** user‑cannot‑confirm
**Result:** Estimated 40% savings on licensing costs.

```

The `user-cannot-confirm` marker forces this into its own bucket, ensuring downstream generators treat it as unsupported narrative only.

## Summary

- **Career-ops** enforces a strict boundary between primary ([`cv.md`](https://github.com/santifer/career-ops/blob/main/cv.md)) and derived ([`story-bank.md`](https://github.com/santifer/career-ops/blob/main/story-bank.md)) sources to prevent unverified numbers from becoming facts.
- The **`story-provenance-check.mjs`** script classifies quantified claims into four buckets: **existing**, **supportedByResume**, **derived-unverified**, and **user-cannot-confirm**.
- **Provenance markers** (`user-stated YYYY-MM-DD`, `source: cv.md`, `user-cannot-confirm`) manually override automatic classification.
- The checking logic is **read-only**, ensuring that verification requires explicit user action through a separate UX flow.
- Numeric pattern extraction covers percentages, scale references, hour ranges, and other common resume quantification formats.

## Frequently Asked Questions

### How does career-ops prevent made-up numbers from appearing in generated CVs?

The system treats [`interview-prep/story-bank.md`](https://github.com/santifer/career-ops/blob/main/interview-prep/story-bank.md) as a derived source with narrative trust only. Any quantified claim found only in this file is classified as `derived-unverified` by the provenance checker in `story-provenance-check.mjs`. Downstream generators ignore these claims unless the user explicitly confirms them and adds a provenance marker like `user-stated YYYY-MM-DD`.

### What is the difference between supportedByResume and derived-unverified buckets?

**supportedByResume** indicates that while the exact number isn't in [`cv.md`](https://github.com/santifer/career-ops/blob/main/cv.md), the surrounding context words overlap with the primary CV, suggesting the claim is plausible. **derived-unverified** means the number appears exclusively in the story-bank with no textual overlap or provenance marker, indicating it may be an estimate invented during interview preparation.

### Can the provenance checker automatically update my story-bank.md file?

No. The provenance check implemented in `story-provenance-check.mjs` is strictly read-only. It reports classifications but never writes to source files. A separate UX workflow handles user confirmations and writes `**Provenance:**` markers back to [`story-bank.md`](https://github.com/santifer/career-ops/blob/main/story-bank.md), ensuring that `user-cannot-confirm` claims cannot be automatically reclassified.

### Why does the system use four buckets instead of three?

While the companion script `jd-skill-gap.mjs` uses three buckets for skill mentions, numeric claims require an additional `user-cannot-confirm` tier. This fourth bucket explicitly handles cases where users acknowledge they cannot verify a number, preventing accidental "laundering" of guesses into facts through subsequent heuristic checks.