How to Check the Provenance of Figures in the story-bank: Complete Verification Guide

TLDR: Run node story-provenance-check.mjs to audit numeric claims in interview-prep/story-bank.md against your cv.md, which categorizes each figure into one of four provenance buckets ranging from verified to explicitly unknown.

The career-ops repository maintains a zero-LLM verification system that ensures every statistic in your interview preparation stories traces back to a primary source. The story-provenance-check.mjs script performs a read-only analysis of interview-prep/story-bank.md, comparing numeric figures against your canonical cv.md to prevent unverified claims from entering formal job applications. Learning how to check the provenance of figures in the story-bank guarantees that every number in your CV, cover letters, and interview responses is either verified or explicitly flagged.

How the Provenance Checker Classifies Numeric Claims

The classification logic is defined in the header of story-provenance-check.mjs (lines 4-70). Each figure is assigned to exactly one of four provenance buckets based on automated heuristics and explicit user markers.

Existing (Verified in CV)

Figures marked as existing appear verbatim in your cv.md with matching surrounding context, or carry an explicit provenance marker. The checker validates these through scoped number matching, which verifies both the digits and the semantic context align with your resume.

Supported by Resume

Claims classified as supportedByResume reference concepts appearing in your CV but lack the exact numeric match. The checker identifies these through context-overlap heuristics, comparing words surrounding the figure against your entire cv.md content.

Derived-Unverified

The derived-unverified bucket collects figures found only in the story-bank with no corresponding cv.md entry and no provenance marker. These represent potential fabrications or remembered statistics requiring verification before use.

User-Cannot-Confirm

When you mark a claim with **Provenance:** user-cannot-confirm, the checker permanently assigns it to the user-cannot-confirm bucket. This marker overrides all automated heuristics and prevents the figure from ever being promoted to verified status (lines 65-71).

Running the story-provenance-check.mjs Script

The script operates entirely in read-only mode, ensuring it never modifies your source files during analysis.

Command-Line Usage

Execute the checker from the repository root using Node.js:


# Default JSON output showing all four buckets

node story-provenance-check.mjs

# Human-readable summary view

node story-provenance-check.mjs --summary

# Verify script integrity against bundled test suite

node story-provenance-check.mjs --self-test

By default, the script locates interview-prep/story-bank.md and cv.md in the current directory. Override these paths using the --story-bank and --cv flags (defined in lines 55-58):

node story-provenance-check.mjs --story-bank ./custom/stories.md --cv ./resume/cv.md

Interpreting JSON Output

The default output is a JSON object containing four arrays corresponding to the provenance buckets. Each entry includes the story title, raw claim text, regex pattern detected, and classification reason:

{
  "existing": [
    { "story": "Onboarding Workflow", "claim": "8 hours to 2 hours", "pattern": "hour-range", "reason": "matched scoped number" }
  ],
  "supportedByResume": [
    { "story": "LMS Migration Team", "claim": "15-person team", "pattern": "scale-hyphen", "contextWords": ["team","migration"] }
  ],
  "derivedUnverified": [
    { "story": "Statewide Rollout", "claim": "500+ employees", "pattern": "plus-noun" }
  ],
  "userCannotConfirm": [
    { "story": "Vendor Negotiation", "claim": "40% savings", "pattern": "percent", "reason": "explicit Provenance marker (user-cannot-confirm)" }
  ]
}

Adding Provenance Markers to story-bank.md

Story blocks follow the ### heading convention established by match-star.mjs. Within each block, add a **Provenance:** line to manually set verification status.

Marker Syntax and Examples

Accepted formats include source citations, user confirmations, and explicit uncertainty markers:


### [Scale] Statewide Rollout

**Provenance:** source: cv.md
**Situation:** We needed to deploy across all offices.
**Task:** Lead the statewide initiative.
**Action:** Coordinated with regional managers.
**Result:** 500+ employees completed training within the quarter.

Valid marker values:

  • **Provenance:** source: cv.md — Treats the claim as verified based on CV content
  • **Provenance:** user-stated 2023-11-05 — Confirms manual verification on the specified date
  • **Provenance:** derived-unverified — Explicitly marks as unverified (default if omitted)
  • **Provenance:** user-cannot-confirm — Permanently excludes the figure from the verified bucket

Automated Verification Workflows

Integrate the checker into custom tooling to block unverified figures from publication pipelines.

Shell Integration

Create a wrapper script for repeated execution:

// utils/checkProvenance.js
import { execSync } from 'child_process';

function runProvenanceCheck() {
  const raw = execSync('node story-provenance-check.mjs --summary', {
    encoding: 'utf8',
  });
  console.log(raw);
}

runProvenanceCheck();

Programmatic Access

Parse the JSON output to enforce policies in build systems:

import { execFileSync } from 'child_process';

function getUnverifiedClaims() {
  const output = execFileSync(
    'node',
    ['story-provenance-check.mjs'],
    { encoding: 'utf8' }
  );
  const result = JSON.parse(output);
  return result.derivedUnverified;
}

// Fail CI if unverified figures exist
const unverified = getUnverifiedClaims();
if (unverified.length > 0) {
  console.error('Build blocked: Unverified figures detected', unverified);
  process.exit(1);
}

Summary

  • The story-provenance-check.mjs script provides read-only verification of numeric claims in interview-prep/story-bank.md against your cv.md
  • Four provenance buckets classify figures: existing, supportedByResume, derived-unverified, and user-cannot-confirm
  • Run node story-provenance-check.mjs for JSON output or append --summary for human-readable reports
  • Add **Provenance:** markers to story blocks to manually verify figures or mark them as unverifiable
  • Use --story-bank and --cv flags to process files in non-standard locations

Frequently Asked Questions

What file paths does the provenance checker use by default?

The script defaults to interview-prep/story-bank.md for story content and cv.md for the canonical source of truth. Override either path using the --story-bank <path> and --cv <path> command-line arguments.

How do I permanently mark a figure as unverifiable?

Add **Provenance:** user-cannot-confirm to the story block containing the figure. This marker takes precedence over all automated heuristics and ensures the claim always sorts into the userCannotConfirm bucket, preventing accidental verification in future runs.

Can the provenance checker modify my story files?

No. The story-provenance-check.mjs script operates strictly in read-only mode. It parses your markdown files and outputs JSON or summary text to stdout without writing changes back to disk, ensuring your source files remain untouched.

What is the difference between "existing" and "supportedByResume" classifications?

An existing classification means the exact number appears in your CV with matching context or carries an explicit provenance marker. supportedByResume indicates the concept appears in your CV but the specific figure does not, suggesting the number may be approximate or requires additional verification before use in formal documents.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →