How to Check the Provenance of Figures in the story-bank: Complete Verification Guide
TLDR: Run node story-provenance-check.mjs to audit numeric claims in interview-prep/story-bank.md against your cv.md, which categorizes each figure into one of four provenance buckets ranging from verified to explicitly unknown.
The career-ops repository maintains a zero-LLM verification system that ensures every statistic in your interview preparation stories traces back to a primary source. The story-provenance-check.mjs script performs a read-only analysis of interview-prep/story-bank.md, comparing numeric figures against your canonical cv.md to prevent unverified claims from entering formal job applications. Learning how to check the provenance of figures in the story-bank guarantees that every number in your CV, cover letters, and interview responses is either verified or explicitly flagged.
How the Provenance Checker Classifies Numeric Claims
The classification logic is defined in the header of story-provenance-check.mjs (lines 4-70). Each figure is assigned to exactly one of four provenance buckets based on automated heuristics and explicit user markers.
Existing (Verified in CV)
Figures marked as existing appear verbatim in your cv.md with matching surrounding context, or carry an explicit provenance marker. The checker validates these through scoped number matching, which verifies both the digits and the semantic context align with your resume.
Supported by Resume
Claims classified as supportedByResume reference concepts appearing in your CV but lack the exact numeric match. The checker identifies these through context-overlap heuristics, comparing words surrounding the figure against your entire cv.md content.
Derived-Unverified
The derived-unverified bucket collects figures found only in the story-bank with no corresponding cv.md entry and no provenance marker. These represent potential fabrications or remembered statistics requiring verification before use.
User-Cannot-Confirm
When you mark a claim with **Provenance:** user-cannot-confirm, the checker permanently assigns it to the user-cannot-confirm bucket. This marker overrides all automated heuristics and prevents the figure from ever being promoted to verified status (lines 65-71).
Running the story-provenance-check.mjs Script
The script operates entirely in read-only mode, ensuring it never modifies your source files during analysis.
Command-Line Usage
Execute the checker from the repository root using Node.js:
# Default JSON output showing all four buckets
node story-provenance-check.mjs
# Human-readable summary view
node story-provenance-check.mjs --summary
# Verify script integrity against bundled test suite
node story-provenance-check.mjs --self-test
By default, the script locates interview-prep/story-bank.md and cv.md in the current directory. Override these paths using the --story-bank and --cv flags (defined in lines 55-58):
node story-provenance-check.mjs --story-bank ./custom/stories.md --cv ./resume/cv.md
Interpreting JSON Output
The default output is a JSON object containing four arrays corresponding to the provenance buckets. Each entry includes the story title, raw claim text, regex pattern detected, and classification reason:
{
"existing": [
{ "story": "Onboarding Workflow", "claim": "8 hours to 2 hours", "pattern": "hour-range", "reason": "matched scoped number" }
],
"supportedByResume": [
{ "story": "LMS Migration Team", "claim": "15-person team", "pattern": "scale-hyphen", "contextWords": ["team","migration"] }
],
"derivedUnverified": [
{ "story": "Statewide Rollout", "claim": "500+ employees", "pattern": "plus-noun" }
],
"userCannotConfirm": [
{ "story": "Vendor Negotiation", "claim": "40% savings", "pattern": "percent", "reason": "explicit Provenance marker (user-cannot-confirm)" }
]
}
Adding Provenance Markers to story-bank.md
Story blocks follow the ### heading convention established by match-star.mjs. Within each block, add a **Provenance:** line to manually set verification status.
Marker Syntax and Examples
Accepted formats include source citations, user confirmations, and explicit uncertainty markers:
### [Scale] Statewide Rollout
**Provenance:** source: cv.md
**Situation:** We needed to deploy across all offices.
**Task:** Lead the statewide initiative.
**Action:** Coordinated with regional managers.
**Result:** 500+ employees completed training within the quarter.
Valid marker values:
**Provenance:** source: cv.md— Treats the claim as verified based on CV content**Provenance:** user-stated 2023-11-05— Confirms manual verification on the specified date**Provenance:** derived-unverified— Explicitly marks as unverified (default if omitted)**Provenance:** user-cannot-confirm— Permanently excludes the figure from the verified bucket
Automated Verification Workflows
Integrate the checker into custom tooling to block unverified figures from publication pipelines.
Shell Integration
Create a wrapper script for repeated execution:
// utils/checkProvenance.js
import { execSync } from 'child_process';
function runProvenanceCheck() {
const raw = execSync('node story-provenance-check.mjs --summary', {
encoding: 'utf8',
});
console.log(raw);
}
runProvenanceCheck();
Programmatic Access
Parse the JSON output to enforce policies in build systems:
import { execFileSync } from 'child_process';
function getUnverifiedClaims() {
const output = execFileSync(
'node',
['story-provenance-check.mjs'],
{ encoding: 'utf8' }
);
const result = JSON.parse(output);
return result.derivedUnverified;
}
// Fail CI if unverified figures exist
const unverified = getUnverifiedClaims();
if (unverified.length > 0) {
console.error('Build blocked: Unverified figures detected', unverified);
process.exit(1);
}
Summary
- The
story-provenance-check.mjsscript provides read-only verification of numeric claims ininterview-prep/story-bank.mdagainst yourcv.md - Four provenance buckets classify figures: existing, supportedByResume, derived-unverified, and user-cannot-confirm
- Run
node story-provenance-check.mjsfor JSON output or append--summaryfor human-readable reports - Add
**Provenance:**markers to story blocks to manually verify figures or mark them as unverifiable - Use
--story-bankand--cvflags to process files in non-standard locations
Frequently Asked Questions
What file paths does the provenance checker use by default?
The script defaults to interview-prep/story-bank.md for story content and cv.md for the canonical source of truth. Override either path using the --story-bank <path> and --cv <path> command-line arguments.
How do I permanently mark a figure as unverifiable?
Add **Provenance:** user-cannot-confirm to the story block containing the figure. This marker takes precedence over all automated heuristics and ensures the claim always sorts into the userCannotConfirm bucket, preventing accidental verification in future runs.
Can the provenance checker modify my story files?
No. The story-provenance-check.mjs script operates strictly in read-only mode. It parses your markdown files and outputs JSON or summary text to stdout without writing changes back to disk, ensuring your source files remain untouched.
What is the difference between "existing" and "supportedByResume" classifications?
An existing classification means the exact number appears in your CV with matching context or carries an explicit provenance marker. supportedByResume indicates the concept appears in your CV but the specific figure does not, suggesting the number may be approximate or requires additional verification before use in formal documents.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →