How thesis-drift Detects Changes in Investment Thesis Documents: A Technical Deep Dive
thesis-drift automatically identifies material factual changes in investment thesis documents by extracting structured evidence from markdown reports, recomputing numerical data with specialized tools, and applying a deterministic three-state assessment across five predefined dimensions.
The thesis-drift skill in the xbtlin/ai-berkshire repository provides rigorous change detection for investment research. Unlike simple text diffing, this Claude Code skill distinguishes substantive factual drift from superficial wording edits by leveraging structured data extraction and evidence-based validation.
Three Operational Modes for Comparison
The skill adapts its workflow based on how it is invoked, parsed from $ARGUMENTS as defined in skills/thesis-drift.md (lines 27-34).
Mode A: Explicit File Comparison
When users provide specific file paths, the skill loads two explicit markdown reports and executes a direct comparison.
/thesis-drift Apple reports/Apple-thesis-2023-03-15.md reports/Apple-thesis-2024-03-15.md
Mode B: Automatic Snapshot Discovery
When invoked with only a company name, the skill automatically discovers thesis snapshots in the reports/ directory, selecting the earliest complete file as the baseline and the latest as the new report.
/thesis-drift Tesla
Mode C: Missing Baseline Handling
If only a current thesis exists, the skill switches to Mode C and instructs the user to establish a historical baseline.
/thesis-drift Zoom
When no baseline is found, the skill responds with a prompt to first run:
/thesis-tracker Zoom
The Evidence-Driven Detection Pipeline
The core detection logic follows a strict, sequential workflow designed to eliminate linguistic ambiguity and enforce numerical rigor.
Structured Data Extraction
In skills/thesis-drift.md (lines 40-52), the skill parses each markdown report to extract a deterministic set of fields:
- Date and ticker
- Core hypothesis list
- Red-line list
- Valuation anchors
- Tracking table
- Management quality assessment
- Competitive moat assessment
- Recommended action
Missing sections are flagged as "结构缺失" (structurally missing) rather than fabricated, allowing the comparison to proceed with available data.
Evidence Normalization
The skill constructs a side-by-side evidence table (lines 54-66) that aligns every dimension with the factual evidence extracted from both the old and new reports. This normalization step separates what changed from how it was worded, ensuring that semantic drift does not trigger false positives.
Numerical Verification with financial_rigor.py
To avoid "LLM mental math," every numerical change—whether in price, EPS, BVPS, FCF per share, or market capitalization—is recomputed using tools/financial_rigor.py (lines 70-88). This tool cross-validates fields against multiple sources and ensures reproducibility.
python3 tools/financial_rigor.py verify-valuation \
--price 255.30 \
--eps 3.45 \
--bvps 12.80 \
--fcf-per-share 2.10
Dimension-Wise Drift Assessment
For each of the five fixed dimensions—valuation anchors, core hypothesis, red-line list, management quality, and competitive moat—the skill applies a deterministic three-state decision matrix (lines 93-101):
- Improved
- Unchanged
- Weakened
This explicit decision matrix removes ambiguity and enables reliable downstream automation, such as triggering portfolio rebalancing workflows when specific thresholds are breached.
Evidence-Driven Rule Enforcement
Any verdict other than Unchanged must cite concrete new evidence (lines 105-113). Acceptable evidence includes quarterly earnings figures, regulatory filings, management changes, or quantifiable price-valuation shifts. If no supporting evidence exists, the change is marked Unchanged or Unable to determine ("无法判断"), preventing hallucinated conclusions.
Safety Mechanisms and Validation Rules
The skill implements strict guardrails (lines 33-38, 78-86) to maintain analytical integrity:
- Cross-company refusal: The skill refuses to compare thesis documents from different companies.
- Traceability requirement: All conclusions must be traceable to verifiable sources.
- Graceful degradation: When structural data is missing, the skill marks the gap explicitly rather than fabricating comparisons.
Key Source Files and Implementation
| File | Role |
|---|---|
skills/thesis-drift.md |
Main skill definition containing the argument parser, evidence table builder, and decision logic |
tools/financial_rigor.py |
Numerical validation library for valuation multiples and market-cap calculations |
skills/thesis-tracker.md |
Generates the structured thesis snapshots consumed by thesis-drift |
Summary
- thesis-drift operates in three modes (explicit files, auto-discovery, and missing baseline) to accommodate different user workflows.
- The skill extracts structured data from markdown reports and normalizes evidence into comparable dimensions.
- All numerical calculations are delegated to
tools/financial_rigor.pyto ensure mathematical accuracy. - A deterministic three-state decision matrix (Improved/Unchanged/Weakened) evaluates five fixed dimensions: valuation, hypothesis, red-lines, management, and moat.
- Evidence-driven rules require concrete citations for any non-Unchanged verdict, preventing false drift detection.
Frequently Asked Questions
What input format does thesis-drift require?
The skill expects markdown files generated by the /thesis-tracker command in the ai-berkshire repository. These files must contain specific structured sections including core hypothesis lists, red-line lists, valuation anchors, and management assessments. While the skill attempts extraction from non-standard formats, optimal results require the structured output from skills/thesis-tracker.md.
How does thesis-drift prevent false positives from wording changes?
By normalizing evidence into a side-by-side table that aligns factual data rather than text strings, the skill separates linguistic variation from substantive change. The comparison focuses on extracted evidence values (e.g., specific EPS figures, management ratings) rather than semantic similarity, ensuring that reworded content does not trigger false drift alerts.
Can thesis-drift compare thesis documents from different companies?
No. The skill explicitly refuses cross-company comparisons as a safety measure. If the ticker symbols extracted from the two documents do not match, the skill aborts the operation and requests corrected inputs, preventing erroneous analytical conclusions that could arise from comparing fundamentally different investment contexts.
What happens if numerical data is inconsistent between sources?
When tools/financial_rigor.py detects inconsistencies during verification, it flags the discrepancy in the drift report. The skill marks the specific dimension as Unable to determine ("无法判断") rather than selecting the most recent figure, requiring the analyst to manually resolve the data conflict before finalizing the drift assessment.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →