# How thesis-drift Detects Changes in Investment Thesis Documents: A Technical Deep Dive

> Discover how thesis-drift technically detects changes in investment thesis documents. It extracts evidence, recomputes data, and applies a deterministic assessment across five dimensions.

- Repository: [Xbt Lin/ai-berkshire](https://github.com/xbtlin/ai-berkshire)
- Tags: deep-dive
- Published: 2026-07-26

---

**thesis-drift automatically identifies material factual changes in investment thesis documents by extracting structured evidence from markdown reports, recomputing numerical data with specialized tools, and applying a deterministic three-state assessment across five predefined dimensions.**

The `thesis-drift` skill in the **xbtlin/ai-berkshire** repository provides rigorous change detection for investment research. Unlike simple text diffing, this Claude Code skill distinguishes substantive factual drift from superficial wording edits by leveraging structured data extraction and evidence-based validation.

## Three Operational Modes for Comparison

The skill adapts its workflow based on how it is invoked, parsed from `$ARGUMENTS` as defined in [`skills/thesis-drift.md`](https://github.com/xbtlin/ai-berkshire/blob/main/skills/thesis-drift.md) [(lines 27-34)](https://github.com/xbtlin/ai-berkshire/blob/main/skills/thesis-drift.md#L27-L34).

### Mode A: Explicit File Comparison

When users provide specific file paths, the skill loads two explicit markdown reports and executes a direct comparison.

```text
/thesis-drift Apple reports/Apple-thesis-2023-03-15.md reports/Apple-thesis-2024-03-15.md

```

### Mode B: Automatic Snapshot Discovery

When invoked with only a company name, the skill automatically discovers thesis snapshots in the `reports/` directory, selecting the earliest complete file as the baseline and the latest as the new report.

```text
/thesis-drift Tesla

```

### Mode C: Missing Baseline Handling

If only a current thesis exists, the skill switches to Mode C and instructs the user to establish a historical baseline.

```text
/thesis-drift Zoom

```

When no baseline is found, the skill responds with a prompt to first run:

```text
/thesis-tracker Zoom

```

## The Evidence-Driven Detection Pipeline

The core detection logic follows a strict, sequential workflow designed to eliminate linguistic ambiguity and enforce numerical rigor.

### Structured Data Extraction

In [`skills/thesis-drift.md`](https://github.com/xbtlin/ai-berkshire/blob/main/skills/thesis-drift.md) [(lines 40-52)](https://github.com/xbtlin/ai-berkshire/blob/main/skills/thesis-drift.md#L40-L52), the skill parses each markdown report to extract a deterministic set of fields:

- Date and ticker
- Core hypothesis list
- Red-line list
- Valuation anchors
- Tracking table
- Management quality assessment
- Competitive moat assessment
- Recommended action

Missing sections are flagged as "结构缺失" (structurally missing) rather than fabricated, allowing the comparison to proceed with available data.

### Evidence Normalization

The skill constructs a side-by-side evidence table [(lines 54-66)](https://github.com/xbtlin/ai-berkshire/blob/main/skills/thesis-drift.md#L54-L66) that aligns every dimension with the factual evidence extracted from both the old and new reports. This normalization step separates *what* changed from *how* it was worded, ensuring that semantic drift does not trigger false positives.

### Numerical Verification with financial_rigor.py

To avoid "LLM mental math," every numerical change—whether in price, EPS, BVPS, FCF per share, or market capitalization—is recomputed using **tools/financial_rigor.py** [(lines 70-88)](https://github.com/xbtlin/ai-berkshire/blob/main/skills/thesis-drift.md#L70-L88). This tool cross-validates fields against multiple sources and ensures reproducibility.

```bash
python3 tools/financial_rigor.py verify-valuation \
  --price 255.30 \
  --eps 3.45 \
  --bvps 12.80 \
  --fcf-per-share 2.10

```

### Dimension-Wise Drift Assessment

For each of the five fixed dimensions—**valuation anchors**, **core hypothesis**, **red-line list**, **management quality**, and **competitive moat**—the skill applies a deterministic three-state decision matrix [(lines 93-101)](https://github.com/xbtlin/ai-berkshire/blob/main/skills/thesis-drift.md#L93-L101):

- **Improved**
- **Unchanged**
- **Weakened**

This explicit decision matrix removes ambiguity and enables reliable downstream automation, such as triggering portfolio rebalancing workflows when specific thresholds are breached.

### Evidence-Driven Rule Enforcement

Any verdict other than **Unchanged** must cite concrete new evidence [(lines 105-113)](https://github.com/xbtlin/ai-berkshire/blob/main/skills/thesis-drift.md#L105-L113). Acceptable evidence includes quarterly earnings figures, regulatory filings, management changes, or quantifiable price-valuation shifts. If no supporting evidence exists, the change is marked **Unchanged** or **Unable to determine** ("无法判断"), preventing hallucinated conclusions.

## Safety Mechanisms and Validation Rules

The skill implements strict guardrails [(lines 33-38, 78-86)](https://github.com/xbtlin/ai-berkshire/blob/main/skills/thesis-drift.md#L33-L38) to maintain analytical integrity:

- **Cross-company refusal**: The skill refuses to compare thesis documents from different companies.
- **Traceability requirement**: All conclusions must be traceable to verifiable sources.
- **Graceful degradation**: When structural data is missing, the skill marks the gap explicitly rather than fabricating comparisons.

## Key Source Files and Implementation

| File | Role |
|------|------|
| [`skills/thesis-drift.md`](https://github.com/xbtlin/ai-berkshire/blob/main/skills/thesis-drift.md) | Main skill definition containing the argument parser, evidence table builder, and decision logic |
| [`tools/financial_rigor.py`](https://github.com/xbtlin/ai-berkshire/blob/main/tools/financial_rigor.py) | Numerical validation library for valuation multiples and market-cap calculations |
| [`skills/thesis-tracker.md`](https://github.com/xbtlin/ai-berkshire/blob/main/skills/thesis-tracker.md) | Generates the structured thesis snapshots consumed by `thesis-drift` |

## Summary

- **thesis-drift** operates in three modes (explicit files, auto-discovery, and missing baseline) to accommodate different user workflows.
- The skill extracts structured data from markdown reports and normalizes evidence into comparable dimensions.
- All numerical calculations are delegated to [`tools/financial_rigor.py`](https://github.com/xbtlin/ai-berkshire/blob/main/tools/financial_rigor.py) to ensure mathematical accuracy.
- A deterministic three-state decision matrix (Improved/Unchanged/Weakened) evaluates five fixed dimensions: valuation, hypothesis, red-lines, management, and moat.
- Evidence-driven rules require concrete citations for any non-Unchanged verdict, preventing false drift detection.

## Frequently Asked Questions

### What input format does thesis-drift require?

The skill expects markdown files generated by the `/thesis-tracker` command in the ai-berkshire repository. These files must contain specific structured sections including core hypothesis lists, red-line lists, valuation anchors, and management assessments. While the skill attempts extraction from non-standard formats, optimal results require the structured output from [`skills/thesis-tracker.md`](https://github.com/xbtlin/ai-berkshire/blob/main/skills/thesis-tracker.md).

### How does thesis-drift prevent false positives from wording changes?

By normalizing evidence into a side-by-side table that aligns factual data rather than text strings, the skill separates linguistic variation from substantive change. The comparison focuses on extracted evidence values (e.g., specific EPS figures, management ratings) rather than semantic similarity, ensuring that reworded content does not trigger false drift alerts.

### Can thesis-drift compare thesis documents from different companies?

No. The skill explicitly refuses cross-company comparisons as a safety measure. If the ticker symbols extracted from the two documents do not match, the skill aborts the operation and requests corrected inputs, preventing erroneous analytical conclusions that could arise from comparing fundamentally different investment contexts.

### What happens if numerical data is inconsistent between sources?

When [`tools/financial_rigor.py`](https://github.com/xbtlin/ai-berkshire/blob/main/tools/financial_rigor.py) detects inconsistencies during verification, it flags the discrepancy in the drift report. The skill marks the specific dimension as **Unable to determine** ("无法判断") rather than selecting the most recent figure, requiring the analyst to manually resolve the data conflict before finalizing the drift assessment.