How Multi-Source Verification Works for Market Cap Data in AI-Berkshire

The ai-berkshire repository validates critical market cap figures through a two-step workflow that checks arithmetic integrity against share price data and cross-validates values across independent financial sources using decimal-precision calculations and configurable tolerance thresholds.

The xbtlin/ai-berkshire project implements rigorous financial data validation through its Financial Rigor Toolkit, ensuring that critical metrics like market capitalization are accurate before they feed into valuation models. This multi-source verification system combines direct calculation checks with cross-provider consensus validation to eliminate errors from unit mismatches, stale share counts, or single-source data corruption.

The Two-Step Verification Architecture

The verification system operates through a layered pipeline defined in tools/financial_rigor.py, requiring both mathematical consistency and cross-source agreement before data is deemed reliable.

Step 1: Direct Arithmetic Verification

The verify_market_cap function (lines 61-90) performs a primary sanity check by multiplying the reported share price by the total shares outstanding and comparing the result against the market cap disclosed in filings or data providers. The function calculates a deviation percentage and raises a hard error if the gap exceeds 5% or a soft alert if it exceeds 1%. This step catches simple arithmetic errors, unit mismatches (e.g., millions vs. billions), or incorrect share count updates before they propagate downstream.

Step 2: Cross-Source Consensus Validation

The cross_validate function (lines 67-104) implements multi-source verification by gathering the same data point from several independent providers—such as the company’s annual report (年报), Yahoo Finance, and StockAnalysis. The implementation:

  • Converts each value to an exact Decimal type to prevent floating-point drift (implemented in lines 31-38)
  • Computes the median of all supplied values as the consensus reference
  • Measures each source’s deviation from this median
  • Flags any source exceeding the default 2% tolerance threshold while printing the consensus value for downstream calculations

By requiring multiple independent confirmations and applying tight tolerances, the toolkit reduces reliance on any single data provider and surfaces discrepancies stemming from currency conversion errors, reporting delays, or data entry mistakes.

Implementation Examples

You can invoke these verification steps directly from the command line. First, validate the raw calculation integrity:

python3 tools/financial_rigor.py verify-market-cap \
    --price 510 --shares 9.11e9 --reported 4.65e12 --currency HKD

Then cross-validate the market cap across multiple independent sources:

python3 tools/financial_rigor.py cross-validate \
    --field "Market Cap" \
    --values '{"年报": 4.62e12, "Yahoo": 4.65e12, "StockAnalysis": 4.68e12}' \
    --unit "HKD" --tolerance 2.0

The second command outputs a consensus report similar to:


✅ Yahoo               : 4.65e12 HKD  (偏差 0.64%)
✅ StockAnalysis       : 4.68e12 HKD  (偏差 1.30%)
✅ 年报                : 4.62e12 HKD  (偏差 0.64%)
✅ 所有来源偏差 ≤ 2.0%, 数据一致
共识值 (加权中位数): 4.65e12 HKD

If any source deviates beyond the configured tolerance, the tool emits a warning and suggests preferring the company’s official filing or exchange data over third-party aggregators.

Integration with Research Workflows

According to the source code in skills/investment-research.md, research agents automatically invoke this toolkit at critical validation checkpoints during the investment research process. Additionally, tools/report_audit.py orchestrates comprehensive audit runs that include both market cap verification and cross-source checks as part of standardized report generation.

This integration ensures that every critical figure entering the valuation pipeline undergoes both numeric verification (price × shares vs. reported cap) and consistency reporting (identifying which sources agree and which require manual review).

Summary

  • Two-step validation: The verify_market_cap function checks arithmetic integrity (5% hard threshold, 1% soft threshold), while cross_validate ensures multi-source consensus (2% tolerance).
  • Precision handling: All calculations use Decimal types (lines 31-38 of financial_rigor.py) to eliminate floating-point drift.
  • Median consensus: The system uses median values across sources rather than averages to resist outlier influence.
  • Automated integration: Research skills in skills/investment-research.md and audit tools in report_audit.py embed these checks into standard workflows.

Frequently Asked Questions

What triggers a hard error versus a soft alert in the market cap verification?

The verify_market_cap function raises a hard error when the calculated market cap (share price × shares outstanding) deviates from the reported value by more than 5%, indicating potential unit mismatches or data corruption. It issues a soft alert for deviations exceeding 1%, suggesting the discrepancy warrants investigation but may not block downstream processing.

Why does the toolkit use Decimal instead of float for financial calculations?

As implemented in lines 31-38 of tools/financial_rigor.py, the cross_validate function converts all input values to Python’s Decimal type before comparison. This prevents floating-point representation errors that accumulate in standard binary floating-point arithmetic, ensuring that deviation percentages and median calculations remain exact to the last significant digit.

How does the system handle disagreements between official filings and third-party data providers?

The cross_validate function treats the median of all sources as the consensus truth rather than prioritizing any single provider. If the company’s annual report (年报) disagrees with Yahoo Finance or StockAnalysis beyond the 2% tolerance, the tool flags the outlier and prints the median value for downstream use. The researcher retains discretion to overweight official filings, but the system quantifies exactly how much each source diverges from the consensus.

Which files orchestrate the complete multi-source verification workflow?

The primary implementation resides in tools/financial_rigor.py, which contains both verify_market_cap and cross_validate. The skills/investment-research.md file defines when research agents invoke these checks, while tools/report_audit.py executes full audit runs that bundle market cap verification with other financial rigor tests as part of comprehensive report validation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →