# How to Cross-Validate Financial Data from Multiple Sources Using the AI‑Berkshire Project

> Cross-validate financial data from multiple sources with the AI-Berkshire project. This toolkit calculates median consensus and flags deviations to ensure data accuracy.

- Repository: [Xbt Lin/ai-berkshire](https://github.com/xbtlin/ai-berkshire)
- Tags: how-to-guide
- Published: 2026-07-25

---

**The AI‑Berkshire repository provides a Financial Rigor Toolkit that validates numeric financial metrics across independent sources by calculating median consensus and flagging deviations that exceed a configurable tolerance threshold.**

The AI‑Berkshire project includes a specialized validation framework designed to ensure data integrity when aggregating financial metrics from disparate providers. Located in the [`tools/financial_rigor.py`](https://github.com/xbtlin/ai-berkshire/blob/main/tools/financial_rigor.py) module, the toolkit offers a robust method to cross-validate financial data from multiple sources, preventing costly errors caused by inconsistent reporting periods or currency mismatches.

## The Cross-Validation Architecture

### Precision-First Arithmetic

The implementation in [`tools/financial_rigor.py`](https://github.com/xbtlin/ai-berkshire/blob/main/tools/financial_rigor.py) converts all incoming values to Python `Decimal` objects with a fixed 28-digit precision context. This eliminates floating-point drift that commonly corrupts financial calculations involving high-precision decimals or currency conversions.

### Consensus Calculation Logic

The **`cross_validate`** function computes the median of all provided source values to establish a reference "consensus" point. It then calculates each source's percent deviation from this median, flagging any data point that exceeds the configurable tolerance threshold (default **2%**).

## Implementation Methods

### Command-Line Interface

The toolkit operates as a self-contained CLI, accessible directly from the repository root. The `cross_validate` routine is exposed through the `cross-validate` subcommand.

```python

# Example: CLI usage from repository root

python3 tools/financial_rigor.py cross-validate \
    --field revenue \
    --values '{"AnnualReport": 7518, "Yahoo": 7500, "StockAnalysis": 7520}' \
    --unit "億" \
    --tolerance 2.5

```

### Programmatic Integration

For Jupyter notebooks or automated pipelines, import the module directly and invoke the function with a dictionary mapping source names to numeric values.

```python
from tools.financial_rigor import cross_validate

# Collect values from independent providers

source_values = {
    "AnnualReport": 7518,      # Company 2023 filing (单位: 億)

    "YahooFinance": 7500,      # Yahoo Finance API

    "StockAnalysis": 7520,     # StockAnalysis.com aggregation

}

# Execute validation with 2% tolerance

result = cross_validate(
    field_name="Revenue",
    source_values=source_values,
    unit="億",
    tolerance_pct=2.0
)

print("Consensus:", result["consensus"])
print("Consistent:", result["all_consistent"])

```

## Output Interpretation and Discrepancy Handling

The function returns a dictionary containing the consensus value and a Boolean **`all_consistent`** flag, while simultaneously printing a formatted table to stdout. Sources within tolerance display with ✅, while outliers exceeding the threshold show ❌.

When discrepancies occur, the tool recommends prioritizing primary sources such as company annual reports or exchange-registered filings over aggregated third-party estimates. Common root causes include different fiscal year definitions, currency conversion timestamps, or non-GAAP adjustments applied selectively by certain providers.

## Project Structure and Supporting Files

According to the AI‑Berkshire source code, the validation ecosystem includes:

- **[`tools/financial_rigor.py`](https://github.com/xbtlin/ai-berkshire/blob/main/tools/financial_rigor.py)**: Core implementation of `cross_validate` and CLI entry point
- **[`scripts/sync-codex-skills.py`](https://github.com/xbtlin/ai-berkshire/blob/main/scripts/sync-codex-skills.py)**: Maintains synchronization between CLI updates and Codex-compatible wrappers
- **[`AGENTS.md`](https://github.com/xbtlin/ai-berkshire/blob/main/AGENTS.md)**: Documents the overall architecture and validation workflow integration

## Summary

- **Use [`tools/financial_rigor.py`](https://github.com/xbtlin/ai-berkshire/blob/main/tools/financial_rigor.py)** to access the `cross_validate` function for any numeric financial metric.
- **Supply a dictionary** mapping source names to values, along with an optional unit string and tolerance percentage (default 2%).
- **Leverage Decimal arithmetic** implemented in the toolkit to avoid floating-point errors during consensus calculation.
- **Interpret results** via the returned consensus value and `all_consistent` Boolean, augmented by the printed deviation table.
- **Prioritize primary sources** when resolving flagged discrepancies identified by the validation routine.

## Frequently Asked Questions

### What tolerance percentage should I use for volatile metrics like cryptocurrency market cap?

For highly volatile assets, increase the `tolerance_pct` parameter to 5% or higher when calling `cross_validate`. The default 2% suits stable metrics like audited revenue, but rapidly fluctuating values require wider bands to avoid false positives while still catching significant data errors.

### Can the toolkit handle currency conversion or different units automatically?

No, the `cross_validate` function assumes all input values share the same unit and currency. You must normalize currencies through external conversion before building the `source_values` dictionary. The `unit` parameter is strictly for display purposes in the output table.

### How does the algorithm handle an even number of sources?

When provided with an even number of data points, the median calculation in [`tools/financial_rigor.py`](https://github.com/xbtlin/ai-berkshire/blob/main/tools/financial_rigor.py) uses the standard statistical median (average of the two middle values). All deviations are then calculated against this computed consensus point, maintaining consistency regardless of source count.

### Is the CLI available as a pip-installable package?

Currently, the toolkit runs as a standalone script within the xbtlin/ai-berkshire repository. You must clone the repository and execute `python3 tools/financial_rigor.py` from the project root; there is no PyPI distribution, though the module can be imported directly into Python scripts within the project environment.