# How to Cross-Validate Financial Data from Multiple Sources with tools/financial_rigor.py

> Learn to cross-validate financial data from multiple sources using the financial_rigor.py script. Ensure data accuracy by identifying discrepancies with a configurable tolerance threshold.

- Repository: [Xbt Lin/ai-berkshire](https://github.com/xbtlin/ai-berkshire)
- Tags: how-to-guide
- Published: 2026-07-11

---

**The `cross_validate` function in [`tools/financial_rigor.py`](https://github.com/xbtlin/ai-berkshire/blob/main/tools/financial_rigor.py) computes a median-based consensus from multiple data providers and flags any source exceeding a configurable tolerance threshold.**

The `ai-berkshire` repository provides a lightweight, **standard-library-only** toolkit for investment research validation. The `cross_validate` function in [`tools/financial_rigor.py`](https://github.com/xbtlin/ai-berkshire/blob/main/tools/financial_rigor.py) offers a robust way to cross-validate financial data from multiple sources by comparing values against a calculated median and identifying outliers that deviate beyond specified tolerance limits.

## Understanding the Cross-Validation Architecture

The cross-validation system relies on three core components to ensure precision without floating-point drift. All logic resides in [`tools/financial_rigor.py`](https://github.com/xbtlin/ai-berkshire/blob/main/tools/financial_rigor.py), specifically within the `cross_validate` function (lines 167-204).

### High-Precision Decimal Arithmetic

All numeric processing uses a dedicated Decimal context (`_CTX`) and the `exact()` conversion function. This guarantees that financial ratios and large numbers (like market caps in billions) undergo arithmetic without floating-point precision errors, which is critical when comparing values from different data providers.

### Median-Based Consensus Calculation

The `cross_validate` function calculates the **median** of all provided values rather than the mean. This approach provides robust central tendency that tolerates occasional extreme values or outliers common in financial estimates from sources like annual reports, Bloomberg, or Yahoo Finance.

### Tolerance-Based Deviation Detection

Each source value is compared against the median consensus. Any deviation exceeding the configurable `tolerance_pct` (default **2%**) triggers a warning flag. The function returns a dictionary containing the consensus value and a boolean `all_consistent` indicator.

## CLI Usage for Cross-Validating Financial Data

The most common approach uses the command-line interface exposed through the `main` function with `argparse`. The `cross-validate` sub-command accepts JSON-formatted source values and optional unit specifications.

```bash
python3 tools/financial_rigor.py cross-validate \
    --field revenue \
    --values '{"年报": 7518, "Yahoo": 7500, "StockAnalysis": 7520}' \
    --unit 亿 \
    --tolerance 2.0

```

This command validates revenue figures from three sources, using "亿" (hundred millions) as the display unit. The output displays each source's deviation percentage from the median using the `fmt_number` formatter, shows visual checkmarks (✅) for valid sources, and indicates overall consistency.

## Programmatic Integration in Python

Import the `cross_validate` function directly to embed validation logic into data pipelines or research notebooks.

```python
from tools.financial_rigor import cross_validate

# Example data from different data providers

source_vals = {
    "AnnualReport": 7.518e9,
    "YahooFinance": 7.500e9,
    "Morningstar": 7.520e9,
}

result = cross_validate(
    field_name="revenue",
    source_values=source_vals,
    unit="亿",
    tolerance_pct=2.0
)

print(result)

# {'consensus': 7518000000.0, 'all_consistent': True}

```

The function converts all inputs using `exact()` to ensure precision, computes the median, and returns both the consensus value and consistency status.

## Handling Volatile Data with Custom Tolerance

For sectors with inherently noisy reporting or when comparing estimates across different fiscal calendars, increase the tolerance threshold to reduce false positives.

```bash
python3 tools/financial_rigor.py cross-validate \
    --field net_income \
    --values '{"Report": 1200, "Yahoo": 1150, "FactSet": 1300}' \
    --unit 亿 \
    --tolerance 5.0

```

Setting `--tolerance 5.0` allows deviations up to 5% from the median, accommodating legitimate variances in reporting methodologies while still catching significant data entry errors or currency conversion mistakes.

## Integration with AI-Berkshire Research Workflow

According to the repository's [`AGENTS.md`](https://github.com/xbtlin/ai-berkshire/blob/main/AGENTS.md) file, the CLI is automatically invoked by Claude Code Skills whenever a validation checkpoint is reached. This integration ensures that financial data extracted during AI-assisted research undergoes rigorous cross-validation before entering analysis workflows. Because [`financial_rigor.py`](https://github.com/xbtlin/ai-berkshire/blob/main/financial_rigor.py) uses only Python standard library features (available in Python ≥3.7), it runs in any environment without pip install requirements.

## Summary

- The `cross_validate` function in [`tools/financial_rigor.py`](https://github.com/xbtlin/ai-berkshire/blob/main/tools/financial_rigor.py) (lines 167-204) provides median-based consensus calculation for financial data from multiple sources.
- **Decimal precision** is enforced through the `_CTX` context and `exact()` conversion, eliminating floating-point arithmetic errors.
- The default **2% tolerance** identifies outliers while tolerating minor reporting differences, adjustable via `--tolerance` (CLI) or `tolerance_pct` (Python).
- The tool returns a dictionary with `consensus` and `all_consistent` keys, enabling automated validation checks in research pipelines.
- Zero external dependencies make this utility portable across any Python 3.7+ environment.

## Frequently Asked Questions

### What is the default tolerance threshold for cross-validation?

The default tolerance is **2%**, meaning any source value deviating more than 2% from the calculated median is flagged as inconsistent. You can adjust this threshold using the `--tolerance` flag in the CLI or the `tolerance_pct` parameter when calling `cross_validate()` programmatically.

### Why does the tool use median instead of mean for consensus calculation?

Financial estimates from different providers often contain outliers due to varying reporting standards or timing differences. The **median** offers a robust central tendency that tolerates occasional extreme values, making the validation step resilient to single-source errors that would skew a mean calculation.

### Can I use this tool without installing external dependencies?

Yes. The [`financial_rigor.py`](https://github.com/xbtlin/ai-berkshire/blob/main/financial_rigor.py) module is designed as a **standard-library-only** toolkit with no external dependencies beyond Python 3.7+. You can copy the file to any environment and run it immediately without a requirements.txt or pip install process.

### How does the tool handle different numerical units?

The `fmt_number` helper function converts large numbers into readable units like `亿` (hundred millions), `B` (billions), or `T` (trillions) for display purposes. When calling `cross_validate`, specify the unit via the `--unit` CLI flag or `unit` parameter; the underlying calculation always uses the raw Decimal values to maintain precision.