# How to Use the financial-data Skill for Data Validation in AI Berkshire

> Learn how to use the financial-data skill for data validation in AI Berkshire. Ensure accuracy with its two-source validation rule and automatic discrepancy flagging.

- Repository: [Xbt Lin/ai-berkshire](https://github.com/xbtlin/ai-berkshire)
- Tags: how-to-guide
- Published: 2026-07-25

---

**The financial-data skill enforces a mandatory two-source validation rule that requires every financial metric to be verified against independent providers, automatically flagging discrepancies exceeding 1%.**

The AI Berkshire project implements rigorous data quality controls through its `financial-data` skill, a modular component designed to eliminate single-source errors in financial analysis workflows. This skill orchestrates dual-provider verification across multiple markets, ensuring that downstream AI-generated reports rely on vetted, accurate figures. Mastering the financial-data skill for data validation is essential for maintaining the project's research quality standards in algorithmic investment analysis.

## Skill Architecture and Data Flow

The validation system operates through three distinct layers that work together to guarantee data integrity before any financial metric reaches an analysis report.

### Skill Specification Layer

The canonical workflow definition resides in [`skills/financial-data.md`](https://github.com/xbtlin/ai-berkshire/blob/main/skills/financial-data.md). This markdown file specifies the required source pairs for each geographic market:

- **U.S. stocks**: Macrotrends + StockAnalysis
- **Hong Kong**: Aastocks + Macrotrends
- **Mainland China**: East Money + Giant Shares
- **Taiwan**: FinMind + Goodinfo

The specification mandates that every key metric must originate from at least two independent providers, with any variance larger than 1% triggering an automatic alert.

### Tool Implementation Layer

Low-level data fetchers live under the `tools/` directory, providing concrete implementations for each provider:

- [`tools/twstock_data.py`](https://github.com/xbtlin/ai-berkshire/blob/main/tools/twstock_data.py) — Accesses FinMind and Goodinfo for Taiwanese equities
- [`tools/morningstar_fair_value.py`](https://github.com/xbtlin/ai-berkshire/blob/main/tools/morningstar_fair_value.py) — Implements Macrotrends scraping for U.S. stocks
- [`tools/ashare_data.py`](https://github.com/xbtlin/ai-berkshire/blob/main/tools/ashare_data.py) — Handles East Money and Giant Shares for A-shares
- [`tools/stock_screener.py`](https://github.com/xbtlin/ai-berkshire/blob/main/tools/stock_screener.py) — Additional market-specific data aggregation

### Runtime Enforcement

Downstream skills invoke the financial-data validation through explicit references. Files such as [`skills/investment-research.md`](https://github.com/xbtlin/ai-berkshire/blob/main/skills/investment-research.md) and [`codex-skills/income-investment/SKILL.md`](https://github.com/xbtlin/ai-berkshire/blob/main/codex-skills/income-investment/SKILL.md) contain comments directing the system to "Step 2 — 对清单每项从可靠信源取数 (参见 skills/financial-data.md)". This ensures that whenever a skill needs a figure, the runtime automatically orchestrates the two-provider check, computes relative differences, and inserts warning flags when deviations exceed the 1% threshold.

## The Two-Source Validation Workflow

The financial-data skill executes a six-step validation pipeline for every requested metric:

1. **Resolve target** — Identify the ticker symbol and corresponding market region from skill input
2. **Primary query** — Fetch the metric from the first designated provider (e.g., Macrotrends for U.S. equities)
3. **Secondary query** — Retrieve the identical metric from the second provider (e.g., StockAnalysis)
4. **Difference calculation** — Compute absolute and relative percentage differences between the two values
5. **Threshold check** — If the relative difference exceeds 1%, insert a "⚠️ Discrepancy > 1%" flag into the generated report
6. **Reconciliation** — Return the reconciled value (typically the average or preferred source) for downstream use

This workflow satisfies the project's "Research Quality Rules" as defined in [`AGENTS.md`](https://github.com/xbtlin/ai-berkshire/blob/main/AGENTS.md), ensuring AI-generated analyses rest on a dual-source data foundation.

## Code Implementation Examples

### Direct Python Integration

You can invoke the validation logic directly by combining the tool implementations:

```python

# utils/financial.py (illustrative helper)

from tools.twstock_data import fetch_twstock
from tools.morningstar_fair_value import fetch_us_stock

def get_validated_metric(ticker: str, market: str, metric: str):
    """Retrieve `metric` for `ticker` from two independent sources and validate."""
    if market == "US":
        src1 = fetch_us_stock(ticker, provider="macrotrends")[metric]
        src2 = fetch_us_stock(ticker, provider="stockanalysis")[metric]
    elif market == "TW":
        src1 = fetch_twstock(ticker, provider="finmind")[metric]
        src2 = fetch_twstock(ticker, provider="goodinfo")[metric]
    else:
        raise ValueError("Unsupported market")

    # Cross-validation (1% threshold)

    diff = abs(src1 - src2) / ((src1 + src2) / 2)
    if diff > 0.01:
        print(f"⚠️ Discrepancy >1% for {ticker}:{metric} ({src1:.2f} vs {src2:.2f})")
    # Return the average as the reconciled value

    return (src1 + src2) / 2

```

### Claude/Codex Prompt Integration

When using the AI Berkshire system through Claude or Codex, reference the skill explicitly:

```text
You are AI Berkshire.  
Please evaluate the P/E ratio of AAPL.  
Apply the `financial-data` skill (see /skills/financial-data.md) to fetch the value from Macrotrends and StockAnalysis, cross-validate, and flag any deviation >1%.

```

The runtime loads the skill specification from [`codex-prompts/financial-data.md`](https://github.com/xbtlin/ai-berkshire/blob/main/codex-prompts/financial-data.md) and injects the validation logic automatically.

### Command-Line Execution

Execute a full research checklist that includes validation:

```bash

# From the repository root

python3 scripts/sync-codex-skills.py   # ensure the skill is up-to-date

ai-berkshire run investment-checklist --ticker TSLA

```

The `investment-checklist` skill internally references `financial-data` (as noted in [`codex-skills/investment-checklist/SKILL.md`](https://github.com/xbtlin/ai-berkshire/blob/main/codex-skills/investment-checklist/SKILL.md)), so the generated output contains validated figures with discrepancy warnings where applicable.

## Key Files in the Validation Pipeline

- [`skills/financial-data.md`](https://github.com/xbtlin/ai-berkshire/blob/main/skills/financial-data.md) — Canonical specification of the validation workflow and source matrix
- [`tools/twstock_data.py`](https://github.com/xbtlin/ai-berkshire/blob/main/tools/twstock_data.py) — Taiwanese data providers (FinMind & Goodinfo)
- [`tools/morningstar_fair_value.py`](https://github.com/xbtlin/ai-berkshire/blob/main/tools/morningstar_fair_value.py) — U.S. stock data scraper for Macrotrends
- [`codex-prompts/financial-data.md`](https://github.com/xbtlin/ai-berkshire/blob/main/codex-prompts/financial-data.md) — Guidance for loading the skill in Codex-generated prompts
- [`skills/investment-research.md`](https://github.com/xbtlin/ai-berkshire/blob/main/skills/investment-research.md) — Demonstrates cross-skill invocation patterns
- [`codex-skills/income-investment/SKILL.md`](https://github.com/xbtlin/ai-berkshire/blob/main/codex-skills/income-investment/SKILL.md) — Production example of validation integration

## Summary

- The **financial-data skill** mandates dual-source verification for all financial metrics in the AI Berkshire ecosystem.
- A **1% discrepancy threshold** triggers automatic alerts when independent providers diverge beyond acceptable tolerances.
- Market-specific provider pairs are defined in [`skills/financial-data.md`](https://github.com/xbtlin/ai-berkshire/blob/main/skills/financial-data.md), with implementations in [`tools/twstock_data.py`](https://github.com/xbtlin/ai-berkshire/blob/main/tools/twstock_data.py) and related fetchers.
- Downstream skills like `investment-research` and `income-investment` invoke the validation layer automatically through standardized references.
- The system supports direct Python usage, Codex prompt integration, and command-line execution patterns.

## Frequently Asked Questions

### What triggers the discrepancy warning in the financial-data skill?

The skill calculates the relative percentage difference between values from two independent providers. When this difference exceeds **1%**, the system inserts a "⚠️ Discrepancy > 1%" warning flag into the analysis output. This threshold is hardcoded in the validation logic referenced across skills like [`investment-research.md`](https://github.com/xbtlin/ai-berkshire/blob/main/investment-research.md) and [`income-investment/SKILL.md`](https://github.com/xbtlin/ai-berkshire/blob/main/income-investment/SKILL.md).

### Which data providers does the financial-data skill use for each market?

According to [`skills/financial-data.md`](https://github.com/xbtlin/ai-berkshire/blob/main/skills/financial-data.md), the skill maps specific provider pairs to geographic markets: **U.S. stocks** use Macrotrends and StockAnalysis; **Hong Kong** equities use Aastocks and Macrotrends; **Mainland China** uses East Money and Giant Shares; and **Taiwan** relies on FinMind and Goodinfo. These pairings ensure independent data sources for cross-validation.

### How do other skills invoke the financial-data validation layer?

Skills reference the validation workflow through standardized comments such as "Step 2 — 对清单每项从可靠信源取数 (参见 skills/financial-data.md)" found in files like [`skills/investment-research.md`](https://github.com/xbtlin/ai-berkshire/blob/main/skills/investment-research.md). This convention allows the AI Berkshire runtime to automatically inject the two-source check whenever financial data is requested, without requiring repetitive validation code in each skill.

### Can I modify the 1% validation threshold or add new markets?

The threshold and provider mappings are defined in the skill specification at [`skills/financial-data.md`](https://github.com/xbtlin/ai-berkshire/blob/main/skills/financial-data.md). While the core validation logic expects the 1% standard for consistency with the project's Research Quality Rules, you can extend support for additional markets by updating the provider matrix in the skill definition and implementing corresponding fetchers in the `tools/` directory (following patterns established in [`tools/twstock_data.py`](https://github.com/xbtlin/ai-berkshire/blob/main/tools/twstock_data.py)).