How to Use the financial-data Skill for Data Validation in AI Berkshire

The financial-data skill enforces a mandatory two-source validation rule that requires every financial metric to be verified against independent providers, automatically flagging discrepancies exceeding 1%.

The AI Berkshire project implements rigorous data quality controls through its financial-data skill, a modular component designed to eliminate single-source errors in financial analysis workflows. This skill orchestrates dual-provider verification across multiple markets, ensuring that downstream AI-generated reports rely on vetted, accurate figures. Mastering the financial-data skill for data validation is essential for maintaining the project's research quality standards in algorithmic investment analysis.

Skill Architecture and Data Flow

The validation system operates through three distinct layers that work together to guarantee data integrity before any financial metric reaches an analysis report.

Skill Specification Layer

The canonical workflow definition resides in skills/financial-data.md. This markdown file specifies the required source pairs for each geographic market:

  • U.S. stocks: Macrotrends + StockAnalysis
  • Hong Kong: Aastocks + Macrotrends
  • Mainland China: East Money + Giant Shares
  • Taiwan: FinMind + Goodinfo

The specification mandates that every key metric must originate from at least two independent providers, with any variance larger than 1% triggering an automatic alert.

Tool Implementation Layer

Low-level data fetchers live under the tools/ directory, providing concrete implementations for each provider:

Runtime Enforcement

Downstream skills invoke the financial-data validation through explicit references. Files such as skills/investment-research.md and codex-skills/income-investment/SKILL.md contain comments directing the system to "Step 2 — 对清单每项从可靠信源取数 (参见 skills/financial-data.md)". This ensures that whenever a skill needs a figure, the runtime automatically orchestrates the two-provider check, computes relative differences, and inserts warning flags when deviations exceed the 1% threshold.

The Two-Source Validation Workflow

The financial-data skill executes a six-step validation pipeline for every requested metric:

  1. Resolve target — Identify the ticker symbol and corresponding market region from skill input
  2. Primary query — Fetch the metric from the first designated provider (e.g., Macrotrends for U.S. equities)
  3. Secondary query — Retrieve the identical metric from the second provider (e.g., StockAnalysis)
  4. Difference calculation — Compute absolute and relative percentage differences between the two values
  5. Threshold check — If the relative difference exceeds 1%, insert a "⚠️ Discrepancy > 1%" flag into the generated report
  6. Reconciliation — Return the reconciled value (typically the average or preferred source) for downstream use

This workflow satisfies the project's "Research Quality Rules" as defined in AGENTS.md, ensuring AI-generated analyses rest on a dual-source data foundation.

Code Implementation Examples

Direct Python Integration

You can invoke the validation logic directly by combining the tool implementations:


# utils/financial.py (illustrative helper)

from tools.twstock_data import fetch_twstock
from tools.morningstar_fair_value import fetch_us_stock

def get_validated_metric(ticker: str, market: str, metric: str):
    """Retrieve `metric` for `ticker` from two independent sources and validate."""
    if market == "US":
        src1 = fetch_us_stock(ticker, provider="macrotrends")[metric]
        src2 = fetch_us_stock(ticker, provider="stockanalysis")[metric]
    elif market == "TW":
        src1 = fetch_twstock(ticker, provider="finmind")[metric]
        src2 = fetch_twstock(ticker, provider="goodinfo")[metric]
    else:
        raise ValueError("Unsupported market")

    # Cross-validation (1% threshold)

    diff = abs(src1 - src2) / ((src1 + src2) / 2)
    if diff > 0.01:
        print(f"⚠️ Discrepancy >1% for {ticker}:{metric} ({src1:.2f} vs {src2:.2f})")
    # Return the average as the reconciled value

    return (src1 + src2) / 2

Claude/Codex Prompt Integration

When using the AI Berkshire system through Claude or Codex, reference the skill explicitly:

You are AI Berkshire.  
Please evaluate the P/E ratio of AAPL.  
Apply the `financial-data` skill (see /skills/financial-data.md) to fetch the value from Macrotrends and StockAnalysis, cross-validate, and flag any deviation >1%.

The runtime loads the skill specification from codex-prompts/financial-data.md and injects the validation logic automatically.

Command-Line Execution

Execute a full research checklist that includes validation:


# From the repository root

python3 scripts/sync-codex-skills.py   # ensure the skill is up-to-date

ai-berkshire run investment-checklist --ticker TSLA

The investment-checklist skill internally references financial-data (as noted in codex-skills/investment-checklist/SKILL.md), so the generated output contains validated figures with discrepancy warnings where applicable.

Key Files in the Validation Pipeline

Summary

  • The financial-data skill mandates dual-source verification for all financial metrics in the AI Berkshire ecosystem.
  • A 1% discrepancy threshold triggers automatic alerts when independent providers diverge beyond acceptable tolerances.
  • Market-specific provider pairs are defined in skills/financial-data.md, with implementations in tools/twstock_data.py and related fetchers.
  • Downstream skills like investment-research and income-investment invoke the validation layer automatically through standardized references.
  • The system supports direct Python usage, Codex prompt integration, and command-line execution patterns.

Frequently Asked Questions

What triggers the discrepancy warning in the financial-data skill?

The skill calculates the relative percentage difference between values from two independent providers. When this difference exceeds 1%, the system inserts a "⚠️ Discrepancy > 1%" warning flag into the analysis output. This threshold is hardcoded in the validation logic referenced across skills like investment-research.md and income-investment/SKILL.md.

Which data providers does the financial-data skill use for each market?

According to skills/financial-data.md, the skill maps specific provider pairs to geographic markets: U.S. stocks use Macrotrends and StockAnalysis; Hong Kong equities use Aastocks and Macrotrends; Mainland China uses East Money and Giant Shares; and Taiwan relies on FinMind and Goodinfo. These pairings ensure independent data sources for cross-validation.

How do other skills invoke the financial-data validation layer?

Skills reference the validation workflow through standardized comments such as "Step 2 — 对清单每项从可靠信源取数 (参见 skills/financial-data.md)" found in files like skills/investment-research.md. This convention allows the AI Berkshire runtime to automatically inject the two-source check whenever financial data is requested, without requiring repetitive validation code in each skill.

Can I modify the 1% validation threshold or add new markets?

The threshold and provider mappings are defined in the skill specification at skills/financial-data.md. While the core validation logic expects the 1% standard for consistency with the project's Research Quality Rules, you can extend support for additional markets by updating the provider matrix in the skill definition and implementing corresponding fetchers in the tools/ directory (following patterns established in tools/twstock_data.py).

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →