How Stock Data Is Fetched in Daily Stock Analysis: Architecture and Implementation

The daily_stock_analysis repository implements a resilient strategy-pattern architecture where DataFetcherManager orchestrates multiple provider-specific fetchers (Yahoo Finance, Tushare, Akshare, E-Finance) with automatic failover, routing US/HK stocks to specialized sources while normalizing all data to a standard schema.

The ZhuLinsen/daily_stock_analysis project solves data reliability challenges in quantitative finance by abstracting multiple third-party APIs behind a unified interface. Understanding how stock data is fetched requires examining the three-layer abstraction in data_provider/base.py and its concrete implementations, which handle everything from rate limiting to cross-market symbol normalization.

The Strategy-Pattern Data Fetching Stack

The repository organizes data acquisition into distinct layers that separate interface definition from provider-specific implementation:

  • Fetcher Base (data_provider/base.py): Defines the abstract BaseFetcher class with standardized methods like _fetch_raw_data, _normalize_data, get_daily_data, and get_realtime_quote. This layer handles cross-cutting concerns including rate-limit management, exception summarization, and column standardization to the STANDARD_COLUMNS set (date, open, high, low, close, volume, amount, pct_chg).

  • Concrete Fetchers: Seven provider-specific implementations handle raw data acquisition and market-specific quirks:

    • YfinanceFetcher (Yahoo Finance fallback and US/HK indices)
    • TushareFetcher (Tushare Pro API for high-quality Chinese data)
    • AkshareFetcher (Multi-portal Chinese data via Akshare library)
    • EfinanceFetcher (East Finance/E-Finance primary source)
    • PytdxFetcher (通达信/PyTDX historical data)
    • BaostockFetcher (Baostock API fallback)
    • LongbridgeFetcher (Longbridge API for US/HK real-time quotes)
  • Fetcher Manager (DataFetcherManager): Maintains priority-ordered fetcher lists, implements automatic failover logic, and exposes the primary façade used by downstream analysis modules. This singleton orchestrator lives in data_provider/base.py and is instantiated in main.py and service layers.

Step-by-Step Data Fetching Workflow

When the system requests daily K-line data for a stock like 600519 (Kweichow Moutai), DataFetcherManager.get_daily_data executes a seven-stage pipeline:

1. Symbol Normalization

The manager first calls normalize_stock_code to strip exchange prefixes (SH, SZ, BJ) and append market-appropriate suffixes. For US markets, tickers pass through unchanged, while HK stocks receive suffix handling via _is_hk_market checks. Market-specific helpers like YfinanceFetcher._convert_stock_code or TushareFetcher._convert_stock_code handle provider-specific formatting requirements.

2. Fetcher Selection

The manager applies market-aware routing logic:

  • US Stocks: Routes directly to YfinanceFetcher or LongbridgeFetcher (preferred when API credentials configure longbridge source priority)
  • HK Stocks: Uses dual-source routing between LongbridgeFetcher and AkshareFetcher
  • CN A-Shares: Iterates through a default priority chain: EFinance → Akshare → PyTDX → Tushare → Baostock → Yfinance → Longbridge

When a Tushare token is present in configuration, that source receives priority elevation (priority -1) in the ordered list.

3. Invocation and Serialization

DataFetcherManager._call_fetcher_method serializes calls per fetcher instance to prevent rate-limit violations, then invokes the selected fetcher's get_daily_data method.

4. Raw Data Acquisition

Each fetcher implements _fetch_raw_data using provider-specific clients:

  • YfinanceFetcher calls yfinance.download()
  • TushareFetcher invokes the Tushare Pro HTTP client
  • AkshareFetcher delegates to the Akshare library functions

5. Schema Normalization

Fetchers transform provider-specific columns into the common schema via _normalize_data. This mapping ensures date, open, high, low, close, volume, amount, and pct_chg fields exist consistently, calculating derived metrics like percentage change where sources omit them.

6. Data Cleaning and Enrichment

BaseFetcher._clean_data coerces data types, drops NaN values, and sorts chronologically. The _calculate_indicators method then appends technical analysis fields including SMA-5/10/20 and volume-ratio calculations.

7. Result Return

The manager returns a tuple containing the final DataFrame and the string name of the successful fetcher source. If all sources fail, the system raises DataFetchError after recording detailed error summaries via summarize_exception.

Market-Specific Routing Logic

The fetcher selection algorithm distinguishes between historical daily data and real-time quote requirements.

CN Market Priorities

For mainland Chinese securities, the default priority prioritizes free sources:

  1. E-Finance (efinance_fetcher.py) - Highest priority when no Tushare token exists
  2. Akshare - Multi-source aggregator for Sina, Tencent, and East Money portals
  3. PyTDX - 通达信 data for historical backfill
  4. Tushare - Elevated to first position when TUSHARE_TOKEN is configured
  5. Baostock - Secondary fallback
  6. Yahoo Finance - Tertiary fallback for delisted or thinly traded symbols

US and HK Market Handling

Cross-border securities trigger specialized routing in get_daily_data:

  • Longbridge API serves as primary when longbridge appears in configuration priorities and valid API keys exist
  • Yahoo Finance provides universal fallback coverage for US equities and indices
  • Akshare supplies HK real-time data when Longbridge is unavailable

The us_index_mapping.py file maintains symbol translations (e.g., SPX → ^GSPC) required by Yahoo Finance for major indices.

Real-Time Quote Handling

Real-time data follows a parallel but distinct failover chain through get_realtime_quote:

US/HK Markets:

  • Primary: LongbridgeFetcher (if configured with valid credentials)
  • Fallback: YfinanceFetcher for US; AkshareFetcher for HK

CN Markets: The manager consults realtime_source_priority configuration (typically efinance, akshare_em, akshare_sina, tushare), returning a UnifiedRealtimeQuote object standardizing fields like price, change_pct, volume, amount, and pe_ratio.

Missing fields in partial responses are supplemented by subsequent lower-priority sources before final return.

Practical Implementation Examples

Fetching 30 Days of Historical Data

from data_provider.base import DataFetcherManager

manager = DataFetcherManager()
df, source = manager.get_daily_data('600519')  # Kweichow Moutai

print(f"Fetched {len(df)} rows from {source}")
print(df.head())

The manager automatically selects E-Finance by default (or Tushare if token is present) and returns normalized columns.

Retrieving US Stock Real-Time Quotes

manager = DataFetcherManager()
quote = manager.get_realtime_quote('AAPL')
if quote:
    print(f"AAPL price={quote.price:.2f}, change={quote.change_pct}%")

Routes to Longbridge if API keys exist; otherwise falls back to Yahoo Finance delayed data.

Accessing Main Indices

manager = DataFetcherManager()
indices = manager.get_main_indices(region='cn')
for idx in indices:
    print(f"{idx['code']}: {idx['current']} ({idx['change_pct']}%)")

Delegates to YfinanceFetcher.get_main_indices, which maps internal codes to Yahoo symbols like 000001 → ^SSEC.

Direct Low-Level Fetcher Usage

from data_provider.yfinance_fetcher import YfinanceFetcher

yf = YfinanceFetcher()
df = yf.get_daily_data('000001')  # Shanghai Composite via Yahoo

df = yf._normalize_data(df, '000001')
print(df[['date', 'close']].tail())

Bypasses manager failover but preserves normalization logic for custom analysis pipelines.

Summary

  • Strategy Pattern Architecture: BaseFetcher defines the contract while seven concrete implementations handle provider-specific logic in files like yfinance_fetcher.py and tushare_fetcher.py.
  • Intelligent Routing: DataFetcherManager automatically routes US stocks to Yahoo/Longbridge, HK stocks to Longbridge/Akshare, and CN stocks through a priority chain starting with E-Finance or Tushare (when tokenized).
  • Automatic Failover: The system attempts fetchers sequentially, catching exceptions via summarize_exception until successful or the list exhausts.
  • Schema Unification: All sources normalize to standard columns (date, open, high, low, close, volume, amount, pct_chg) through _normalize_data methods.
  • Real-Time Support: Separate priority chains in get_realtime_quote handle intraday quotes with UnifiedRealtimeQuote normalization.

Frequently Asked Questions

How does the system handle rate limits from data providers?

Each concrete fetcher implements rate-limit handling within _fetch_raw_data according to provider specifications. TushareFetcher includes specific Tushare Pro API rate-limit logic, while DataFetcherManager._call_fetcher_method serializes calls per fetcher instance to prevent concurrent request flooding. Failed requests trigger automatic failover to the next priority source rather than waiting for rate-limit resets.

What happens if all data sources fail for a symbol?

If the ordered priority list exhausts without successful data retrieval, DataFetcherManager raises a DataFetchError containing summarized exception details from all attempted sources. The error object generated via summarize_exception includes the specific failure reasons (network timeouts, authentication errors, or missing symbols) for debugging.

Can I force a specific data source instead of using automatic selection?

Yes. While DataFetcherManager provides automatic routing, you can instantiate concrete fetchers directly from their respective modules (e.g., from data_provider.tushare_fetcher import TushareFetcher). Direct instantiation bypasses the manager's failover chain but retains the _normalize_data and _clean_data processing pipeline specific to that source.

Why are there different priority lists for daily data versus real-time quotes?

Daily historical data and real-time quotes have different reliability characteristics and API limitations. Daily data prioritizes free, high-volume historical sources (E-Finance, Akshare) while real-time data prioritizes low-latency sources (Longbridge, E-Finance realtime endpoints). The realtime_source_priority configuration allows independent tuning for latency-sensitive intraday trading versus backtesting workflows.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →