How Stock Data Is Fetched in Daily Stock Analysis: Architecture and Implementation
The daily_stock_analysis repository implements a resilient strategy-pattern architecture where DataFetcherManager orchestrates multiple provider-specific fetchers (Yahoo Finance, Tushare, Akshare, E-Finance) with automatic failover, routing US/HK stocks to specialized sources while normalizing all data to a standard schema.
The ZhuLinsen/daily_stock_analysis project solves data reliability challenges in quantitative finance by abstracting multiple third-party APIs behind a unified interface. Understanding how stock data is fetched requires examining the three-layer abstraction in data_provider/base.py and its concrete implementations, which handle everything from rate limiting to cross-market symbol normalization.
The Strategy-Pattern Data Fetching Stack
The repository organizes data acquisition into distinct layers that separate interface definition from provider-specific implementation:
-
Fetcher Base (
data_provider/base.py): Defines the abstractBaseFetcherclass with standardized methods like_fetch_raw_data,_normalize_data,get_daily_data, andget_realtime_quote. This layer handles cross-cutting concerns including rate-limit management, exception summarization, and column standardization to theSTANDARD_COLUMNSset (date,open,high,low,close,volume,amount,pct_chg). -
Concrete Fetchers: Seven provider-specific implementations handle raw data acquisition and market-specific quirks:
YfinanceFetcher(Yahoo Finance fallback and US/HK indices)TushareFetcher(Tushare Pro API for high-quality Chinese data)AkshareFetcher(Multi-portal Chinese data via Akshare library)EfinanceFetcher(East Finance/E-Finance primary source)PytdxFetcher(通达信/PyTDX historical data)BaostockFetcher(Baostock API fallback)LongbridgeFetcher(Longbridge API for US/HK real-time quotes)
-
Fetcher Manager (
DataFetcherManager): Maintains priority-ordered fetcher lists, implements automatic failover logic, and exposes the primary façade used by downstream analysis modules. This singleton orchestrator lives indata_provider/base.pyand is instantiated inmain.pyand service layers.
Step-by-Step Data Fetching Workflow
When the system requests daily K-line data for a stock like 600519 (Kweichow Moutai), DataFetcherManager.get_daily_data executes a seven-stage pipeline:
1. Symbol Normalization
The manager first calls normalize_stock_code to strip exchange prefixes (SH, SZ, BJ) and append market-appropriate suffixes. For US markets, tickers pass through unchanged, while HK stocks receive suffix handling via _is_hk_market checks. Market-specific helpers like YfinanceFetcher._convert_stock_code or TushareFetcher._convert_stock_code handle provider-specific formatting requirements.
2. Fetcher Selection
The manager applies market-aware routing logic:
- US Stocks: Routes directly to
YfinanceFetcherorLongbridgeFetcher(preferred when API credentials configurelongbridgesource priority) - HK Stocks: Uses dual-source routing between
LongbridgeFetcherandAkshareFetcher - CN A-Shares: Iterates through a default priority chain: EFinance → Akshare → PyTDX → Tushare → Baostock → Yfinance → Longbridge
When a Tushare token is present in configuration, that source receives priority elevation (priority -1) in the ordered list.
3. Invocation and Serialization
DataFetcherManager._call_fetcher_method serializes calls per fetcher instance to prevent rate-limit violations, then invokes the selected fetcher's get_daily_data method.
4. Raw Data Acquisition
Each fetcher implements _fetch_raw_data using provider-specific clients:
YfinanceFetchercallsyfinance.download()TushareFetcherinvokes the Tushare Pro HTTP clientAkshareFetcherdelegates to the Akshare library functions
5. Schema Normalization
Fetchers transform provider-specific columns into the common schema via _normalize_data. This mapping ensures date, open, high, low, close, volume, amount, and pct_chg fields exist consistently, calculating derived metrics like percentage change where sources omit them.
6. Data Cleaning and Enrichment
BaseFetcher._clean_data coerces data types, drops NaN values, and sorts chronologically. The _calculate_indicators method then appends technical analysis fields including SMA-5/10/20 and volume-ratio calculations.
7. Result Return
The manager returns a tuple containing the final DataFrame and the string name of the successful fetcher source. If all sources fail, the system raises DataFetchError after recording detailed error summaries via summarize_exception.
Market-Specific Routing Logic
The fetcher selection algorithm distinguishes between historical daily data and real-time quote requirements.
CN Market Priorities
For mainland Chinese securities, the default priority prioritizes free sources:
- E-Finance (
efinance_fetcher.py) - Highest priority when no Tushare token exists - Akshare - Multi-source aggregator for Sina, Tencent, and East Money portals
- PyTDX - 通达信 data for historical backfill
- Tushare - Elevated to first position when
TUSHARE_TOKENis configured - Baostock - Secondary fallback
- Yahoo Finance - Tertiary fallback for delisted or thinly traded symbols
US and HK Market Handling
Cross-border securities trigger specialized routing in get_daily_data:
- Longbridge API serves as primary when
longbridgeappears in configuration priorities and valid API keys exist - Yahoo Finance provides universal fallback coverage for US equities and indices
- Akshare supplies HK real-time data when Longbridge is unavailable
The us_index_mapping.py file maintains symbol translations (e.g., SPX → ^GSPC) required by Yahoo Finance for major indices.
Real-Time Quote Handling
Real-time data follows a parallel but distinct failover chain through get_realtime_quote:
US/HK Markets:
- Primary:
LongbridgeFetcher(if configured with valid credentials) - Fallback:
YfinanceFetcherfor US;AkshareFetcherfor HK
CN Markets:
The manager consults realtime_source_priority configuration (typically efinance, akshare_em, akshare_sina, tushare), returning a UnifiedRealtimeQuote object standardizing fields like price, change_pct, volume, amount, and pe_ratio.
Missing fields in partial responses are supplemented by subsequent lower-priority sources before final return.
Practical Implementation Examples
Fetching 30 Days of Historical Data
from data_provider.base import DataFetcherManager
manager = DataFetcherManager()
df, source = manager.get_daily_data('600519') # Kweichow Moutai
print(f"Fetched {len(df)} rows from {source}")
print(df.head())
The manager automatically selects E-Finance by default (or Tushare if token is present) and returns normalized columns.
Retrieving US Stock Real-Time Quotes
manager = DataFetcherManager()
quote = manager.get_realtime_quote('AAPL')
if quote:
print(f"AAPL price={quote.price:.2f}, change={quote.change_pct}%")
Routes to Longbridge if API keys exist; otherwise falls back to Yahoo Finance delayed data.
Accessing Main Indices
manager = DataFetcherManager()
indices = manager.get_main_indices(region='cn')
for idx in indices:
print(f"{idx['code']}: {idx['current']} ({idx['change_pct']}%)")
Delegates to YfinanceFetcher.get_main_indices, which maps internal codes to Yahoo symbols like 000001 → ^SSEC.
Direct Low-Level Fetcher Usage
from data_provider.yfinance_fetcher import YfinanceFetcher
yf = YfinanceFetcher()
df = yf.get_daily_data('000001') # Shanghai Composite via Yahoo
df = yf._normalize_data(df, '000001')
print(df[['date', 'close']].tail())
Bypasses manager failover but preserves normalization logic for custom analysis pipelines.
Summary
- Strategy Pattern Architecture:
BaseFetcherdefines the contract while seven concrete implementations handle provider-specific logic in files likeyfinance_fetcher.pyandtushare_fetcher.py. - Intelligent Routing:
DataFetcherManagerautomatically routes US stocks to Yahoo/Longbridge, HK stocks to Longbridge/Akshare, and CN stocks through a priority chain starting with E-Finance or Tushare (when tokenized). - Automatic Failover: The system attempts fetchers sequentially, catching exceptions via
summarize_exceptionuntil successful or the list exhausts. - Schema Unification: All sources normalize to standard columns (
date,open,high,low,close,volume,amount,pct_chg) through_normalize_datamethods. - Real-Time Support: Separate priority chains in
get_realtime_quotehandle intraday quotes withUnifiedRealtimeQuotenormalization.
Frequently Asked Questions
How does the system handle rate limits from data providers?
Each concrete fetcher implements rate-limit handling within _fetch_raw_data according to provider specifications. TushareFetcher includes specific Tushare Pro API rate-limit logic, while DataFetcherManager._call_fetcher_method serializes calls per fetcher instance to prevent concurrent request flooding. Failed requests trigger automatic failover to the next priority source rather than waiting for rate-limit resets.
What happens if all data sources fail for a symbol?
If the ordered priority list exhausts without successful data retrieval, DataFetcherManager raises a DataFetchError containing summarized exception details from all attempted sources. The error object generated via summarize_exception includes the specific failure reasons (network timeouts, authentication errors, or missing symbols) for debugging.
Can I force a specific data source instead of using automatic selection?
Yes. While DataFetcherManager provides automatic routing, you can instantiate concrete fetchers directly from their respective modules (e.g., from data_provider.tushare_fetcher import TushareFetcher). Direct instantiation bypasses the manager's failover chain but retains the _normalize_data and _clean_data processing pipeline specific to that source.
Why are there different priority lists for daily data versus real-time quotes?
Daily historical data and real-time quotes have different reliability characteristics and API limitations. Daily data prioritizes free, high-volume historical sources (E-Finance, Akshare) while real-time data prioritizes low-latency sources (Longbridge, E-Finance realtime endpoints). The realtime_source_priority configuration allows independent tuning for latency-sensitive intraday trading versus backtesting workflows.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →