How the Stock Analysis Algorithm Works in daily_stock_analysis: Complete Pipeline Guide
The stock analysis algorithm orchestrates a 12-stage pipeline that aggregates multi-source market data, performs technical trend calculations, enriches context with chip distribution and fundamentals, and generates a structured decision dashboard via LLM.
The daily_stock_analysis repository by ZhuLinsen implements a production-grade stock analysis algorithm designed for both individual investors and automated trading workflows. At its core, the StockAnalysisPipeline class in src/core/pipeline.py coordinates data fetching, technical analysis, and LLM generation, while analyzer_service.py provides high-level convenience functions for CLI, Web UI, and bot integrations. This guide examines the complete algorithm flow, from SQLite-cached OHLCV data to the final persisted analysis result.
The 12-Stage Stock Analysis Pipeline
The algorithm follows a strictly ordered execution path within StockAnalysisPipeline.analyze_stock(), where each stage is resilient to failures and continues with available data.
1. Data Acquisition and Caching
The pipeline first checks the local SQLite cache via src/storage.py for existing daily bars. If data is stale or force_refresh=True, DataFetcherManager pulls 30-day OHLCV history from providers including Tencent, Akshare, E-Finance, and Tushare.
# From src/core/pipeline.py
pipeline.fetch_and_save_stock_data(stock_code) # Lines 79-106
2. Real-Time Quote Enrichment
When config.enable_realtime_quote is enabled, the algorithm injects live price, volume-ratio, and turnover-rate data. Failures gracefully fall back to the most recent close price without aborting the pipeline.
3. Chip Distribution Analysis
The algorithm retrieves shareholder structure data (profit ratio, concentration, average cost) through fetcher_manager.get_chip_distribution(). Missing chip data is logged but does not block downstream stages.
4. Fundamental Context Aggregation
Under a configurable timeout (config.fundamental_stage_timeout_seconds), the pipeline calls fetcher_manager.get_fundamental_context() to gather valuation metrics and financial ratios. Timeout failures result in a placeholder context to ensure continuity.
5. Technical Trend Calculation
Historical bars (approximately 90 days) are loaded from SQLite and optionally augmented with real-time quotes. The StockTrendAnalyzer in src/stock_analyzer.py computes moving averages, price bias, and trend strength signals.
# Technical analysis invocation (pipeline.py lines 144-162)
analyzer = StockTrendAnalyzer(price_df)
trend_result = analyzer.analyze()
6. Agent Mode Routing
If config.agent_mode is active, control transfers to an Agent executor in src/agent/factory.py at lines 264-277, enabling multi-step tool use while consuming the same enriched dataset.
7. Multi-Dimensional News Search
When SearchService is available, the algorithm issues up to five parallel searches (news, risk, earnings) and formats results via search_service.format_intel_report() (lines 178-199).
8. Social Sentiment Integration
For US tickers (is_us_stock_code), the SocialSentimentService aggregates Reddit, X (Twitter), and Polymarket sentiment data, merging it into the news context block (lines 200-209).
9. Context Enhancement
All disparate data sources—realtime quotes, chip distribution, trend results, fundamentals, and news—are unified in _enhance_context() (lines 222-272). This step calculates derived fields including ma_status, price_change_ratio, and the is_index_etf flag.
10. LLM Generation
The enriched dictionary is passed to GeminiAnalyzer.analyze() in src/analyzer.py, a LiteLLM wrapper that streams tokens and enforces a strict JSON dashboard schema defined in LEGACY_DEFAULT_SYSTEM_PROMPT.
11. Post-Processing and Fallbacks
After LLM generation, the algorithm applies integrity checks and fills mandatory missing fields:
fill_chip_structure_if_needed()- Populates default chip metricsfill_price_position_if_needed()- Calculates MA5/MA20 position datacheck_content_integrity()→apply_placeholder_fill()- Optional placeholder replacement
12. Persistence and Notification
The final AnalysisResult is serialized to SQLite via db.save_analysis_history() (lines 314-322), enabling historical comparison, while NotificationService optionally pushes results to WeChat, Feishu, or Telegram.
Core Components and Source Files
| Component | Responsibility | Source File |
|---|---|---|
| StockAnalysisPipeline | Orchestrates the 12-stage workflow, timeout handling, and error recovery. | src/core/pipeline.py |
| GeminiAnalyzer | LiteLLM wrapper for streaming JSON dashboard generation and schema validation. | src/analyzer.py |
| DataFetcherManager | Unified interface with automatic failover across Chinese and US data providers. | data_provider/__init__.py |
| StockTrendAnalyzer | Technical indicator calculation (MA, bias, signals) for trend determination. | src/stock_analyzer.py |
| SearchService | Multi-provider web search abstraction (Bocha, Tavily, Anspire, SearXNG). | src/search_service.py |
| SocialSentimentService | US-specific sentiment aggregation from Reddit/X/Polymarket. | src/services/social_sentiment_service.py |
| NotificationService | Multi-channel report distribution layer. | src/notification.py |
Code Examples: Running the Algorithm
Single Stock Analysis
Use the high-level analyze_stock function from analyzer_service.py for individual ticker analysis:
from analyzer_service import analyze_stock
from src.config import get_config
# Analyze Kweichow Moutai with full dashboard
result = analyze_stock(
stock_code="600519",
config=get_config(),
full_report=True,
)
if result and result.success:
print("核心结论:", result.get_core_conclusion())
print("操作建议:", result.operation_advice)
print("MA5价格:", result.dashboard["data_perspective"]["price_position"]["ma5"])
Batch Analysis
Process multiple securities efficiently using analyze_stocks:
from analyzer_service import analyze_stocks
codes = ["AAPL", "600519", "hk00700"]
reports = analyze_stocks(
stock_codes=codes,
full_report=False, # Simple summary mode
)
for r in reports:
print(f"{r.code}: {r.operation_advice} – {r.trend_prediction}")
Direct Pipeline Access
Access the pipeline directly for advanced configuration or custom query IDs:
from src.core.pipeline import StockAnalysisPipeline
from src.config import get_config
from src.enums import ReportType
cfg = get_config()
pipeline = StockAnalysisPipeline(
config=cfg,
query_id="custom_analysis_001",
query_source="cli"
)
result = pipeline.analyze_stock(
code="000001",
report_type=ReportType.FULL,
skip_analysis=False,
)
print(result.dashboard["core_conclusion"]["one_sentence"])
Market Review Generation
Generate macro-level market recaps using the same algorithm components:
from analyzer_service import perform_market_review
review_text = perform_market_review(config=get_config())
print("今日大盘复盘:\n", review_text)
Summary
- The algorithm executes through a resilient 12-stage pipeline orchestrated by
StockAnalysisPipelineinsrc/core/pipeline.py. - Data fetching supports multiple providers (Tencent, Akshare, Tushare, YFinance) with automatic failover and local SQLite caching.
- Technical analysis leverages
StockTrendAnalyzerfor moving-average calculations across 90-day historical windows. - Context enrichment combines chip distribution, fundamental data, multi-source news, and US social sentiment into a unified dictionary.
- LLM generation uses
GeminiAnalyzer(LiteLLM) to produce JSON-structured decision dashboards with strict schema validation. - Post-processing ensures data integrity via fallback functions before persisting to SQLite and optionally notifying via webhook services.
Frequently Asked Questions
What data sources does the stock analysis algorithm support?
The algorithm supports Chinese market data through Tencent, Akshare, E-Finance, and Tushare providers, while US equities utilize YFinance. The DataFetcherManager automatically cycles through available providers if one fails, and all fetched data is cached in SQLite to minimize API calls.
How does the algorithm handle missing real-time data?
If the real-time quote fetch fails or is disabled, the algorithm gracefully falls back to the most recent cached closing price from the SQLite database. This resilience pattern is implemented in the realtime-quote branch of pipeline.analyze_stock() without aborting subsequent stages.
Can I use a different LLM model with this algorithm?
Yes. The GeminiAnalyzer class in src/analyzer.py is a thin wrapper around LiteLLM, allowing you to configure any supported model (OpenAI, Anthropic, local models) through environment variables or the Config singleton. The system prompt and JSON schema constraints remain consistent regardless of the underlying model provider.
Is this algorithm suitable for high-frequency trading?
No. The algorithm is optimized for daily analysis and decision support, with optional real-time quote enrichment. Data fetching includes 30-90 day historical windows, and the LLM generation stage introduces latency unsuitable for high-frequency strategies. It is designed for swing trading, long-term analysis, and automated daily reporting workflows.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →