How the Stock Analysis Algorithm Works in daily_stock_analysis: Complete Pipeline Guide

The stock analysis algorithm orchestrates a 12-stage pipeline that aggregates multi-source market data, performs technical trend calculations, enriches context with chip distribution and fundamentals, and generates a structured decision dashboard via LLM.

The daily_stock_analysis repository by ZhuLinsen implements a production-grade stock analysis algorithm designed for both individual investors and automated trading workflows. At its core, the StockAnalysisPipeline class in src/core/pipeline.py coordinates data fetching, technical analysis, and LLM generation, while analyzer_service.py provides high-level convenience functions for CLI, Web UI, and bot integrations. This guide examines the complete algorithm flow, from SQLite-cached OHLCV data to the final persisted analysis result.

The 12-Stage Stock Analysis Pipeline

The algorithm follows a strictly ordered execution path within StockAnalysisPipeline.analyze_stock(), where each stage is resilient to failures and continues with available data.

1. Data Acquisition and Caching

The pipeline first checks the local SQLite cache via src/storage.py for existing daily bars. If data is stale or force_refresh=True, DataFetcherManager pulls 30-day OHLCV history from providers including Tencent, Akshare, E-Finance, and Tushare.


# From src/core/pipeline.py

pipeline.fetch_and_save_stock_data(stock_code)  # Lines 79-106

2. Real-Time Quote Enrichment

When config.enable_realtime_quote is enabled, the algorithm injects live price, volume-ratio, and turnover-rate data. Failures gracefully fall back to the most recent close price without aborting the pipeline.

3. Chip Distribution Analysis

The algorithm retrieves shareholder structure data (profit ratio, concentration, average cost) through fetcher_manager.get_chip_distribution(). Missing chip data is logged but does not block downstream stages.

4. Fundamental Context Aggregation

Under a configurable timeout (config.fundamental_stage_timeout_seconds), the pipeline calls fetcher_manager.get_fundamental_context() to gather valuation metrics and financial ratios. Timeout failures result in a placeholder context to ensure continuity.

5. Technical Trend Calculation

Historical bars (approximately 90 days) are loaded from SQLite and optionally augmented with real-time quotes. The StockTrendAnalyzer in src/stock_analyzer.py computes moving averages, price bias, and trend strength signals.


# Technical analysis invocation (pipeline.py lines 144-162)

analyzer = StockTrendAnalyzer(price_df)
trend_result = analyzer.analyze()

6. Agent Mode Routing

If config.agent_mode is active, control transfers to an Agent executor in src/agent/factory.py at lines 264-277, enabling multi-step tool use while consuming the same enriched dataset.

When SearchService is available, the algorithm issues up to five parallel searches (news, risk, earnings) and formats results via search_service.format_intel_report() (lines 178-199).

8. Social Sentiment Integration

For US tickers (is_us_stock_code), the SocialSentimentService aggregates Reddit, X (Twitter), and Polymarket sentiment data, merging it into the news context block (lines 200-209).

9. Context Enhancement

All disparate data sources—realtime quotes, chip distribution, trend results, fundamentals, and news—are unified in _enhance_context() (lines 222-272). This step calculates derived fields including ma_status, price_change_ratio, and the is_index_etf flag.

10. LLM Generation

The enriched dictionary is passed to GeminiAnalyzer.analyze() in src/analyzer.py, a LiteLLM wrapper that streams tokens and enforces a strict JSON dashboard schema defined in LEGACY_DEFAULT_SYSTEM_PROMPT.

11. Post-Processing and Fallbacks

After LLM generation, the algorithm applies integrity checks and fills mandatory missing fields:

  • fill_chip_structure_if_needed() - Populates default chip metrics
  • fill_price_position_if_needed() - Calculates MA5/MA20 position data
  • check_content_integrity() → apply_placeholder_fill() - Optional placeholder replacement

12. Persistence and Notification

The final AnalysisResult is serialized to SQLite via db.save_analysis_history() (lines 314-322), enabling historical comparison, while NotificationService optionally pushes results to WeChat, Feishu, or Telegram.

Core Components and Source Files

Component Responsibility Source File
StockAnalysisPipeline Orchestrates the 12-stage workflow, timeout handling, and error recovery. src/core/pipeline.py
GeminiAnalyzer LiteLLM wrapper for streaming JSON dashboard generation and schema validation. src/analyzer.py
DataFetcherManager Unified interface with automatic failover across Chinese and US data providers. data_provider/__init__.py
StockTrendAnalyzer Technical indicator calculation (MA, bias, signals) for trend determination. src/stock_analyzer.py
SearchService Multi-provider web search abstraction (Bocha, Tavily, Anspire, SearXNG). src/search_service.py
SocialSentimentService US-specific sentiment aggregation from Reddit/X/Polymarket. src/services/social_sentiment_service.py
NotificationService Multi-channel report distribution layer. src/notification.py

Code Examples: Running the Algorithm

Single Stock Analysis

Use the high-level analyze_stock function from analyzer_service.py for individual ticker analysis:

from analyzer_service import analyze_stock
from src.config import get_config

# Analyze Kweichow Moutai with full dashboard

result = analyze_stock(
    stock_code="600519",
    config=get_config(),
    full_report=True,
)

if result and result.success:
    print("核心结论:", result.get_core_conclusion())
    print("操作建议:", result.operation_advice)
    print("MA5价格:", result.dashboard["data_perspective"]["price_position"]["ma5"])

Batch Analysis

Process multiple securities efficiently using analyze_stocks:

from analyzer_service import analyze_stocks

codes = ["AAPL", "600519", "hk00700"]
reports = analyze_stocks(
    stock_codes=codes,
    full_report=False,  # Simple summary mode

)

for r in reports:
    print(f"{r.code}: {r.operation_advice} – {r.trend_prediction}")

Direct Pipeline Access

Access the pipeline directly for advanced configuration or custom query IDs:

from src.core.pipeline import StockAnalysisPipeline
from src.config import get_config
from src.enums import ReportType

cfg = get_config()
pipeline = StockAnalysisPipeline(
    config=cfg,
    query_id="custom_analysis_001",
    query_source="cli"
)

result = pipeline.analyze_stock(
    code="000001",
    report_type=ReportType.FULL,
    skip_analysis=False,
)

print(result.dashboard["core_conclusion"]["one_sentence"])

Market Review Generation

Generate macro-level market recaps using the same algorithm components:

from analyzer_service import perform_market_review

review_text = perform_market_review(config=get_config())
print("今日大盘复盘:\n", review_text)

Summary

  • The algorithm executes through a resilient 12-stage pipeline orchestrated by StockAnalysisPipeline in src/core/pipeline.py.
  • Data fetching supports multiple providers (Tencent, Akshare, Tushare, YFinance) with automatic failover and local SQLite caching.
  • Technical analysis leverages StockTrendAnalyzer for moving-average calculations across 90-day historical windows.
  • Context enrichment combines chip distribution, fundamental data, multi-source news, and US social sentiment into a unified dictionary.
  • LLM generation uses GeminiAnalyzer (LiteLLM) to produce JSON-structured decision dashboards with strict schema validation.
  • Post-processing ensures data integrity via fallback functions before persisting to SQLite and optionally notifying via webhook services.

Frequently Asked Questions

What data sources does the stock analysis algorithm support?

The algorithm supports Chinese market data through Tencent, Akshare, E-Finance, and Tushare providers, while US equities utilize YFinance. The DataFetcherManager automatically cycles through available providers if one fails, and all fetched data is cached in SQLite to minimize API calls.

How does the algorithm handle missing real-time data?

If the real-time quote fetch fails or is disabled, the algorithm gracefully falls back to the most recent cached closing price from the SQLite database. This resilience pattern is implemented in the realtime-quote branch of pipeline.analyze_stock() without aborting subsequent stages.

Can I use a different LLM model with this algorithm?

Yes. The GeminiAnalyzer class in src/analyzer.py is a thin wrapper around LiteLLM, allowing you to configure any supported model (OpenAI, Anthropic, local models) through environment variables or the Config singleton. The system prompt and JSON schema constraints remain consistent regardless of the underlying model provider.

Is this algorithm suitable for high-frequency trading?

No. The algorithm is optimized for daily analysis and decision support, with optional real-time quote enrichment. Data fetching includes 30-90 day historical windows, and the LLM generation stage introduces latency unsuitable for high-frequency strategies. It is designed for swing trading, long-term analysis, and automated daily reporting workflows.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →