TrendRadar Core Functionalities: A Complete Guide to the Modular Python News Aggregator
TrendRadar is a modular Python tool that automates the collection, filtering, ranking, AI enrichment, and notification of trending news from multiple online platforms.
TrendRadar orchestrates an end-to-end pipeline for discovering and distributing trending content. Built with a component-based architecture, it separates concerns across configuration management, scheduling, data crawling, storage, frequency-based analysis, optional AI processing, and multi-channel notifications. This article breaks down each core functionality with direct references to the source code in the sansan0/TrendRadar repository.
Configuration Loading and Environment Management
Every TrendRadar execution begins with unified configuration assembly. The configuration loader in trendradar/core/loader.py reads YAML files from the config/ directory, parses the frequency_words.txt rule file, and merges environment variable overrides into a single dictionary.
Key capabilities include:
- Environment variable overrides for
TIMEZONE,DEBUG,MAX_NEWS_PER_KEYWORD, and other runtime settings - YAML merging that combines
config.yamlwith platform-specific and timeline configurations - Frequency word preprocessing that converts text rules into structured matching directives
from trendradar.core.loader import load_config
cfg = load_config() # reads config/config.yaml by default
print("Enabled platforms:", cfg["PLATFORMS"])
print("Timezone:", cfg["TIMEZONE"])
Timeline-Based Scheduling System
The scheduler in trendradar/core/scheduler.py interprets timeline.yaml to determine which pipeline actions should execute at any given moment. It supports presets, custom periods, daily/weekly plans, and deduplication logic to prevent redundant runs.
The Scheduler class resolves three boolean flags:
collect: Whether to crawl fresh data from platformsanalyze: Whether to process and rank the collected datapush: Whether to distribute notifications to configured channels
from trendradar.core.scheduler import Scheduler
from trendradar.utils.time import get_configured_time
schedule = Scheduler(
schedule_config=cfg["SCHEDULE"],
timeline_data=cfg["_TIMELINE_DATA"],
storage_backend=None,
get_time_func=lambda: get_configured_time(cfg["TIMEZONE"])
)
resolved = schedule.resolve()
print(f"Collect? {resolved.collect}, Analyze? {resolved.analyze}, Push? {resolved.push}")
Multi-Platform Data Crawling
The crawler in trendradar/crawler/__init__.py (organized as the DataFetcher class) performs HTTP requests to configured platforms, respecting per-platform rate limits and optional proxy settings. It returns raw title-rank mappings that feed into downstream analysis.
Crawler capabilities include:
- Request interval enforcement to respect platform rate limits
- Proxy configuration for geographic or network flexibility
- Platform-specific parsers that normalize diverse response formats into consistent structures
Persistent Storage and Incremental Detection
The storage manager in trendradar/storage/__init__.py persists raw crawl results to SQLite and optional TXT/HTML snapshots. It provides read-only access to current and historical data, enabling incremental mode that detects newly added titles between runs.
Storage functions include:
- SQLite persistence for structured querying and deduplication
- Snapshot generation in TXT and HTML formats for manual inspection
- Incremental detection comparing current batches against stored history
Frequency-Word Engine for Filtering and Ranking
The frequency-word engine in trendradar/core/frequency.py parses frequency_words.txt with a flexible syntax supporting required terms (+), filters (!), maximum counts (@), regex patterns (/…/), and display name mapping (=>).
The engine exposes three key functions:
load_frequency_words(): Parses the rule file into structured groupsmatches_word_groups(): Tests titles against all configured groups_word_matches(): Low-level matcher handling individual rule types
Statistical Analysis and Weighted Ranking
The analyzer in trendradar/core/analyzer.py aggregates title matches per keyword, applies multi-factor weighting, and produces final ranked news lists. It also handles RSS feeds through count_rss_frequency and related functions.
Weighting factors include:
- Ranking weight: Original platform position (lower rank = higher weight)
- Frequency weight: How often the keyword appears across platforms
- Hotness weight: Composite momentum score combining recency and velocity
from trendradar.core.analyzer import count_word_frequency
from trendradar.core.frequency import load_frequency_words
word_groups, filter_words, global_filters = load_frequency_words()
stats, total = count_word_frequency(
results, # dict from crawler
word_groups,
filter_words,
id_to_name, # platform_id → readable name
rank_threshold=3,
max_news_per_keyword=5,
sort_by_position_first=True,
quiet=False
)
for group in stats:
print(f'Keyword: {group["word"]} – matched {group["count"]} items')
RSS Feed Integration
TrendRadar treats RSS sources as first-class citizens alongside hot-list platforms. The analyzer's RSS functions apply identical frequency-word filtering and weighting, converting feed items into the same data structure as crawled hot-lists.
This enables unified reporting that blends trending topics from social platforms with curated RSS content, all processed through the same ranking pipeline.
AI-Powered Enrichment and Translation
The AI processor in trendradar/ai/processor.py (and referenced via NewsAnalyzer._run_ai_analysis) provides optional LiteLLM-compatible integration for intelligent content enhancement.
AI capabilities include:
- Content analysis: Summarization, sentiment scoring, and topic extraction
- Translation: Automated conversion to configured target languages
- Intelligent filtering: AI-assisted relevance scoring beyond rule-based matching
Multi-Channel Notification Dispatch
The notification subsystem spans three files in trendradar/notification/:
renderer.py: Formats final reports in Markdown and HTMLsplitter.py: Divides unified reports into channel-specific payloadssenders.py: Implements all platform integrations
Supported channels include Feishu, DingTalk, Telegram, Email, and extensible webhook formats.
from trendradar.notification.senders import FeishuSender
payload = {
"title": "今日热点新闻",
"content": "• 关键字 A: 5 条\n• 关键字 B: 3 条"
}
sender = FeishuSender(
webhook_url="https://open.feishu.cn/open-apis/bot/v2/hook/your-token"
)
sender.send(payload) # HTTP POST to webhook
CLI Entry Point and Workflow Orchestration
The main entry point in trendradar/__main__.py wires all components into a cohesive pipeline. It creates AppContext, executes the scheduler, conditionally runs collection/analysis/notification phases, and generates HTML reports.
Execution modes include:
- Default/daily mode: Full pipeline execution
--mode=current: Only current ranking slice
# Full pipeline (default)
python -m trendradar
# Current ranking only
python -m trendradar --mode=current
Summary
TrendRadar delivers comprehensive trending news automation through these core capabilities:
- Modular configuration with YAML and environment variable overrides
- Timeline scheduling supporting presets, custom periods, and deduplication
- Multi-platform crawling with rate limiting and proxy support
- Persistent storage with SQLite and incremental detection
- Frequency-word engine with flexible rule syntax for filtering and ranking
- Statistical analysis with multi-factor weighting (rank, frequency, hotness)
- RSS integration treated as first-class data sources
- AI enrichment via LiteLLM for summarization, translation, and intelligent filtering
- Multi-channel notifications supporting Feishu, DingTalk, Telegram, Email, and webhooks
- CLI orchestration with configurable execution modes
Frequently Asked Questions
How does TrendRadar prevent duplicate notifications?
The scheduler in trendradar/core/scheduler.py maintains execution records and implements "once-per-day" deduplication. When push is resolved as True, the scheduler checks whether a notification was already sent for the current period before allowing dispatch.
What platforms can TrendRadar crawl for trending topics?
TrendRadar's crawler supports configurable HTTP-based platforms defined in YAML configuration. The extensible DataFetcher class in trendradar/crawler/__init__.py handles per-platform request intervals, proxy settings, and response parsing. Specific platform implementations are loaded based on the PLATFORMS configuration key.
How do frequency-word rules work in TrendRadar?
The frequency-word engine in trendradar/core/frequency.py parses frequency_words.txt with syntax supporting: + (required terms), ! (exclusion filters), @ (maximum match counts), /.../ (regex patterns), and => (display name mapping). The matches_word_groups function tests titles against these rules to categorize and rank content.
Can TrendRadar integrate with custom AI models?
Yes. The AI processor in trendradar/ai/processor.py uses LiteLLM for model compatibility, supporting any provider or self-hosted endpoint that LiteLLM recognizes. Configuration is loaded from YAML with environment-variable overrides for API keys and model selection, enabling integration with OpenAI, Anthropic, local Ollama instances, or custom endpoints.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →