What Is the Architecture of TrendRadar? A Deep Dive into the Modular Python News Aggregator
TrendRadar is a modular Python application with a layered architecture that separates crawling, storage, analysis, AI filtering, and notification into pluggable components coordinated through a central AppContext object.
This article explains the complete architecture of TrendRadar, an open-source news aggregation and reporting system. You'll learn how each layer functions, how data flows through the pipeline, and how the design enables extensibility without core code changes.
Architectural Overview
TrendRadar follows a layered, context-driven architecture with clear separation of concerns. The system coordinates 11 distinct layers through a shared AppContext object that eliminates global state and enables testability.
| Layer | Responsibility | Main File Paths |
|---|---|---|
| Entry Point | CLI/daemon startup, argument parsing | trendradar/__main__.py |
| Configuration & Context | YAML/ENV loading, unified context object | trendradar/core/loader.py, trendradar/context.py |
| Crawler | Platform web scraping and RSS fetching | trendradar/crawler/DataFetcher, trendradar/crawler/rss.py |
| Storage | SQLite persistence, retention, remote sync | trendradar/storage/manager.py, trendradar/storage/sqlite_mixin.py |
| Analysis | Keyword matching, weight calculation, aggregation | trendradar/core/analyzer.py, trendradar/core/frequency.py |
| AI Filter | LLM-based news classification and tagging | trendradar/ai/filter.py, trendradar/ai/client.py |
| Reporting | HTML generation, data structure preparation | trendradar/report/__init__.py, trendradar/report/generator.py |
| Notification | Multi-channel dispatch with batching and translation | trendradar/notification/dispatcher.py, trendradar/notification/renderer.py |
| Scheduler | Timeline-based execution with "once" semantics | trendradar/core/scheduler.py |
Core Design Patterns
Context-Driven Dependency Injection
Every subsystem receives the same AppContext instance, making the codebase testable and free of global state. The AppContext class in trendradar/context.py exposes:
- Configuration access
- Time utilities (
load_frequency_words,count_frequency) - Storage interface
- High-level workflow methods (
prepare_report,generate_html,run_ai_filter)
Pluggable Filtering Architecture
The FILTER.METHOD configuration toggles between keyword matching and AI-driven tagging without changing the surrounding pipeline. This is implemented through:
trendradar/core/analyzer.py– keyword-based word frequency and weight calculationtrendradar/ai/filter.py– LLM classification with tag extraction and versioning
Both return compatible data structures that feed into the same reporting and notification layers.
Data Flow Through the Pipeline
Understanding the complete TrendRadar architecture requires tracing how data moves from source to notification:
1. Startup and Configuration
from trendradar.core.loader import load_config
from trendradar.context import AppContext
# Load configuration (uses CONFIG_PATH env var or defaults)
cfg = load_config()
# Create shared context
ctx = AppContext(cfg)
The load_config() function in trendradar/core/loader.py loads and validates the entire configuration tree from YAML files and environment variables.
2. Scheduling Resolution
# Create scheduler and check if we should run
scheduler = ctx.create_scheduler()
should_run = scheduler.should_crawl() # or should_analyze(), should_push()
The trendradar/core/scheduler.py module implements timeline-based scheduling with cron-like periods and "once" flags to prevent duplicate executions.
3. Data Crawling
# Crawl platforms and RSS feeds
results, id_to_name, failed = ctx.create_scheduler().crawl_all()
The crawler layer in trendradar/crawler/DataFetcher and trendradar/crawler/rss.py fetches raw news from web platforms and RSS feeds, normalizing everything into a common data structure.
4. Storage and Deduplication
Crawled items persist through trendradar/storage/manager.py, which provides:
- SQLite storage via
sqlite_mixin.py - Retention policies for automatic cleanup
- Remote sync capabilities via
trendradar/storage/remote.py
New-title detection happens through ctx.detect_new_titles() and ctx.detect_latest_new_titles(), which identify titles appearing for the first time today.
5. Analysis and Filtering
Keyword mode:
stats, total = ctx.count_frequency(
results,
*ctx.load_frequency_words(),
id_to_name,
title_info,
new_titles=new_titles,
mode="daily"
)
The trendradar/core/analyzer.py module handles keyword-group matching, weight calculation, and platform-wise aggregation. Weights combine rank, frequency, and hotness scores.
AI mode:
# Enable AI filtering
cfg["FILTER"]["METHOD"] = "ai"
ctx = AppContext(cfg)
# Run AI filter
ai_result = ctx.run_ai_filter(interests_file="my_interests.txt")
# Convert to report format
hotlist_stats, rss_stats = ctx.convert_ai_filter_to_report_data(ai_result, mode="daily")
The trendradar/ai/filter.py pipeline classifies items using large language models, handles tag extraction and versioning, and returns a unified result set compatible with the reporting layer.
6. Report Generation
# Generate HTML report
html_path = ctx.generate_html(
stats,
total,
new_titles=new_titles,
id_to_name=id_to_name,
mode="daily"
)
# Prepare language-agnostic report data
report_data = ctx.prepare_report(hotlist_stats, id_to_name=id_to_name, mode="daily")
The reporting layer in trendradar/report/__init__.py and trendradar/report/generator.py builds data structures for HTML and text rendering. Templates reside in docs/.
7. Notification Dispatch
dispatcher = ctx.create_notification_dispatcher()
dispatcher.dispatch_all(
report_data=report_data,
report_type="热点分析报告",
update_info=None,
proxy_url=None,
mode="daily",
html_file_path=None,
rss_items=rss_stats,
rss_new_items=None,
ai_analysis=None,
standalone_data=None,
skip_translation=False,
)
The trendradar/notification/dispatcher.py sends rendered reports to 10+ channels: Feishu, DingTalk, WeWork, Telegram, Email, Ntfy, Bark, Slack, and generic webhooks. Key supporting modules:
trendradar/notification/renderer.py– channel-specific payload formattingtrendradar/notification/splitter.py– batch splitting for large reportsAITranslator– optional translation to reduce API costs
Key Design Decisions
Flexible Scheduling with Timeline Semantics
The timeline.yaml configuration defines daily/weekly "periods" (e.g., "morning", "evening") with "once" flags to prevent duplicate executions. This is implemented in trendradar/core/scheduler.py through the ResolvedSchedule class.
Storage Abstraction
StorageManager in trendradar/storage/manager.py hides implementation details behind a uniform API:
| Backend | Implementation |
|---|---|
| Local SQLite | sqlite_mixin.py |
| Plain text | Direct file I/O |
| Remote S3 | trendradar/storage/remote.py |
| Pull sync | Sync protocol in remote.py |
This allows swapping backends without touching business logic.
Modular Notification System
Each notification channel has its own renderer and batch size configuration. The dispatcher can skip translation for already-translated data, reducing API costs when multiple channels share the same content.
Summary
- TrendRadar architecture is a layered, modular Python application with 11 distinct functional layers coordinated through
AppContext - Context-driven dependency injection eliminates global state and enables comprehensive testing
- Pluggable filtering supports both keyword-based and AI-driven analysis without pipeline changes
- Timeline-based scheduling with "once" semantics prevents duplicate executions
- Storage abstraction unifies SQLite, file, and remote S3 backends behind a common API
- Multi-channel notifications support 10+ platforms with batching, translation, and channel-specific rendering
Frequently Asked Questions
What programming pattern does TrendRadar use for dependency management?
TrendRadar uses context-driven dependency injection. Every subsystem receives the same AppContext instance passed through constructors or method parameters. This pattern eliminates global state, makes components individually testable with mocked contexts, and centralizes access to configuration, storage, and utilities in trendradar/context.py.
How does TrendRadar switch between keyword and AI filtering?
TrendRadar toggles filtering methods through the FILTER.METHOD configuration key. Setting "keyword" activates the trendradar/core/analyzer.py pipeline for word frequency and weight calculation. Setting "ai" activates trendradar/ai/filter.py for LLM-based classification. Both pipelines return compatible data structures, so trendradar/report/__init__.py and trendradar/notification/dispatcher.py work unchanged regardless of filtering method.
Can TrendRadar store data on remote servers instead of locally?
Yes. The StorageManager class in trendradar/storage/manager.py abstracts storage backends behind a uniform API. While local SQLite is the default via trendradar/storage/sqlite_mixin.py, the system also supports remote S3 storage and pull-sync protocols through trendradar/storage/remote.py. You can swap backends by changing configuration without modifying business logic in crawlers, analyzers, or notifiers.
What notification channels does TrendRadar support?
TrendRadar supports 10 notification channels: Feishu, DingTalk, WeWork, Telegram, Email, Ntfy, Bark, Slack, and generic webhooks. The trendradar/notification/dispatcher.py coordinates delivery, while trendradar/notification/renderer.py formats channel-specific payloads. The trendradar/notification/splitter.py handles batching for large reports, and optional AITranslator integration reduces API costs when multiple channels share content.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →