How to Parse TrendRadar Output: SQLite, Snapshots, and Python API Guide
Parse TrendRadar output by reading SQLite .db files from the output/news/ or output/rss/ directories, using ParserService for high-level access, or loading plain-text and HTML snapshots from output/txt/ and output/html/ for historic views.
TrendRadar is an open-source trend monitoring system that stores daily crawl results in a structured, date-based file hierarchy. Understanding how to parse this output is essential for custom analytics, reporting, or integration with external systems. This guide covers the complete output format, SQLite schemas, and the Python API provided by the mcp_server/services/parser_service.py module.
TrendRadar Output Directory Structure
TrendRadar organizes all output under a top-level output/ directory with five primary subdirectories:
output/
├─ news/ # news data, one SQLite DB per day
│ └─ 2025-12-28.db
├─ rss/ # RSS data, one SQLite DB per day
│ └─ 2025-12-28.db
├─ txt/ # plain-text snapshots (date-folder → *.txt)
│ └─ 2025-12-28/
│ └─ 12-34-56.txt
├─ html/ # rendered HTML snapshots (date-folder → *.html)
│ └─ 2025-12-28/
│ └─ 12-34-56.html
└─ meta/ # auxiliary meta files (e.g. doctor_report.json)
Each data type (news, rss) receives its own SQLite database per day. Snapshot files preserve the exact rendered output at specific crawl times, useful for auditing or display purposes.
SQLite Database Schema for News and RSS Data
The SQLite .db files follow a relational schema designed for efficient trend tracking and historical analysis.
News Database Schema
| Table | Purpose | Key Columns |
|---|---|---|
news_items |
One row per headline | id, platform_id, title, rank, url, mobile_url |
platforms |
Maps platform_id to human-readable name |
id, name |
rank_history |
Historic rank changes per news_item_id |
news_item_id, rank, timestamp |
crawl_records |
Timestamps of each crawl run | id, started_at, completed_at |
RSS Database Schema
The RSS schema mirrors the news structure with renamed tables:
| Table | RSS Equivalent |
|---|---|
rss_items |
news_items |
rss_feeds |
platforms |
rss_crawl_records |
crawl_records |
The reading logic is implemented in ParserService._read_news_from_sqlite and ParserService._read_rss_from_sqlite within mcp_server/services/parser_service.py.
Using ParserService to Parse TrendRadar Output
The ParserService class in mcp_server/services/parser_service.py provides the high-level API for accessing TrendRadar output without manual SQL queries.
Initializing ParserService
from mcp_server.services.parser_service import ParserService
# Auto-detect project root and initialize cache
parser = ParserService()
The constructor accepts an optional project_root parameter if your working directory differs from the repository structure.
Reading Titles for a Specific Date
from datetime import datetime, timedelta
yesterday = datetime.now() - timedelta(days=1)
# Read all news titles from yesterday
titles, id_name_map, timestamps = parser.read_all_titles_for_date(
date=yesterday,
db_type="news" # or "rss"
)
print(f"Found {len(titles)} platforms with news on {yesterday.date()}")
for src_id, item_dict in titles.items():
platform_name = id_name_map.get(src_id, src_id)
print(f"📰 {platform_name} – {len(item_dict)} headlines")
The read_all_titles_for_date method handles:
- Cache management for repeated access
- SQLite database opening and closing
- Fallback to
Noneif no data exists for the requested date
Listing Available Dates
# Get all dates with news data
available_dates = parser.get_available_dates("news")
print(f"News data available for {len(available_dates)} days")
# Check specific date availability
target_date = "2025-12-28"
if target_date in available_dates:
print(f"Data available for {target_date}")
Dates are returned as strings in YYYY-MM-DD format, matching the database filenames.
Converting Raw Crawler Output to Structured Data
When working with raw crawler results rather than stored databases, use convert_crawl_results_to_news_data from trendradar/storage/base.py.
from mcp_server.services.parser_service import convert_crawl_results_to_news_data
# Raw crawler output format: {source_id: {title: {properties}}}
raw_crawler_output = {
"newsnow": {
"Example headline": {
"ranks": [1, 2],
"url": "https://example.com/article",
"mobileUrl": ""
},
"Another trending story": {
"ranks": [3],
"url": "https://example.com/other",
"mobileUrl": "https://m.example.com/other"
}
}
}
id_to_name = {"newsnow": "NewsNow"}
failed_ids = [] # Empty list if no crawl failures
news_data = convert_crawl_results_to_news_data(
results=raw_crawler_output,
id_to_name=id_to_name,
failed_ids=failed_ids,
crawl_time="09:15",
crawl_date="2025-12-28"
)
# Access structured data
first_item = news_data.items["newsnow"][0]
print(f"Title: {first_item.title}")
print(f"Source: {first_item.source_name}")
print(f"Current rank: {first_item.rank}")
print(f"Rank history: {first_item.ranks}")
The resulting NewsData object contains:
dateandcrawl_time: When the data was collecteditems: Dictionary mappingsource_idto lists ofNewsItemobjectsid_to_name: Mapping of source IDs to display namesfailed_ids: List of sources that failed to crawl
Each NewsItem includes title, source_id, source_name, rank, url, mobile_url, crawl_time, ranks (history), first_time, last_time, and count.
Parsing Historic Snapshot Files
For the exact rendered output shown to users at a specific time, read the snapshot files directly from output/txt/ or output/html/.
Loading Text Snapshots
from pathlib import Path
def load_text_snapshot(date: str, time: str) -> str:
"""
Load a plain-text snapshot.
Date format: YYYY-MM-DD
Time format: HH-MM-SS (as stored in filename)
"""
path = Path("output") / "txt" / date / f"{time}.txt"
return path.read_text(encoding="utf-8")
# Example: Load snapshot from 2025-12-28 at 09:15:00
text_content = load_text_snapshot("2025-12-28", "09-15-00")
print(text_content)
Loading HTML Snapshots
from pathlib import Path
def load_html_snapshot(date: str, time: str) -> str:
"""
Load a rendered HTML snapshot.
"""
path = Path("output") / "html" / date / f"{time}.html"
return path.read_text(encoding="utf-8")
# Example: Load HTML report
html_content = load_html_snapshot("2025-12-28", "09-15-00")
# Can be written to file, served via HTTP, or parsed with BeautifulSoup
Snapshot directories are created by generate_html_report in trendradar/report/generator.py. The filename timestamp uses hyphen separators (HH-MM-SS) rather than colons for filesystem compatibility.
Summary
- TrendRadar output resides in the
output/directory with separate subdirectories fornews/,rss/,txt/,html/, andmeta/data - SQLite databases in
news/andrss/contain structured data withnews_items/rss_items,platforms/rss_feeds, and crawl history tables - ParserService in
mcp_server/services/parser_service.pyprovides the high-level API:read_all_titles_for_date(),get_available_dates(), andconvert_crawl_results_to_news_data() - Snapshot files in
txt/andhtml/preserve exact rendered output for specific crawl times, readable directly as text files - Data models
NewsDataandNewsItemintrendradar/storage/base.pystructure the parsed output with rank history, URLs, and metadata
Frequently Asked Questions
What database format does TrendRadar use?
TrendRadar uses SQLite .db files for structured data storage. Each day generates separate databases for news and RSS data in output/news/YYYY-MM-DD.db and output/rss/YYYY-MM-DD.db. The schema includes relational tables for items, platforms, and crawl history.
How do I check which dates have available data?
Use ParserService.get_available_dates(db_type) to list all dates with existing databases. Pass "news" or "rss" as the db_type parameter. The method returns a list of date strings in YYYY-MM-DD format extracted from database filenames in the respective output directory.
Can I access historical snapshots without parsing SQLite?
Yes. TrendRadar saves plain-text and HTML snapshots in output/txt/YYYY-MM-DD/HH-MM-SS.txt and output/html/YYYY-MM-DD/HH-MM-SS.html. These files contain the exact rendered output from specific crawl times. Read them directly with standard file I/O—no database parsing required.
What is the difference between raw crawler output and NewsData?
Raw crawler output is a nested dictionary mapping source IDs to title dictionaries with properties like ranks, url, and mobileUrl. NewsData is a structured Pydantic model with typed fields: date, crawl_time, items (dict of NewsItem lists), id_to_name mapping, and failed_ids. Use convert_crawl_results_to_news_data() to transform raw output into the structured model.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →