How to Parse TrendRadar Output: SQLite, Snapshots, and Python API Guide

Parse TrendRadar output by reading SQLite .db files from the output/news/ or output/rss/ directories, using ParserService for high-level access, or loading plain-text and HTML snapshots from output/txt/ and output/html/ for historic views.

TrendRadar is an open-source trend monitoring system that stores daily crawl results in a structured, date-based file hierarchy. Understanding how to parse this output is essential for custom analytics, reporting, or integration with external systems. This guide covers the complete output format, SQLite schemas, and the Python API provided by the mcp_server/services/parser_service.py module.

TrendRadar Output Directory Structure

TrendRadar organizes all output under a top-level output/ directory with five primary subdirectories:


output/
├─ news/                # news data, one SQLite DB per day

│   └─ 2025-12-28.db
├─ rss/                 # RSS data, one SQLite DB per day

│   └─ 2025-12-28.db
├─ txt/                 # plain-text snapshots (date-folder → *.txt)

│   └─ 2025-12-28/
│       └─ 12-34-56.txt
├─ html/                # rendered HTML snapshots (date-folder → *.html)

│   └─ 2025-12-28/
│       └─ 12-34-56.html
└─ meta/                # auxiliary meta files (e.g. doctor_report.json)

Each data type (news, rss) receives its own SQLite database per day. Snapshot files preserve the exact rendered output at specific crawl times, useful for auditing or display purposes.

SQLite Database Schema for News and RSS Data

The SQLite .db files follow a relational schema designed for efficient trend tracking and historical analysis.

News Database Schema

Table Purpose Key Columns
news_items One row per headline id, platform_id, title, rank, url, mobile_url
platforms Maps platform_id to human-readable name id, name
rank_history Historic rank changes per news_item_id news_item_id, rank, timestamp
crawl_records Timestamps of each crawl run id, started_at, completed_at

RSS Database Schema

The RSS schema mirrors the news structure with renamed tables:

Table RSS Equivalent
rss_items news_items
rss_feeds platforms
rss_crawl_records crawl_records

The reading logic is implemented in ParserService._read_news_from_sqlite and ParserService._read_rss_from_sqlite within mcp_server/services/parser_service.py.

Using ParserService to Parse TrendRadar Output

The ParserService class in mcp_server/services/parser_service.py provides the high-level API for accessing TrendRadar output without manual SQL queries.

Initializing ParserService

from mcp_server.services.parser_service import ParserService

# Auto-detect project root and initialize cache

parser = ParserService()

The constructor accepts an optional project_root parameter if your working directory differs from the repository structure.

Reading Titles for a Specific Date

from datetime import datetime, timedelta

yesterday = datetime.now() - timedelta(days=1)

# Read all news titles from yesterday

titles, id_name_map, timestamps = parser.read_all_titles_for_date(
    date=yesterday,
    db_type="news"  # or "rss"

)

print(f"Found {len(titles)} platforms with news on {yesterday.date()}")
for src_id, item_dict in titles.items():
    platform_name = id_name_map.get(src_id, src_id)
    print(f"📰 {platform_name} – {len(item_dict)} headlines")

The read_all_titles_for_date method handles:

  • Cache management for repeated access
  • SQLite database opening and closing
  • Fallback to None if no data exists for the requested date

Listing Available Dates


# Get all dates with news data

available_dates = parser.get_available_dates("news")
print(f"News data available for {len(available_dates)} days")

# Check specific date availability

target_date = "2025-12-28"
if target_date in available_dates:
    print(f"Data available for {target_date}")

Dates are returned as strings in YYYY-MM-DD format, matching the database filenames.

Converting Raw Crawler Output to Structured Data

When working with raw crawler results rather than stored databases, use convert_crawl_results_to_news_data from trendradar/storage/base.py.

from mcp_server.services.parser_service import convert_crawl_results_to_news_data

# Raw crawler output format: {source_id: {title: {properties}}}

raw_crawler_output = {
    "newsnow": {
        "Example headline": {
            "ranks": [1, 2],
            "url": "https://example.com/article",
            "mobileUrl": ""
        },
        "Another trending story": {
            "ranks": [3],
            "url": "https://example.com/other",
            "mobileUrl": "https://m.example.com/other"
        }
    }
}

id_to_name = {"newsnow": "NewsNow"}
failed_ids = []  # Empty list if no crawl failures

news_data = convert_crawl_results_to_news_data(
    results=raw_crawler_output,
    id_to_name=id_to_name,
    failed_ids=failed_ids,
    crawl_time="09:15",
    crawl_date="2025-12-28"
)

# Access structured data

first_item = news_data.items["newsnow"][0]
print(f"Title: {first_item.title}")
print(f"Source: {first_item.source_name}")
print(f"Current rank: {first_item.rank}")
print(f"Rank history: {first_item.ranks}")

The resulting NewsData object contains:

  • date and crawl_time: When the data was collected
  • items: Dictionary mapping source_id to lists of NewsItem objects
  • id_to_name: Mapping of source IDs to display names
  • failed_ids: List of sources that failed to crawl

Each NewsItem includes title, source_id, source_name, rank, url, mobile_url, crawl_time, ranks (history), first_time, last_time, and count.

Parsing Historic Snapshot Files

For the exact rendered output shown to users at a specific time, read the snapshot files directly from output/txt/ or output/html/.

Loading Text Snapshots

from pathlib import Path

def load_text_snapshot(date: str, time: str) -> str:
    """
    Load a plain-text snapshot.
    Date format: YYYY-MM-DD
    Time format: HH-MM-SS (as stored in filename)
    """
    path = Path("output") / "txt" / date / f"{time}.txt"
    return path.read_text(encoding="utf-8")

# Example: Load snapshot from 2025-12-28 at 09:15:00

text_content = load_text_snapshot("2025-12-28", "09-15-00")
print(text_content)

Loading HTML Snapshots

from pathlib import Path

def load_html_snapshot(date: str, time: str) -> str:
    """
    Load a rendered HTML snapshot.
    """
    path = Path("output") / "html" / date / f"{time}.html"
    return path.read_text(encoding="utf-8")

# Example: Load HTML report

html_content = load_html_snapshot("2025-12-28", "09-15-00")

# Can be written to file, served via HTTP, or parsed with BeautifulSoup

Snapshot directories are created by generate_html_report in trendradar/report/generator.py. The filename timestamp uses hyphen separators (HH-MM-SS) rather than colons for filesystem compatibility.

Summary

  • TrendRadar output resides in the output/ directory with separate subdirectories for news/, rss/, txt/, html/, and meta/ data
  • SQLite databases in news/ and rss/ contain structured data with news_items/rss_items, platforms/rss_feeds, and crawl history tables
  • ParserService in mcp_server/services/parser_service.py provides the high-level API: read_all_titles_for_date(), get_available_dates(), and convert_crawl_results_to_news_data()
  • Snapshot files in txt/ and html/ preserve exact rendered output for specific crawl times, readable directly as text files
  • Data models NewsData and NewsItem in trendradar/storage/base.py structure the parsed output with rank history, URLs, and metadata

Frequently Asked Questions

What database format does TrendRadar use?

TrendRadar uses SQLite .db files for structured data storage. Each day generates separate databases for news and RSS data in output/news/YYYY-MM-DD.db and output/rss/YYYY-MM-DD.db. The schema includes relational tables for items, platforms, and crawl history.

How do I check which dates have available data?

Use ParserService.get_available_dates(db_type) to list all dates with existing databases. Pass "news" or "rss" as the db_type parameter. The method returns a list of date strings in YYYY-MM-DD format extracted from database filenames in the respective output directory.

Can I access historical snapshots without parsing SQLite?

Yes. TrendRadar saves plain-text and HTML snapshots in output/txt/YYYY-MM-DD/HH-MM-SS.txt and output/html/YYYY-MM-DD/HH-MM-SS.html. These files contain the exact rendered output from specific crawl times. Read them directly with standard file I/O—no database parsing required.

What is the difference between raw crawler output and NewsData?

Raw crawler output is a nested dictionary mapping source IDs to title dictionaries with properties like ranks, url, and mobileUrl. NewsData is a structured Pydantic model with typed fields: date, crawl_time, items (dict of NewsItem lists), id_to_name mapping, and failed_ids. Use convert_crawl_results_to_news_data() to transform raw output into the structured model.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →