# How to Parse TrendRadar Output: SQLite, Snapshots, and Python API Guide

> Learn how to parse TrendRadar output using SQLite, text/HTML snapshots, or the Python API. Access news and RSS data efficiently from the sansan0/TrendRadar repo.

- Repository: [sansan/TrendRadar](https://github.com/sansan0/TrendRadar)
- Tags: how-to-guide
- Published: 2026-04-22

---

**Parse TrendRadar output by reading SQLite `.db` files from the `output/news/` or `output/rss/` directories, using `ParserService` for high-level access, or loading plain-text and HTML snapshots from `output/txt/` and `output/html/` for historic views.**

TrendRadar is an open-source trend monitoring system that stores daily crawl results in a structured, date-based file hierarchy. Understanding how to parse this output is essential for custom analytics, reporting, or integration with external systems. This guide covers the complete output format, SQLite schemas, and the Python API provided by the [`mcp_server/services/parser_service.py`](https://github.com/sansan0/TrendRadar/blob/main/mcp_server/services/parser_service.py) module.

## TrendRadar Output Directory Structure

TrendRadar organizes all output under a top-level `output/` directory with five primary subdirectories:

```

output/
├─ news/                # news data, one SQLite DB per day

│   └─ 2025-12-28.db
├─ rss/                 # RSS data, one SQLite DB per day

│   └─ 2025-12-28.db
├─ txt/                 # plain-text snapshots (date-folder → *.txt)

│   └─ 2025-12-28/
│       └─ 12-34-56.txt
├─ html/                # rendered HTML snapshots (date-folder → *.html)

│   └─ 2025-12-28/
│       └─ 12-34-56.html
└─ meta/                # auxiliary meta files (e.g. doctor_report.json)

```

Each data type (`news`, `rss`) receives its own SQLite database per day. Snapshot files preserve the exact rendered output at specific crawl times, useful for auditing or display purposes.

## SQLite Database Schema for News and RSS Data

The SQLite `.db` files follow a relational schema designed for efficient trend tracking and historical analysis.

### News Database Schema

| Table | Purpose | Key Columns |
|-------|---------|-------------|
| `news_items` | One row per headline | `id`, `platform_id`, `title`, `rank`, `url`, `mobile_url` |
| `platforms` | Maps `platform_id` to human-readable name | `id`, `name` |
| `rank_history` | Historic rank changes per `news_item_id` | `news_item_id`, `rank`, `timestamp` |
| `crawl_records` | Timestamps of each crawl run | `id`, `started_at`, `completed_at` |

### RSS Database Schema

The RSS schema mirrors the news structure with renamed tables:

| Table | RSS Equivalent |
|-------|----------------|
| `rss_items` | `news_items` |
| `rss_feeds` | `platforms` |
| `rss_crawl_records` | `crawl_records` |

The reading logic is implemented in `ParserService._read_news_from_sqlite` and `ParserService._read_rss_from_sqlite` within [`mcp_server/services/parser_service.py`](https://github.com/sansan0/TrendRadar/blob/main/mcp_server/services/parser_service.py).

## Using ParserService to Parse TrendRadar Output

The `ParserService` class in [`mcp_server/services/parser_service.py`](https://github.com/sansan0/TrendRadar/blob/main/mcp_server/services/parser_service.py) provides the high-level API for accessing TrendRadar output without manual SQL queries.

### Initializing ParserService

```python
from mcp_server.services.parser_service import ParserService

# Auto-detect project root and initialize cache

parser = ParserService()

```

The constructor accepts an optional `project_root` parameter if your working directory differs from the repository structure.

### Reading Titles for a Specific Date

```python
from datetime import datetime, timedelta

yesterday = datetime.now() - timedelta(days=1)

# Read all news titles from yesterday

titles, id_name_map, timestamps = parser.read_all_titles_for_date(
    date=yesterday,
    db_type="news"  # or "rss"

)

print(f"Found {len(titles)} platforms with news on {yesterday.date()}")
for src_id, item_dict in titles.items():
    platform_name = id_name_map.get(src_id, src_id)
    print(f"📰 {platform_name} – {len(item_dict)} headlines")

```

The `read_all_titles_for_date` method handles:
- Cache management for repeated access
- SQLite database opening and closing
- Fallback to `None` if no data exists for the requested date

### Listing Available Dates

```python

# Get all dates with news data

available_dates = parser.get_available_dates("news")
print(f"News data available for {len(available_dates)} days")

# Check specific date availability

target_date = "2025-12-28"
if target_date in available_dates:
    print(f"Data available for {target_date}")

```

Dates are returned as strings in `YYYY-MM-DD` format, matching the database filenames.

## Converting Raw Crawler Output to Structured Data

When working with raw crawler results rather than stored databases, use `convert_crawl_results_to_news_data` from [`trendradar/storage/base.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/storage/base.py).

```python
from mcp_server.services.parser_service import convert_crawl_results_to_news_data

# Raw crawler output format: {source_id: {title: {properties}}}

raw_crawler_output = {
    "newsnow": {
        "Example headline": {
            "ranks": [1, 2],
            "url": "https://example.com/article",
            "mobileUrl": ""
        },
        "Another trending story": {
            "ranks": [3],
            "url": "https://example.com/other",
            "mobileUrl": "https://m.example.com/other"
        }
    }
}

id_to_name = {"newsnow": "NewsNow"}
failed_ids = []  # Empty list if no crawl failures

news_data = convert_crawl_results_to_news_data(
    results=raw_crawler_output,
    id_to_name=id_to_name,
    failed_ids=failed_ids,
    crawl_time="09:15",
    crawl_date="2025-12-28"
)

# Access structured data

first_item = news_data.items["newsnow"][0]
print(f"Title: {first_item.title}")
print(f"Source: {first_item.source_name}")
print(f"Current rank: {first_item.rank}")
print(f"Rank history: {first_item.ranks}")

```

The resulting `NewsData` object contains:
- `date` and `crawl_time`: When the data was collected
- `items`: Dictionary mapping `source_id` to lists of `NewsItem` objects
- `id_to_name`: Mapping of source IDs to display names
- `failed_ids`: List of sources that failed to crawl

Each `NewsItem` includes `title`, `source_id`, `source_name`, `rank`, `url`, `mobile_url`, `crawl_time`, `ranks` (history), `first_time`, `last_time`, and `count`.

## Parsing Historic Snapshot Files

For the exact rendered output shown to users at a specific time, read the snapshot files directly from `output/txt/` or `output/html/`.

### Loading Text Snapshots

```python
from pathlib import Path

def load_text_snapshot(date: str, time: str) -> str:
    """
    Load a plain-text snapshot.
    Date format: YYYY-MM-DD
    Time format: HH-MM-SS (as stored in filename)
    """
    path = Path("output") / "txt" / date / f"{time}.txt"
    return path.read_text(encoding="utf-8")

# Example: Load snapshot from 2025-12-28 at 09:15:00

text_content = load_text_snapshot("2025-12-28", "09-15-00")
print(text_content)

```

### Loading HTML Snapshots

```python
from pathlib import Path

def load_html_snapshot(date: str, time: str) -> str:
    """
    Load a rendered HTML snapshot.
    """
    path = Path("output") / "html" / date / f"{time}.html"
    return path.read_text(encoding="utf-8")

# Example: Load HTML report

html_content = load_html_snapshot("2025-12-28", "09-15-00")

# Can be written to file, served via HTTP, or parsed with BeautifulSoup

```

Snapshot directories are created by `generate_html_report` in [`trendradar/report/generator.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/report/generator.py). The filename timestamp uses hyphen separators (`HH-MM-SS`) rather than colons for filesystem compatibility.

## Summary

- **TrendRadar output** resides in the `output/` directory with separate subdirectories for `news/`, `rss/`, `txt/`, `html/`, and `meta/` data
- **SQLite databases** in `news/` and `rss/` contain structured data with `news_items`/`rss_items`, `platforms`/`rss_feeds`, and crawl history tables
- **ParserService** in [`mcp_server/services/parser_service.py`](https://github.com/sansan0/TrendRadar/blob/main/mcp_server/services/parser_service.py) provides the high-level API: `read_all_titles_for_date()`, `get_available_dates()`, and `convert_crawl_results_to_news_data()`
- **Snapshot files** in `txt/` and `html/` preserve exact rendered output for specific crawl times, readable directly as text files
- **Data models** `NewsData` and `NewsItem` in [`trendradar/storage/base.py`](https://github.com/sansan0/TrendRadar/blob/main/trendradar/storage/base.py) structure the parsed output with rank history, URLs, and metadata

## Frequently Asked Questions

### What database format does TrendRadar use?

TrendRadar uses **SQLite** `.db` files for structured data storage. Each day generates separate databases for news and RSS data in `output/news/YYYY-MM-DD.db` and `output/rss/YYYY-MM-DD.db`. The schema includes relational tables for items, platforms, and crawl history.

### How do I check which dates have available data?

Use `ParserService.get_available_dates(db_type)` to list all dates with existing databases. Pass `"news"` or `"rss"` as the `db_type` parameter. The method returns a list of date strings in `YYYY-MM-DD` format extracted from database filenames in the respective output directory.

### Can I access historical snapshots without parsing SQLite?

Yes. TrendRadar saves plain-text and HTML snapshots in [`output/txt/YYYY-MM-DD/HH-MM-SS.txt`](https://github.com/sansan0/TrendRadar/blob/main/output/txt/YYYY-MM-DD/HH-MM-SS.txt) and [`output/html/YYYY-MM-DD/HH-MM-SS.html`](https://github.com/sansan0/TrendRadar/blob/main/output/html/YYYY-MM-DD/HH-MM-SS.html). These files contain the exact rendered output from specific crawl times. Read them directly with standard file I/O—no database parsing required.

### What is the difference between raw crawler output and NewsData?

Raw crawler output is a nested dictionary mapping source IDs to title dictionaries with properties like `ranks`, `url`, and `mobileUrl`. `NewsData` is a structured Pydantic model with typed fields: `date`, `crawl_time`, `items` (dict of `NewsItem` lists), `id_to_name` mapping, and `failed_ids`. Use `convert_crawl_results_to_news_data()` to transform raw output into the structured model.