How News Subscriptions and Automated Research Digests Work in Local Deep Research

Local Deep Research automatically runs scheduled searches through configurable news subscriptions, stores results as research history, and surfaces curated findings in a unified digest feed.

Local Deep Research provides a comprehensive news subscription system that transforms static search queries into self-maintaining research pipelines. This open-source framework (learningcircuit/local-deep-research) enables automated monitoring of topics through periodic search execution against configured engines like SerpAPI or DuckDuckGo. Results aggregate into a personalized automated research digest accessible via REST API endpoints.

Core Architecture Components

The system implements a four-layer architecture that separates web interface concerns from background processing logic.

API Layer

The Flask routes in src/local_deep_research/web/routes/news_routes.py expose RESTful endpoints for subscription management. Key routes include POST /api/news/subscriptions for creation, GET /api/news/subscriptions for listing, GET /api/news/feed for digest retrieval, and POST /api/news/feedback for voting. These endpoints handle authentication and delegate to the business logic layer.

Business Logic Layer

Located in src/local_deep_research/news/api.py, this layer implements create_subscription, update_subscription, delete_subscription, get_news_feed, and get_subscription_history. The create_subscription function builds NewsSubscription ORM objects and explicitly notifies the scheduler via _notify_scheduler_about_subscription_change("created") to trigger immediate pickup.

Subscription Model Layer

The abstract base class BaseSubscription in src/local_deep_research/news/subscription_manager/base_subscription.py defines the contract for all subscriptions. Concrete implementations SearchSubscription (search_subscription.py) and TopicSubscription (topic_subscription.py) implement generate_search_query() to produce query strings like "site:nytimes.com {topic}". The model tracks next_refresh, refresh_interval_minutes, is_active state, and provides lifecycle hooks: on_refresh_start(), on_refresh_success(), and on_refresh_error().

Scheduler and Storage Layer

NewsScheduler in src/local_deep_research/news/subscription_manager/scheduler.py runs as a background worker, executing run_once() on a timer (typically every minute). It queries SQLSubscriptionStorage (storage.py) to load active subscriptions via methods like list_by_user. The storage layer handles create, update_refresh_time, increment_stats, and delete operations against SQLite or PostgreSQL backends.

End-to-End Subscription Workflow

The platform converts a "search every 4 hours" definition into a continuous research pipeline through six distinct stages:

  1. Subscription Creation: The Flask route invokes news_api.create_subscription, which persists a NewsSubscription record (defined in src/local_deep_research/database/models/news.py) via SQLSubscriptionStorage and notifies the scheduler.

  2. Scheduler Detection: NewsScheduler.run_once() loads all active subscriptions where is_active is true. For each subscription, subscription.should_refresh() evaluates whether the current UTC time exceeds the stored next_refresh timestamp.

  3. Query Generation and Execution: Upon refresh trigger, the concrete subscription class executes generate_search_query(). The scheduler forwards this query to the generic search-engine factory (src/local_deep_research/web_search_engines/...), executing against configured backends. Results transform into ResearchHistory entries stored via get_user_db_session(user_id).

  4. Post-Refresh Lifecycle Handling: Success invokes subscription.on_refresh_success(results), incrementing refresh_count via increment_stats and recalculating next_refresh. Failure triggers subscription.on_refresh_error(exc), implementing exponential back-off and disabling the subscription after 10 consecutive failures.

  5. Digest Generation: The GET /api/news/feed endpoint calls news_api.get_news_feed, filtering ResearchHistory (exposed in src/local_deep_research/database/models/__init__.py) for entries with news metadata (generated_headline, is_news_search). The function returns up to 20 news cards containing headlines, summaries, and source links.

  6. User Interaction and Feedback: Users vote on individual news cards via POST /api/news/feedback, which updates ratings in the UserRating table. Aggregated votes are retrieved via get_votes_for_cards to influence future rankings.

Working with the News API

Creating a Subscription

import requests

payload = {
    "query": "quantum computing breakthroughs",
    "type": "search",
    "refresh_minutes": 180,
    "name": "Quantum Updates",
    "is_active": True,
    "search_engine": "serper",
}
resp = requests.post(
    "https://your-instance/api/news/subscriptions",
    json=payload,
    cookies={"session": "..."}
)
print(resp.json())

Retrieving the Automated Digest

resp = requests.get(
    "https://your-instance/api/news/feed?limit=10",
    cookies={"session": "..."}
)
feed = resp.json()["news_items"]
for card in feed:
    print(f"{card['headline']} – {card['summary']}")

Updating Subscription Parameters

update_payload = {"refresh_minutes": 240}
resp = requests.patch(
    "https://your-instance/api/news/subscriptions/abcd-1234",
    json=update_payload,
    cookies={"session": "..."}
)

Background Scheduler Implementation

The NewsScheduler orchestrates automated refreshes without user interaction.

from local_deep_research.news.subscription_manager.scheduler import get_news_scheduler

scheduler = get_news_scheduler()
while True:
    scheduler.run_once()
    time.sleep(60)

Internally, run_once() iterates through active subscriptions, validates refresh timing with should_refresh(), executes searches, and persists history records to the database.

Key Implementation Files

Summary

  • News subscriptions in Local Deep Research are persistent search definitions that execute on configurable intervals through BaseSubscription subclasses like SearchSubscription and TopicSubscription.
  • The NewsScheduler background worker checks should_refresh() every minute, executing queries through the generic search-engine wrapper and storing results as ResearchHistory entries.
  • Automated research digests surface via GET /api/news/feed, which filters history for news-specific metadata and returns curated cards with headlines, summaries, and voting capabilities.
  • The system implements robust error handling with exponential back-off, disabling subscriptions after 10 consecutive failures via on_refresh_error().
  • User interactions including pause/resume operations and card voting persist through SQLSubscriptionStorage and the UserRating table.

Frequently Asked Questions

How does the scheduler know when to run a subscription?

The NewsScheduler loads all active subscriptions from the database every minute and calls subscription.should_refresh(), which compares the current UTC timestamp against the next_refresh field stored in the subscription record.

What happens if a scheduled search fails?

The system invokes subscription.on_refresh_error(exc), which increments an internal error counter using increment_stats and applies exponential back-off to the next refresh time. After 10 consecutive failures, the subscription automatically sets is_active to false to prevent further resource consumption.

Can I use different search engines for different subscriptions?

Yes. Each subscription stores a search_engine parameter (e.g., "serper", "duckduckgo"). When the scheduler executes generate_search_query(), it passes the query to the search-engine factory configured for that specific subscription type, allowing per-subscription engine selection.

How does the research digest differ from regular search history?

While both store results in ResearchHistory, the news feed specifically filters for entries containing news metadata flags like generated_headline and is_news_search. The get_news_feed function transforms these into structured cards with voting capabilities via the UserRating table, whereas standard search history appears in the general research log without digest formatting.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →