How News Subscriptions and Automated Research Digests Work in Local Deep Research
Local Deep Research automatically runs scheduled searches through configurable news subscriptions, stores results as research history, and surfaces curated findings in a unified digest feed.
Local Deep Research provides a comprehensive news subscription system that transforms static search queries into self-maintaining research pipelines. This open-source framework (learningcircuit/local-deep-research) enables automated monitoring of topics through periodic search execution against configured engines like SerpAPI or DuckDuckGo. Results aggregate into a personalized automated research digest accessible via REST API endpoints.
Core Architecture Components
The system implements a four-layer architecture that separates web interface concerns from background processing logic.
API Layer
The Flask routes in src/local_deep_research/web/routes/news_routes.py expose RESTful endpoints for subscription management. Key routes include POST /api/news/subscriptions for creation, GET /api/news/subscriptions for listing, GET /api/news/feed for digest retrieval, and POST /api/news/feedback for voting. These endpoints handle authentication and delegate to the business logic layer.
Business Logic Layer
Located in src/local_deep_research/news/api.py, this layer implements create_subscription, update_subscription, delete_subscription, get_news_feed, and get_subscription_history. The create_subscription function builds NewsSubscription ORM objects and explicitly notifies the scheduler via _notify_scheduler_about_subscription_change("created") to trigger immediate pickup.
Subscription Model Layer
The abstract base class BaseSubscription in src/local_deep_research/news/subscription_manager/base_subscription.py defines the contract for all subscriptions. Concrete implementations SearchSubscription (search_subscription.py) and TopicSubscription (topic_subscription.py) implement generate_search_query() to produce query strings like "site:nytimes.com {topic}". The model tracks next_refresh, refresh_interval_minutes, is_active state, and provides lifecycle hooks: on_refresh_start(), on_refresh_success(), and on_refresh_error().
Scheduler and Storage Layer
NewsScheduler in src/local_deep_research/news/subscription_manager/scheduler.py runs as a background worker, executing run_once() on a timer (typically every minute). It queries SQLSubscriptionStorage (storage.py) to load active subscriptions via methods like list_by_user. The storage layer handles create, update_refresh_time, increment_stats, and delete operations against SQLite or PostgreSQL backends.
End-to-End Subscription Workflow
The platform converts a "search every 4 hours" definition into a continuous research pipeline through six distinct stages:
-
Subscription Creation: The Flask route invokes
news_api.create_subscription, which persists aNewsSubscriptionrecord (defined insrc/local_deep_research/database/models/news.py) viaSQLSubscriptionStorageand notifies the scheduler. -
Scheduler Detection:
NewsScheduler.run_once()loads all active subscriptions whereis_activeis true. For each subscription,subscription.should_refresh()evaluates whether the current UTC time exceeds the storednext_refreshtimestamp. -
Query Generation and Execution: Upon refresh trigger, the concrete subscription class executes
generate_search_query(). The scheduler forwards this query to the generic search-engine factory (src/local_deep_research/web_search_engines/...), executing against configured backends. Results transform intoResearchHistoryentries stored viaget_user_db_session(user_id). -
Post-Refresh Lifecycle Handling: Success invokes
subscription.on_refresh_success(results), incrementingrefresh_countviaincrement_statsand recalculatingnext_refresh. Failure triggerssubscription.on_refresh_error(exc), implementing exponential back-off and disabling the subscription after 10 consecutive failures. -
Digest Generation: The
GET /api/news/feedendpoint callsnews_api.get_news_feed, filteringResearchHistory(exposed insrc/local_deep_research/database/models/__init__.py) for entries with news metadata (generated_headline,is_news_search). The function returns up to 20 news cards containing headlines, summaries, and source links. -
User Interaction and Feedback: Users vote on individual news cards via
POST /api/news/feedback, which updates ratings in theUserRatingtable. Aggregated votes are retrieved viaget_votes_for_cardsto influence future rankings.
Working with the News API
Creating a Subscription
import requests
payload = {
"query": "quantum computing breakthroughs",
"type": "search",
"refresh_minutes": 180,
"name": "Quantum Updates",
"is_active": True,
"search_engine": "serper",
}
resp = requests.post(
"https://your-instance/api/news/subscriptions",
json=payload,
cookies={"session": "..."}
)
print(resp.json())
Retrieving the Automated Digest
resp = requests.get(
"https://your-instance/api/news/feed?limit=10",
cookies={"session": "..."}
)
feed = resp.json()["news_items"]
for card in feed:
print(f"{card['headline']} – {card['summary']}")
Updating Subscription Parameters
update_payload = {"refresh_minutes": 240}
resp = requests.patch(
"https://your-instance/api/news/subscriptions/abcd-1234",
json=update_payload,
cookies={"session": "..."}
)
Background Scheduler Implementation
The NewsScheduler orchestrates automated refreshes without user interaction.
from local_deep_research.news.subscription_manager.scheduler import get_news_scheduler
scheduler = get_news_scheduler()
while True:
scheduler.run_once()
time.sleep(60)
Internally, run_once() iterates through active subscriptions, validates refresh timing with should_refresh(), executes searches, and persists history records to the database.
Key Implementation Files
src/local_deep_research/web/routes/news_routes.py: Flask endpoints for subscription CRUD, feed access, and feedback submission.src/local_deep_research/news/api.py: Core business logic includingcreate_subscription,get_news_feed, andupdate_subscription.src/local_deep_research/news/subscription_manager/base_subscription.py: AbstractBaseSubscriptionclass with timing logic and error handling hooks.src/local_deep_research/news/subscription_manager/search_subscription.py: Concrete implementation for free-text search subscriptions.src/local_deep_research/news/subscription_manager/topic_subscription.py: Concrete implementation for topic-based subscriptions.src/local_deep_research/news/subscription_manager/scheduler.py: Background worker implementingNewsSchedulerandrun_once().src/local_deep_research/news/subscription_manager/storage.py:SQLSubscriptionStorageclass handling persistence and statistics.src/local_deep_research/database/models/news.py: ORM definitions forNewsSubscriptionandUserRatingtables.src/local_deep_research/database/models/__init__.py: ExposesResearchHistoryused by the feed generator.
Summary
- News subscriptions in Local Deep Research are persistent search definitions that execute on configurable intervals through
BaseSubscriptionsubclasses likeSearchSubscriptionandTopicSubscription. - The
NewsSchedulerbackground worker checksshould_refresh()every minute, executing queries through the generic search-engine wrapper and storing results asResearchHistoryentries. - Automated research digests surface via
GET /api/news/feed, which filters history for news-specific metadata and returns curated cards with headlines, summaries, and voting capabilities. - The system implements robust error handling with exponential back-off, disabling subscriptions after 10 consecutive failures via
on_refresh_error(). - User interactions including pause/resume operations and card voting persist through
SQLSubscriptionStorageand theUserRatingtable.
Frequently Asked Questions
How does the scheduler know when to run a subscription?
The NewsScheduler loads all active subscriptions from the database every minute and calls subscription.should_refresh(), which compares the current UTC timestamp against the next_refresh field stored in the subscription record.
What happens if a scheduled search fails?
The system invokes subscription.on_refresh_error(exc), which increments an internal error counter using increment_stats and applies exponential back-off to the next refresh time. After 10 consecutive failures, the subscription automatically sets is_active to false to prevent further resource consumption.
Can I use different search engines for different subscriptions?
Yes. Each subscription stores a search_engine parameter (e.g., "serper", "duckduckgo"). When the scheduler executes generate_search_query(), it passes the query to the search-engine factory configured for that specific subscription type, allowing per-subscription engine selection.
How does the research digest differ from regular search history?
While both store results in ResearchHistory, the news feed specifically filters for entries containing news metadata flags like generated_headline and is_news_search. The get_news_feed function transforms these into structured cards with voting capabilities via the UserRating table, whereas standard search history appears in the general research log without digest formatting.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →