How Cross-Platform Convergence Detection Works in last30days-skill
Cross-platform convergence detection in last30days-skill identifies identical stories across Reddit, X, Hacker News, and other platforms by computing a hybrid Jaccard similarity score between normalized text snippets and bidirectionally linking items that exceed a 0.40 threshold.
The last30days-skill repository aggregates trending posts from disparate social platforms into a unified feed. To prevent the same story from appearing multiple times under different sources, the library implements cross-platform convergence detection that algorithmically matches content using token and character-level similarity metrics.
Text Normalization for Cross-Source Comparison
Before calculating similarity, the system normalizes content from each platform to create comparable text representations. In scripts/lib/dedupe.py, the _get_cross_source_text() function handles platform-specific preprocessing:
- X/TikTok/Instagram posts are truncated to 100 characters to remove platform-specific fluff
- Hacker News titles strip prefixes like "Show HN:" or "Ask HN:"
- Other items use standard text extraction via
get_item_text()
# scripts/lib/dedupe.py
def _get_cross_source_text(item: AnyItem) -> str:
if isinstance(item, schema.XItem):
return item.text[:100]
if isinstance(item, schema.TikTokItem):
return item.text[:100]
if isinstance(item, schema.HackerNewsItem):
title = item.title
if title.startswith("Show HN:"):
title = title[8:].strip()
return title
return get_item_text(item)
This normalization ensures that the same story posted with slightly different formatting across platforms can be meaningfully compared.
Hybrid Similarity Scoring
The detection relies on two complementary similarity algorithms implemented in scripts/lib/dedupe.py. The _hybrid_similarity() function returns the maximum of token-level Jaccard similarity and character trigram overlap, making the system robust to both semantic matches and noisy text variations.
Token-Level Comparison
The _token_jaccard() function tokenizes text by removing punctuation, converting to lowercase, and filtering out stopwords from the STOPWORDS set before computing Jaccard overlap:
# scripts/lib/dedupe.py
def _token_jaccard(text_a: str, text_b: str) -> float:
# Tokenization removes punctuation, lowercases, and filters stopwords
tokens_a = set([w for w in tokenize(text_a) if w not in STOPWORDS])
tokens_b = set([w for w in tokenize(text_b) if w not in STOPWORDS])
return jaccard_similarity(tokens_a, tokens_b)
Character Trigram Analysis
For catching partial matches and handling text variations, get_ngrams() extracts overlapping character trigrams (3-grams):
# scripts/lib/dedupe.py
def get_ngrams(text: str, n: int = 3) -> Set[str]:
# Returns set of character n-grams
...
def _hybrid_similarity(text_a: str, text_b: str) -> float:
trigram_sim = jaccard_similarity(get_ngrams(text_a), get_ngrams(text_b))
token_sim = _token_jaccard(text_a, text_b)
return max(trigram_sim, token_sim)
By taking the maximum of these two scores, the system catches both short keyword matches via tokens and longer substring matches via trigrams.
Cross-Source Linking Algorithm
The cross_source_link() function in scripts/lib/dedupe.py performs the actual convergence detection by comparing every item against every other item from different platforms:
- Excludes intra-source pairs: The algorithm skips comparisons where both items are from the same platform (e.g., Reddit vs Reddit), as deduplication within a single source is handled separately
- Pre-computes normalized text: Each item's comparable text is generated once to avoid redundant processing
- Bidirectional linking: When similarity exceeds the threshold (default 0.40), both items receive references to each other's IDs in their
cross_refslists
# scripts/lib/dedupe.py
def cross_source_link(*source_lists: List[AnyItem], threshold: float = 0.40) -> None:
all_items = [item for sublist in source_lists for item in sublist]
texts = [_get_cross_source_text(item) for item in all_items]
for i in range(len(all_items)):
for j in range(i + 1, len(all_items)):
if type(all_items[i]) is type(all_items[j]):
continue # Skip same source type
similarity = _hybrid_similarity(texts[i], texts[j])
if similarity >= threshold:
# Bidirectional cross-reference
if all_items[j].id not in all_items[i].cross_refs:
all_items[i].cross_refs.append(all_items[j].id)
if all_items[i].id not in all_items[j].cross_refs:
all_items[j].cross_refs.append(all_items[i].id)
Schema Support for Cross-References
Data persistence for converged stories relies on the cross_refs field defined in scripts/lib/schema.py. Every platform-specific item class (e.g., RedditItem, XItem, WebSearchItem) includes:
# scripts/lib/schema.py
@dataclass
class RedditItem:
cross_refs: List[str] = field(default_factory=list)
def to_dict(self) -> dict:
d = {...}
if self.cross_refs:
d['cross_refs'] = self.cross_refs
return d
When to_dict() serializes an item with matched convergences, the cross_refs array contains the IDs of related items from other platforms, enabling downstream rendering to display "also on Hacker News" or "also on X" indicators.
Practical Usage Examples
Linking Reddit and Hacker News Items
from scripts.lib import reddit, hackernews, dedupe
reddit_items = reddit.fetch_latest(...) # List[RedditItem]
hn_items = hackernews.fetch_latest(...) # List[HackerNewsItem]
# Perform cross-platform convergence detection
dedupe.cross_source_link(reddit_items, hn_items)
# Access cross-references
print(reddit_items[0].cross_refs) # e.g., ['HN123']
print(hn_items[0].cross_refs) # e.g., ['R456']
Adjusting Sensitivity
Increase the threshold to reduce false positives when dealing with generic titles:
# Stricter matching with 0.55 threshold
dedupe.cross_source_link(
reddit_items,
hn_items,
youtube_items,
threshold=0.55
)
Serializing Converged Items
import json
# to_dict() includes cross_refs when present
output = reddit_items[0].to_dict()
print(json.dumps(output, indent=2))
# Output includes: "cross_refs": ["HN789", "X12345"]
Summary
- Cross-platform convergence detection identifies duplicate stories across Reddit, X, Hacker News, YouTube, and TikTok by comparing normalized text representations
- The hybid similarity algorithm combines token Jaccard and character trigram scores to maximize detection accuracy across both short and long-form content
- Text normalization in
_get_cross_source_text()handles platform-specific formatting like "Show HN:" prefixes and truncates microblogging content to 100 characters - Bidirectional linking updates
cross_refsfields in both matched items when the similarity exceeds the configurable threshold (default 0.40) - The implementation resides primarily in
scripts/lib/dedupe.pywith schema support inscripts/lib/schema.py
Frequently Asked Questions
What is the optimal similarity threshold for convergence detection?
The default threshold of 0.40 balances precision and recall for general news aggregation. Increase it to 0.55 or higher when processing generic topics prone to false matches, or lower it to 0.30 for niche technical content where exact wording varies significantly across platforms.
Why does last30days-skill use both tokens and trigrams for similarity?
Token-level Jaccard excels at identifying semantic matches in longer text, while character trigrams catch substring similarities and handle typos or formatting variations. By taking the maximum of both scores in _hybrid_similarity(), the detector maintains high accuracy across diverse content lengths and styles.
How does the system prevent linking items from the same source?
The cross_source_link() function explicitly checks if type(all_items[i]) is type(all_items[j]) and skips these pairs. Intra-platform deduplication is handled separately by other pipeline stages, ensuring that cross-platform detection only identifies convergence between different services.
Where are cross-platform links stored after detection?
Matched item IDs populate the cross_refs: List[str] field defined in scripts/lib/schema.py. When serialized via to_dict(), these references appear in the JSON output as an array of foreign IDs, allowing downstream consumers to render "also appearing on" indicators for converged stories.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →