How Cross-Platform Convergence Detection Works in last30days-skill

Cross-platform convergence detection in last30days-skill identifies identical stories across Reddit, X, Hacker News, and other platforms by computing a hybrid Jaccard similarity score between normalized text snippets and bidirectionally linking items that exceed a 0.40 threshold.

The last30days-skill repository aggregates trending posts from disparate social platforms into a unified feed. To prevent the same story from appearing multiple times under different sources, the library implements cross-platform convergence detection that algorithmically matches content using token and character-level similarity metrics.

Text Normalization for Cross-Source Comparison

Before calculating similarity, the system normalizes content from each platform to create comparable text representations. In scripts/lib/dedupe.py, the _get_cross_source_text() function handles platform-specific preprocessing:

  • X/TikTok/Instagram posts are truncated to 100 characters to remove platform-specific fluff
  • Hacker News titles strip prefixes like "Show HN:" or "Ask HN:"
  • Other items use standard text extraction via get_item_text()

# scripts/lib/dedupe.py

def _get_cross_source_text(item: AnyItem) -> str:
    if isinstance(item, schema.XItem):
        return item.text[:100]
    if isinstance(item, schema.TikTokItem):
        return item.text[:100]
    if isinstance(item, schema.HackerNewsItem):
        title = item.title
        if title.startswith("Show HN:"):
            title = title[8:].strip()
        return title
    return get_item_text(item)

This normalization ensures that the same story posted with slightly different formatting across platforms can be meaningfully compared.

Hybrid Similarity Scoring

The detection relies on two complementary similarity algorithms implemented in scripts/lib/dedupe.py. The _hybrid_similarity() function returns the maximum of token-level Jaccard similarity and character trigram overlap, making the system robust to both semantic matches and noisy text variations.

Token-Level Comparison

The _token_jaccard() function tokenizes text by removing punctuation, converting to lowercase, and filtering out stopwords from the STOPWORDS set before computing Jaccard overlap:


# scripts/lib/dedupe.py

def _token_jaccard(text_a: str, text_b: str) -> float:
    # Tokenization removes punctuation, lowercases, and filters stopwords

    tokens_a = set([w for w in tokenize(text_a) if w not in STOPWORDS])
    tokens_b = set([w for w in tokenize(text_b) if w not in STOPWORDS])
    return jaccard_similarity(tokens_a, tokens_b)

Character Trigram Analysis

For catching partial matches and handling text variations, get_ngrams() extracts overlapping character trigrams (3-grams):


# scripts/lib/dedupe.py

def get_ngrams(text: str, n: int = 3) -> Set[str]:
    # Returns set of character n-grams

    ...

def _hybrid_similarity(text_a: str, text_b: str) -> float:
    trigram_sim = jaccard_similarity(get_ngrams(text_a), get_ngrams(text_b))
    token_sim   = _token_jaccard(text_a, text_b)
    return max(trigram_sim, token_sim)

By taking the maximum of these two scores, the system catches both short keyword matches via tokens and longer substring matches via trigrams.

Cross-Source Linking Algorithm

The cross_source_link() function in scripts/lib/dedupe.py performs the actual convergence detection by comparing every item against every other item from different platforms:

  1. Excludes intra-source pairs: The algorithm skips comparisons where both items are from the same platform (e.g., Reddit vs Reddit), as deduplication within a single source is handled separately
  2. Pre-computes normalized text: Each item's comparable text is generated once to avoid redundant processing
  3. Bidirectional linking: When similarity exceeds the threshold (default 0.40), both items receive references to each other's IDs in their cross_refs lists

# scripts/lib/dedupe.py

def cross_source_link(*source_lists: List[AnyItem], threshold: float = 0.40) -> None:
    all_items = [item for sublist in source_lists for item in sublist]
    texts = [_get_cross_source_text(item) for item in all_items]
    
    for i in range(len(all_items)):
        for j in range(i + 1, len(all_items)):
            if type(all_items[i]) is type(all_items[j]):
                continue  # Skip same source type

            similarity = _hybrid_similarity(texts[i], texts[j])
            if similarity >= threshold:
                # Bidirectional cross-reference

                if all_items[j].id not in all_items[i].cross_refs:
                    all_items[i].cross_refs.append(all_items[j].id)
                if all_items[i].id not in all_items[j].cross_refs:
                    all_items[j].cross_refs.append(all_items[i].id)

Schema Support for Cross-References

Data persistence for converged stories relies on the cross_refs field defined in scripts/lib/schema.py. Every platform-specific item class (e.g., RedditItem, XItem, WebSearchItem) includes:


# scripts/lib/schema.py

@dataclass
class RedditItem:
    cross_refs: List[str] = field(default_factory=list)
    
    def to_dict(self) -> dict:
        d = {...}
        if self.cross_refs:
            d['cross_refs'] = self.cross_refs
        return d

When to_dict() serializes an item with matched convergences, the cross_refs array contains the IDs of related items from other platforms, enabling downstream rendering to display "also on Hacker News" or "also on X" indicators.

Practical Usage Examples

Linking Reddit and Hacker News Items

from scripts.lib import reddit, hackernews, dedupe

reddit_items = reddit.fetch_latest(...)  # List[RedditItem]

hn_items = hackernews.fetch_latest(...)  # List[HackerNewsItem]

# Perform cross-platform convergence detection

dedupe.cross_source_link(reddit_items, hn_items)

# Access cross-references

print(reddit_items[0].cross_refs)  # e.g., ['HN123']

print(hn_items[0].cross_refs)      # e.g., ['R456']

Adjusting Sensitivity

Increase the threshold to reduce false positives when dealing with generic titles:


# Stricter matching with 0.55 threshold

dedupe.cross_source_link(
    reddit_items, 
    hn_items, 
    youtube_items, 
    threshold=0.55
)

Serializing Converged Items

import json

# to_dict() includes cross_refs when present

output = reddit_items[0].to_dict()
print(json.dumps(output, indent=2))

# Output includes: "cross_refs": ["HN789", "X12345"]

Summary

  • Cross-platform convergence detection identifies duplicate stories across Reddit, X, Hacker News, YouTube, and TikTok by comparing normalized text representations
  • The hybid similarity algorithm combines token Jaccard and character trigram scores to maximize detection accuracy across both short and long-form content
  • Text normalization in _get_cross_source_text() handles platform-specific formatting like "Show HN:" prefixes and truncates microblogging content to 100 characters
  • Bidirectional linking updates cross_refs fields in both matched items when the similarity exceeds the configurable threshold (default 0.40)
  • The implementation resides primarily in scripts/lib/dedupe.py with schema support in scripts/lib/schema.py

Frequently Asked Questions

What is the optimal similarity threshold for convergence detection?

The default threshold of 0.40 balances precision and recall for general news aggregation. Increase it to 0.55 or higher when processing generic topics prone to false matches, or lower it to 0.30 for niche technical content where exact wording varies significantly across platforms.

Why does last30days-skill use both tokens and trigrams for similarity?

Token-level Jaccard excels at identifying semantic matches in longer text, while character trigrams catch substring similarities and handle typos or formatting variations. By taking the maximum of both scores in _hybrid_similarity(), the detector maintains high accuracy across diverse content lengths and styles.

How does the system prevent linking items from the same source?

The cross_source_link() function explicitly checks if type(all_items[i]) is type(all_items[j]) and skips these pairs. Intra-platform deduplication is handled separately by other pipeline stages, ensuring that cross-platform detection only identifies convergence between different services.

Matched item IDs populate the cross_refs: List[str] field defined in scripts/lib/schema.py. When serialized via to_dict(), these references appear in the JSON output as an array of foreign IDs, allowing downstream consumers to render "also appearing on" indicators for converged stories.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →