# How Cross-Platform Convergence Detection Works in last30days-skill

> Learn how cross-platform convergence detection identifies identical stories across Reddit X Hacker News and more using hybrid Jaccard similarity in last30days-skill

- Repository: [Matt Van Horn/last30days-skill](https://github.com/mvanhorn/last30days-skill)
- Tags: internals
- Published: 2026-03-25

---

**Cross-platform convergence detection in last30days-skill identifies identical stories across Reddit, X, Hacker News, and other platforms by computing a hybrid Jaccard similarity score between normalized text snippets and bidirectionally linking items that exceed a 0.40 threshold.**

The last30days-skill repository aggregates trending posts from disparate social platforms into a unified feed. To prevent the same story from appearing multiple times under different sources, the library implements **cross-platform convergence detection** that algorithmically matches content using token and character-level similarity metrics.

## Text Normalization for Cross-Source Comparison

Before calculating similarity, the system normalizes content from each platform to create comparable text representations. In [`scripts/lib/dedupe.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/dedupe.py), the `_get_cross_source_text()` function handles platform-specific preprocessing:

- **X/TikTok/Instagram posts** are truncated to 100 characters to remove platform-specific fluff
- **Hacker News titles** strip prefixes like "Show HN:" or "Ask HN:" 
- Other items use standard text extraction via `get_item_text()`

```python

# scripts/lib/dedupe.py

def _get_cross_source_text(item: AnyItem) -> str:
    if isinstance(item, schema.XItem):
        return item.text[:100]
    if isinstance(item, schema.TikTokItem):
        return item.text[:100]
    if isinstance(item, schema.HackerNewsItem):
        title = item.title
        if title.startswith("Show HN:"):
            title = title[8:].strip()
        return title
    return get_item_text(item)

```

This normalization ensures that the same story posted with slightly different formatting across platforms can be meaningfully compared.

## Hybrid Similarity Scoring

The detection relies on two complementary similarity algorithms implemented in [`scripts/lib/dedupe.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/dedupe.py). The `_hybrid_similarity()` function returns the maximum of **token-level Jaccard similarity** and **character trigram overlap**, making the system robust to both semantic matches and noisy text variations.

### Token-Level Comparison

The `_token_jaccard()` function tokenizes text by removing punctuation, converting to lowercase, and filtering out stopwords from the `STOPWORDS` set before computing Jaccard overlap:

```python

# scripts/lib/dedupe.py

def _token_jaccard(text_a: str, text_b: str) -> float:
    # Tokenization removes punctuation, lowercases, and filters stopwords

    tokens_a = set([w for w in tokenize(text_a) if w not in STOPWORDS])
    tokens_b = set([w for w in tokenize(text_b) if w not in STOPWORDS])
    return jaccard_similarity(tokens_a, tokens_b)

```

### Character Trigram Analysis

For catching partial matches and handling text variations, `get_ngrams()` extracts overlapping character trigrams (3-grams):

```python

# scripts/lib/dedupe.py

def get_ngrams(text: str, n: int = 3) -> Set[str]:
    # Returns set of character n-grams

    ...

def _hybrid_similarity(text_a: str, text_b: str) -> float:
    trigram_sim = jaccard_similarity(get_ngrams(text_a), get_ngrams(text_b))
    token_sim   = _token_jaccard(text_a, text_b)
    return max(trigram_sim, token_sim)

```

By taking the **maximum** of these two scores, the system catches both short keyword matches via tokens and longer substring matches via trigrams.

## Cross-Source Linking Algorithm

The `cross_source_link()` function in [`scripts/lib/dedupe.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/dedupe.py) performs the actual convergence detection by comparing every item against every other item from different platforms:

1. **Excludes intra-source pairs**: The algorithm skips comparisons where both items are from the same platform (e.g., Reddit vs Reddit), as deduplication within a single source is handled separately
2. **Pre-computes normalized text**: Each item's comparable text is generated once to avoid redundant processing
3. **Bidirectional linking**: When similarity exceeds the threshold (default **0.40**), both items receive references to each other's IDs in their `cross_refs` lists

```python

# scripts/lib/dedupe.py

def cross_source_link(*source_lists: List[AnyItem], threshold: float = 0.40) -> None:
    all_items = [item for sublist in source_lists for item in sublist]
    texts = [_get_cross_source_text(item) for item in all_items]
    
    for i in range(len(all_items)):
        for j in range(i + 1, len(all_items)):
            if type(all_items[i]) is type(all_items[j]):
                continue  # Skip same source type

            similarity = _hybrid_similarity(texts[i], texts[j])
            if similarity >= threshold:
                # Bidirectional cross-reference

                if all_items[j].id not in all_items[i].cross_refs:
                    all_items[i].cross_refs.append(all_items[j].id)
                if all_items[i].id not in all_items[j].cross_refs:
                    all_items[j].cross_refs.append(all_items[i].id)

```

## Schema Support for Cross-References

Data persistence for converged stories relies on the `cross_refs` field defined in [`scripts/lib/schema.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/schema.py). Every platform-specific item class (e.g., `RedditItem`, `XItem`, `WebSearchItem`) includes:

```python

# scripts/lib/schema.py

@dataclass
class RedditItem:
    cross_refs: List[str] = field(default_factory=list)
    
    def to_dict(self) -> dict:
        d = {...}
        if self.cross_refs:
            d['cross_refs'] = self.cross_refs
        return d

```

When `to_dict()` serializes an item with matched convergences, the `cross_refs` array contains the IDs of related items from other platforms, enabling downstream rendering to display "also on Hacker News" or "also on X" indicators.

## Practical Usage Examples

### Linking Reddit and Hacker News Items

```python
from scripts.lib import reddit, hackernews, dedupe

reddit_items = reddit.fetch_latest(...)  # List[RedditItem]

hn_items = hackernews.fetch_latest(...)  # List[HackerNewsItem]

# Perform cross-platform convergence detection

dedupe.cross_source_link(reddit_items, hn_items)

# Access cross-references

print(reddit_items[0].cross_refs)  # e.g., ['HN123']

print(hn_items[0].cross_refs)      # e.g., ['R456']

```

### Adjusting Sensitivity

Increase the threshold to reduce false positives when dealing with generic titles:

```python

# Stricter matching with 0.55 threshold

dedupe.cross_source_link(
    reddit_items, 
    hn_items, 
    youtube_items, 
    threshold=0.55
)

```

### Serializing Converged Items

```python
import json

# to_dict() includes cross_refs when present

output = reddit_items[0].to_dict()
print(json.dumps(output, indent=2))

# Output includes: "cross_refs": ["HN789", "X12345"]

```

## Summary

- **Cross-platform convergence detection** identifies duplicate stories across Reddit, X, Hacker News, YouTube, and TikTok by comparing normalized text representations
- The **hybid similarity algorithm** combines token Jaccard and character trigram scores to maximize detection accuracy across both short and long-form content
- **Text normalization** in `_get_cross_source_text()` handles platform-specific formatting like "Show HN:" prefixes and truncates microblogging content to 100 characters
- **Bidirectional linking** updates `cross_refs` fields in both matched items when the similarity exceeds the configurable threshold (default 0.40)
- The implementation resides primarily in [`scripts/lib/dedupe.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/dedupe.py) with schema support in [`scripts/lib/schema.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/schema.py)

## Frequently Asked Questions

### What is the optimal similarity threshold for convergence detection?

The default threshold of **0.40** balances precision and recall for general news aggregation. Increase it to **0.55** or higher when processing generic topics prone to false matches, or lower it to **0.30** for niche technical content where exact wording varies significantly across platforms.

### Why does last30days-skill use both tokens and trigrams for similarity?

Token-level Jaccard excels at identifying semantic matches in longer text, while character trigrams catch substring similarities and handle typos or formatting variations. By taking the **maximum** of both scores in `_hybrid_similarity()`, the detector maintains high accuracy across diverse content lengths and styles.

### How does the system prevent linking items from the same source?

The `cross_source_link()` function explicitly checks `if type(all_items[i]) is type(all_items[j])` and skips these pairs. Intra-platform deduplication is handled separately by other pipeline stages, ensuring that cross-platform detection only identifies convergence between different services.

### Where are cross-platform links stored after detection?

Matched item IDs populate the `cross_refs: List[str]` field defined in [`scripts/lib/schema.py`](https://github.com/mvanhorn/last30days-skill/blob/main/scripts/lib/schema.py). When serialized via `to_dict()`, these references appear in the JSON output as an array of foreign IDs, allowing downstream consumers to render "also appearing on" indicators for converged stories.