What Is PreviouslySeenPostsFilter? Timeline Deduplication in X-Algorithm

PreviouslySeenPostsFilter removes candidate posts that a user has already encountered by checking seen_ids and Bloom filters, ensuring the Home-Mixer timeline surfaces only fresh content.

The PreviouslySeenPostsFilter is a critical deduplication component in the xai-org/x-algorithm repository's Home-Mixer pipeline. It operates early in the candidate generation workflow to prevent duplicate content from appearing in a user's timeline. By filtering out previously seen posts before ranking and mixing stages, the filter maintains content freshness and improves user experience.

How PreviouslySeenPostsFilter Works

The filter implements a strict exclusion policy based on historical user exposure. It accepts a ScoredPostsQuery containing the user's viewing history and a vector of PostCandidate objects, then partitions the candidates into fresh and previously viewed sets.

Input Parameters and Data Structures

According to the source code in home-mixer/models/query.rs, the filter receives:

  • ScoredPostsQuery – Contains the user's seen_ids (explicitly viewed tweet IDs) and optional Bloom filter entries for memory-efficient membership testing
  • PostCandidate vector – The batch of candidate posts generated by earlier pipeline stages, defined in home-mixer/models/candidate.rs

These inputs provide the filter with both the candidate pool and the historical context needed to make exclusion decisions.

The Deduplication Logic

Inside home-mixer/filters/previously_seen_posts_filter.rs, the filter executes the following steps:

  1. Reconstruct seen content sets – The filter builds a set from seen_ids and initializes any Bloom filters supplied in the query
  2. Iterate related post IDs – For each candidate, the filter examines all associated identifiers via related_post_ids_iter, including the original tweet, retweets, and quoted tweets
  3. Membership testing – If any related ID exists in the seen_ids set or returns positive from a Bloom filter, the candidate is marked as previously viewed
  4. Partition results – Candidates pass to the kept bucket if all related IDs are new; otherwise they move to removed

This comprehensive check ensures that variations of the same content (such as different retweets of an original tweet) do not escape filtration.

Output Format

The filter returns a FilterResult struct containing two vectors:

  • kept – Candidates representing fresh content the user has not encountered
  • removed – Candidates matching previously seen IDs that must be excluded from the timeline

Implementation Details in the Codebase

The deduplication logic resides in home-mixer/filters/previously_seen_posts_filter.rs, which implements the filter trait for the candidate processing pipeline. This module depends on data models defined in adjacent files:

The Bloom filter integration provides probabilistic membership testing that reduces memory overhead when handling large seen_ids lists, while the explicit set checking ensures zero false negatives for recently viewed content.

Practical Implementation Example

To integrate PreviouslySeenPostsFilter into a custom pipeline stage:

use home_mixer::filters::previously_seen_posts_filter::PreviouslySeenPostsFilter;
use home_mixer::models::query::ScoredPostsQuery;
use home_mixer::models::candidate::PostCandidate;
use xai_candidate_pipeline::filter::FilterResult;

// Initialize the filter
let filter = PreviouslySeenPostsFilter;

// Execute filtration
// query: ScoredPostsQuery containing seen_ids and Bloom filters
// candidates: Vec<PostCandidate> from upstream generators
let FilterResult { kept, removed } = filter.filter(&query, candidates);

// Process results
// kept: Vec<PostCandidate> containing only fresh content
// removed: Vec<PostCandidate> containing duplicates to discard

This pattern allows the Home-Mixer to maintain separation between fresh and previously viewed content before final timeline composition.

Summary

  • PreviouslySeenPostsFilter eliminates duplicate content by screening candidates against historical user data stored in ScoredPostsQuery
  • The filter checks all related post IDs (originals, retweets, quotes) via related_post_ids_iter to catch content variations
  • It combines explicit seen_ids sets with Bloom filters for efficient deduplication at scale
  • The implementation lives in home-mixer/filters/previously_seen_posts_filter.rs and returns a FilterResult separating kept and removed candidates
  • Early execution in the candidate pipeline ensures timeline freshness before expensive ranking operations

Frequently Asked Questions

What inputs does PreviouslySeenPostsFilter require?

The filter requires two primary inputs: a ScoredPostsQuery containing the user's seen_ids and optional Bloom filters, and a vector of PostCandidate objects. These are defined in home-mixer/models/query.rs and home-mixer/models/candidate.rs respectively, as implemented in the xai-org/x-algorithm repository.

How does PreviouslySeenPostsFilter handle retweets and quoted tweets?

The filter examines all related post identifiers through the related_post_ids_iter method on PostCandidate. This includes the original tweet ID, retweet IDs, and quoted tweet IDs. If any related identifier matches the user's seen_ids or Bloom filter entries, the entire candidate is filtered out, preventing duplicate content from appearing in different structural forms.

Why use Bloom filters alongside explicit seen_ids?

Bloom filters provide memory-efficient probabilistic testing for large historical datasets, reducing the storage footprint required to track seen content over time. The explicit seen_ids set ensures deterministic exclusion for recent interactions. Together, they balance accuracy and resource efficiency in the deduplication pipeline described in home-mixer/filters/previously_seen_posts_filter.rs.

Where is PreviouslySeenPostsFilter located in the X-Algorithm codebase?

The filter implementation resides in home-mixer/filters/previously_seen_posts_filter.rs within the xai-org/x-algorithm repository. It relies on supporting data structures from home-mixer/models/query.rs (for ScoredPostsQuery) and home-mixer/models/candidate.rs (for PostCandidate).

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →