What Is PreviouslySeenPostsFilter? Timeline Deduplication in X-Algorithm
PreviouslySeenPostsFilter removes candidate posts that a user has already encountered by checking seen_ids and Bloom filters, ensuring the Home-Mixer timeline surfaces only fresh content.
The PreviouslySeenPostsFilter is a critical deduplication component in the xai-org/x-algorithm repository's Home-Mixer pipeline. It operates early in the candidate generation workflow to prevent duplicate content from appearing in a user's timeline. By filtering out previously seen posts before ranking and mixing stages, the filter maintains content freshness and improves user experience.
How PreviouslySeenPostsFilter Works
The filter implements a strict exclusion policy based on historical user exposure. It accepts a ScoredPostsQuery containing the user's viewing history and a vector of PostCandidate objects, then partitions the candidates into fresh and previously viewed sets.
Input Parameters and Data Structures
According to the source code in home-mixer/models/query.rs, the filter receives:
ScoredPostsQuery– Contains the user'sseen_ids(explicitly viewed tweet IDs) and optional Bloom filter entries for memory-efficient membership testingPostCandidatevector – The batch of candidate posts generated by earlier pipeline stages, defined inhome-mixer/models/candidate.rs
These inputs provide the filter with both the candidate pool and the historical context needed to make exclusion decisions.
The Deduplication Logic
Inside home-mixer/filters/previously_seen_posts_filter.rs, the filter executes the following steps:
- Reconstruct seen content sets – The filter builds a set from
seen_idsand initializes any Bloom filters supplied in the query - Iterate related post IDs – For each candidate, the filter examines all associated identifiers via
related_post_ids_iter, including the original tweet, retweets, and quoted tweets - Membership testing – If any related ID exists in the
seen_idsset or returns positive from a Bloom filter, the candidate is marked as previously viewed - Partition results – Candidates pass to the
keptbucket if all related IDs are new; otherwise they move toremoved
This comprehensive check ensures that variations of the same content (such as different retweets of an original tweet) do not escape filtration.
Output Format
The filter returns a FilterResult struct containing two vectors:
kept– Candidates representing fresh content the user has not encounteredremoved– Candidates matching previously seen IDs that must be excluded from the timeline
Implementation Details in the Codebase
The deduplication logic resides in home-mixer/filters/previously_seen_posts_filter.rs, which implements the filter trait for the candidate processing pipeline. This module depends on data models defined in adjacent files:
home-mixer/models/query.rs– Defines theScoredPostsQuerystructure that carries exposure metadatahome-mixer/models/candidate.rs– Defines thePostCandidatestructure and therelated_post_ids_itermethod used to traverse post relationships
The Bloom filter integration provides probabilistic membership testing that reduces memory overhead when handling large seen_ids lists, while the explicit set checking ensures zero false negatives for recently viewed content.
Practical Implementation Example
To integrate PreviouslySeenPostsFilter into a custom pipeline stage:
use home_mixer::filters::previously_seen_posts_filter::PreviouslySeenPostsFilter;
use home_mixer::models::query::ScoredPostsQuery;
use home_mixer::models::candidate::PostCandidate;
use xai_candidate_pipeline::filter::FilterResult;
// Initialize the filter
let filter = PreviouslySeenPostsFilter;
// Execute filtration
// query: ScoredPostsQuery containing seen_ids and Bloom filters
// candidates: Vec<PostCandidate> from upstream generators
let FilterResult { kept, removed } = filter.filter(&query, candidates);
// Process results
// kept: Vec<PostCandidate> containing only fresh content
// removed: Vec<PostCandidate> containing duplicates to discard
This pattern allows the Home-Mixer to maintain separation between fresh and previously viewed content before final timeline composition.
Summary
- PreviouslySeenPostsFilter eliminates duplicate content by screening candidates against historical user data stored in
ScoredPostsQuery - The filter checks all related post IDs (originals, retweets, quotes) via
related_post_ids_iterto catch content variations - It combines explicit
seen_idssets with Bloom filters for efficient deduplication at scale - The implementation lives in
home-mixer/filters/previously_seen_posts_filter.rsand returns aFilterResultseparating kept and removed candidates - Early execution in the candidate pipeline ensures timeline freshness before expensive ranking operations
Frequently Asked Questions
What inputs does PreviouslySeenPostsFilter require?
The filter requires two primary inputs: a ScoredPostsQuery containing the user's seen_ids and optional Bloom filters, and a vector of PostCandidate objects. These are defined in home-mixer/models/query.rs and home-mixer/models/candidate.rs respectively, as implemented in the xai-org/x-algorithm repository.
How does PreviouslySeenPostsFilter handle retweets and quoted tweets?
The filter examines all related post identifiers through the related_post_ids_iter method on PostCandidate. This includes the original tweet ID, retweet IDs, and quoted tweet IDs. If any related identifier matches the user's seen_ids or Bloom filter entries, the entire candidate is filtered out, preventing duplicate content from appearing in different structural forms.
Why use Bloom filters alongside explicit seen_ids?
Bloom filters provide memory-efficient probabilistic testing for large historical datasets, reducing the storage footprint required to track seen content over time. The explicit seen_ids set ensures deterministic exclusion for recent interactions. Together, they balance accuracy and resource efficiency in the deduplication pipeline described in home-mixer/filters/previously_seen_posts_filter.rs.
Where is PreviouslySeenPostsFilter located in the X-Algorithm codebase?
The filter implementation resides in home-mixer/filters/previously_seen_posts_filter.rs within the xai-org/x-algorithm repository. It relies on supporting data structures from home-mixer/models/query.rs (for ScoredPostsQuery) and home-mixer/models/candidate.rs (for PostCandidate).
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →