How Pre-Scoring Filters Process Candidate Posts in X Algorithm: DropDuplicatesFilter and AgeFilter Explained
Pre-scoring filters in the X recommendation system generate boolean masks that are logically ANDed together to exclude duplicate posts and aged content before top-k selection, forcing filtered candidate scores to the minimum bfloat16 value to ensure they are never ranked.
The xai-org/x-algorithm repository implements a two-tower recommendation architecture that strictly filters candidate posts before final model scoring occurs. Both DropDuplicatesFilter and AgeFilter operate within the serving pipeline to eliminate invalid candidates at the inference stage, reducing computational overhead and improving result quality. These filters execute in phoenix/xrex/models/recsys_two_tower_serving_filters.py by manipulating binary masks that interact directly with the raw similarity matrix.
The Pre-Scoring Filter Pipeline
The filtering mechanism executes in four distinct phases within the serving infrastructure. According to the source code in recsys_two_tower_serving_filters.py, the pipeline first computes dense similarity scores, then constructs candidate-specific boolean masks, combines them using logical AND operations, and finally applies the combined mask to zero out disqualified candidates before top-k selection.
Computing Initial Candidate Scores
The process begins in the forward_with_filters function, which delegates similarity computation to compute_top_k (lines 60-71). This method calculates a dense matrix all_scores representing the similarity between the user embedding and every candidate post embedding across all shards:
# Inside forward_with_filters → compute_top_k (lines 60-71)
all_scores = compute_top_k(
user_embedding=batch.user_emb,
post_embeddings=post_embs,
# ... additional parameters
)
At this stage, the model has gathered candidates from all shards but has not yet applied content-based filters.
Constructing Boolean Masks in mask_and_top_k
Following score computation, the system enters the mask_and_top_k function (lines 91-93), where individual filters generate masks of identical shape to all_scores. Each mask contains True values for candidates that pass the filter and False for those that fail. The base masks include:
type_mask: Filters by dataset type (e.g., organic posts only)user_eligible_mask: Optional per-user eligibility constraints
These masks are combined using logical AND operations: combined = type_mask & user_eligible_mask.
DropDuplicatesFilter Implementation
The DropDuplicatesFilter eliminates post IDs that appear across multiple shards, ensuring the final candidate list contains only unique content. The filter maintains state using a Bloom filter to track previously seen post IDs during the gathering phase.
Inside mask_and_top_k, the duplicate removal logic checks each candidate against the Bloom filter before final selection:
# Simplified logic from mask_and_top_k
combined = type_mask & user_eligible_mask # Start with base masks
if use_duplicate_filter:
# Bloom filter marks duplicates as False
duplicate_mask = bloom_filter.check(post_ids)
combined = combined & duplicate_mask
Candidates marked as duplicates receive False values in the combined mask, preventing them from advancing to the top-k stage regardless of their similarity scores.
AgeFilter Implementation
The AgeFilter enforces content freshness by examining the postAgeBucketSeq feature defined in phoenix/xrex/data/recsys/feature_config.py (lines 13-28). This filter reads age bucket values directly from the candidate feature vector and compares them against configured maximum thresholds.
The implementation extracts the age bucket column from post_features and creates a boolean mask where only posts newer than the maximum allowed age receive True values:
# Age filtering logic inside mask_and_top_k
age_bucket = post_features[:, postAgeBucketSeq]
age_mask = age_bucket <= MAX_ALLOWED_BUCKET # Keep only recent posts
combined = combined & age_mask
Posts exceeding the age threshold (e.g., older than 30 days) are marked False in the mask, effectively removing aged content from consideration before ranking occurs.
Mask Combination and Score Application
After all filters generate their respective masks, the system combines them into a single boolean tensor and applies it to the raw similarity scores. The critical operation occurs at line 92 of mask_and_top_k, where jnp.where forces filtered scores to the minimum possible bfloat16 value:
# Line 92 in mask_and_top_k
masked_scores = jnp.where(
combined,
all_scores,
jnp.finfo(jnp.bfloat16).min # Minimum value ensures exclusion
)
This operation guarantees that filtered candidates receive scores lower than any valid candidate, ensuring the subsequent top_k_by_key routine (called via local_top_k at line 49) never selects them.
Summary
- Pre-scoring filters execute before final ranking by generating boolean masks that eliminate invalid candidates.
- DropDuplicatesFilter uses a Bloom filter to remove duplicate post IDs that appear across multiple shards, modifying the combined mask in
mask_and_top_k. - AgeFilter references
postAgeBucketSeqfromfeature_config.pyto exclude posts exceeding configured age thresholds. - Mask combination occurs through logical AND operations (
combined = type_mask & user_eligible_mask & age_mask & duplicate_mask). - Score masking uses
jnp.whereto set disqualified candidate scores tojnp.finfo(jnp.bfloat16).min, ensuring they never appear in the final top-k results fromtop_k_by_key.
Frequently Asked Questions
What is the order of operations for pre-scoring filters in X Algorithm?
The filters execute after the model computes the initial similarity matrix in compute_top_k but before the top_k_by_key selection runs. Specifically, mask_and_top_k (lines 91-93) constructs and combines all masks—including duplicate and age filters—then applies them to the scores before returning the final candidates.
How does the DropDuplicatesFilter handle posts across multiple shards?
The filter operates during the candidate gathering phase using a Bloom filter to track seen post IDs. When mask_and_top_k processes candidates from multiple shards, any post ID already present in the Bloom filter is marked False in the duplicate mask, ensuring only the first occurrence of each post advances to scoring.
What happens to candidate scores that fail the age filter?
Candidates failing the age threshold receive False values in the age mask, which propagates to the combined mask. The system then uses jnp.where(combined, all_scores, jnp.finfo(jnp.bfloat16).min) to replace their similarity scores with the minimum bfloat16 value, effectively ranking them below all valid candidates and excluding them from the final output.
Where are the age bucket thresholds defined in the codebase?
The age bucket feature sequence postAgeBucketSeq is declared in phoenix/xrex/data/recsys/feature_config.py (lines 13-28). The specific maximum bucket threshold (MAX_ALLOWED_BUCKET) used by the AgeFilter is configured within the serving logic in phoenix/xrex/models/recsys_two_tower_serving_filters.py, where it compares against the extracted bucket values from the candidate features.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →