How Pre-Scoring Filters Process Candidate Posts in X Algorithm: DropDuplicatesFilter and AgeFilter Explained

Pre-scoring filters in the X recommendation system generate boolean masks that are logically ANDed together to exclude duplicate posts and aged content before top-k selection, forcing filtered candidate scores to the minimum bfloat16 value to ensure they are never ranked.

The xai-org/x-algorithm repository implements a two-tower recommendation architecture that strictly filters candidate posts before final model scoring occurs. Both DropDuplicatesFilter and AgeFilter operate within the serving pipeline to eliminate invalid candidates at the inference stage, reducing computational overhead and improving result quality. These filters execute in phoenix/xrex/models/recsys_two_tower_serving_filters.py by manipulating binary masks that interact directly with the raw similarity matrix.

The Pre-Scoring Filter Pipeline

The filtering mechanism executes in four distinct phases within the serving infrastructure. According to the source code in recsys_two_tower_serving_filters.py, the pipeline first computes dense similarity scores, then constructs candidate-specific boolean masks, combines them using logical AND operations, and finally applies the combined mask to zero out disqualified candidates before top-k selection.

Computing Initial Candidate Scores

The process begins in the forward_with_filters function, which delegates similarity computation to compute_top_k (lines 60-71). This method calculates a dense matrix all_scores representing the similarity between the user embedding and every candidate post embedding across all shards:


# Inside forward_with_filters → compute_top_k (lines 60-71)

all_scores = compute_top_k(
    user_embedding=batch.user_emb,
    post_embeddings=post_embs,
    # ... additional parameters

)

At this stage, the model has gathered candidates from all shards but has not yet applied content-based filters.

Constructing Boolean Masks in mask_and_top_k

Following score computation, the system enters the mask_and_top_k function (lines 91-93), where individual filters generate masks of identical shape to all_scores. Each mask contains True values for candidates that pass the filter and False for those that fail. The base masks include:

  • type_mask: Filters by dataset type (e.g., organic posts only)
  • user_eligible_mask: Optional per-user eligibility constraints

These masks are combined using logical AND operations: combined = type_mask & user_eligible_mask.

DropDuplicatesFilter Implementation

The DropDuplicatesFilter eliminates post IDs that appear across multiple shards, ensuring the final candidate list contains only unique content. The filter maintains state using a Bloom filter to track previously seen post IDs during the gathering phase.

Inside mask_and_top_k, the duplicate removal logic checks each candidate against the Bloom filter before final selection:


# Simplified logic from mask_and_top_k

combined = type_mask & user_eligible_mask  # Start with base masks

if use_duplicate_filter:
    # Bloom filter marks duplicates as False

    duplicate_mask = bloom_filter.check(post_ids)
    combined = combined & duplicate_mask

Candidates marked as duplicates receive False values in the combined mask, preventing them from advancing to the top-k stage regardless of their similarity scores.

AgeFilter Implementation

The AgeFilter enforces content freshness by examining the postAgeBucketSeq feature defined in phoenix/xrex/data/recsys/feature_config.py (lines 13-28). This filter reads age bucket values directly from the candidate feature vector and compares them against configured maximum thresholds.

The implementation extracts the age bucket column from post_features and creates a boolean mask where only posts newer than the maximum allowed age receive True values:


# Age filtering logic inside mask_and_top_k

age_bucket = post_features[:, postAgeBucketSeq]
age_mask = age_bucket <= MAX_ALLOWED_BUCKET  # Keep only recent posts

combined = combined & age_mask

Posts exceeding the age threshold (e.g., older than 30 days) are marked False in the mask, effectively removing aged content from consideration before ranking occurs.

Mask Combination and Score Application

After all filters generate their respective masks, the system combines them into a single boolean tensor and applies it to the raw similarity scores. The critical operation occurs at line 92 of mask_and_top_k, where jnp.where forces filtered scores to the minimum possible bfloat16 value:


# Line 92 in mask_and_top_k

masked_scores = jnp.where(
    combined, 
    all_scores, 
    jnp.finfo(jnp.bfloat16).min  # Minimum value ensures exclusion

)

This operation guarantees that filtered candidates receive scores lower than any valid candidate, ensuring the subsequent top_k_by_key routine (called via local_top_k at line 49) never selects them.

Summary

  • Pre-scoring filters execute before final ranking by generating boolean masks that eliminate invalid candidates.
  • DropDuplicatesFilter uses a Bloom filter to remove duplicate post IDs that appear across multiple shards, modifying the combined mask in mask_and_top_k.
  • AgeFilter references postAgeBucketSeq from feature_config.py to exclude posts exceeding configured age thresholds.
  • Mask combination occurs through logical AND operations (combined = type_mask & user_eligible_mask & age_mask & duplicate_mask).
  • Score masking uses jnp.where to set disqualified candidate scores to jnp.finfo(jnp.bfloat16).min, ensuring they never appear in the final top-k results from top_k_by_key.

Frequently Asked Questions

What is the order of operations for pre-scoring filters in X Algorithm?

The filters execute after the model computes the initial similarity matrix in compute_top_k but before the top_k_by_key selection runs. Specifically, mask_and_top_k (lines 91-93) constructs and combines all masks—including duplicate and age filters—then applies them to the scores before returning the final candidates.

How does the DropDuplicatesFilter handle posts across multiple shards?

The filter operates during the candidate gathering phase using a Bloom filter to track seen post IDs. When mask_and_top_k processes candidates from multiple shards, any post ID already present in the Bloom filter is marked False in the duplicate mask, ensuring only the first occurrence of each post advances to scoring.

What happens to candidate scores that fail the age filter?

Candidates failing the age threshold receive False values in the age mask, which propagates to the combined mask. The system then uses jnp.where(combined, all_scores, jnp.finfo(jnp.bfloat16).min) to replace their similarity scores with the minimum bfloat16 value, effectively ranking them below all valid candidates and excluding them from the final output.

Where are the age bucket thresholds defined in the codebase?

The age bucket feature sequence postAgeBucketSeq is declared in phoenix/xrex/data/recsys/feature_config.py (lines 13-28). The specific maximum bucket threshold (MAX_ALLOWED_BUCKET) used by the AgeFilter is configured within the serving logic in phoenix/xrex/models/recsys_two_tower_serving_filters.py, where it compares against the extracted bucket values from the candidate features.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →