# How Pre-Scoring Filters Process Candidate Posts in X Algorithm: DropDuplicatesFilter and AgeFilter Explained

> Learn how X algorithm pre-scoring filters like DropDuplicatesFilter and AgeFilter process candidate posts. Discover how they exclude duplicates and aged content before ranking.

- Repository: [SpaceXAI Org/x-algorithm](https://github.com/xai-org/x-algorithm)
- Tags: deep-dive
- Published: 2026-09-10

---

**Pre-scoring filters in the X recommendation system generate boolean masks that are logically ANDed together to exclude duplicate posts and aged content before top-k selection, forcing filtered candidate scores to the minimum bfloat16 value to ensure they are never ranked.**

The xai-org/x-algorithm repository implements a two-tower recommendation architecture that strictly filters candidate posts before final model scoring occurs. Both **DropDuplicatesFilter** and **AgeFilter** operate within the serving pipeline to eliminate invalid candidates at the inference stage, reducing computational overhead and improving result quality. These filters execute in [`phoenix/xrex/models/recsys_two_tower_serving_filters.py`](https://github.com/xai-org/x-algorithm/blob/main/phoenix/xrex/models/recsys_two_tower_serving_filters.py) by manipulating binary masks that interact directly with the raw similarity matrix.

## The Pre-Scoring Filter Pipeline

The filtering mechanism executes in four distinct phases within the serving infrastructure. According to the source code in [`recsys_two_tower_serving_filters.py`](https://github.com/xai-org/x-algorithm/blob/main/recsys_two_tower_serving_filters.py), the pipeline first computes dense similarity scores, then constructs candidate-specific boolean masks, combines them using logical AND operations, and finally applies the combined mask to zero out disqualified candidates before top-k selection.

### Computing Initial Candidate Scores

The process begins in the `forward_with_filters` function, which delegates similarity computation to `compute_top_k` (lines 60-71). This method calculates a dense matrix `all_scores` representing the similarity between the user embedding and every candidate post embedding across all shards:

```python

# Inside forward_with_filters → compute_top_k (lines 60-71)

all_scores = compute_top_k(
    user_embedding=batch.user_emb,
    post_embeddings=post_embs,
    # ... additional parameters

)

```

At this stage, the model has gathered candidates from all shards but has not yet applied content-based filters.

### Constructing Boolean Masks in mask_and_top_k

Following score computation, the system enters the `mask_and_top_k` function (lines 91-93), where individual filters generate masks of identical shape to `all_scores`. Each mask contains `True` values for candidates that pass the filter and `False` for those that fail. The base masks include:

- `type_mask`: Filters by dataset type (e.g., organic posts only)
- `user_eligible_mask`: Optional per-user eligibility constraints

These masks are combined using logical AND operations: `combined = type_mask & user_eligible_mask`.

## DropDuplicatesFilter Implementation

The **DropDuplicatesFilter** eliminates post IDs that appear across multiple shards, ensuring the final candidate list contains only unique content. The filter maintains state using a Bloom filter to track previously seen post IDs during the gathering phase.

Inside `mask_and_top_k`, the duplicate removal logic checks each candidate against the Bloom filter before final selection:

```python

# Simplified logic from mask_and_top_k

combined = type_mask & user_eligible_mask  # Start with base masks

if use_duplicate_filter:
    # Bloom filter marks duplicates as False

    duplicate_mask = bloom_filter.check(post_ids)
    combined = combined & duplicate_mask

```

Candidates marked as duplicates receive `False` values in the combined mask, preventing them from advancing to the top-k stage regardless of their similarity scores.

## AgeFilter Implementation

The **AgeFilter** enforces content freshness by examining the `postAgeBucketSeq` feature defined in [`phoenix/xrex/data/recsys/feature_config.py`](https://github.com/xai-org/x-algorithm/blob/main/phoenix/xrex/data/recsys/feature_config.py) (lines 13-28). This filter reads age bucket values directly from the candidate feature vector and compares them against configured maximum thresholds.

The implementation extracts the age bucket column from `post_features` and creates a boolean mask where only posts newer than the maximum allowed age receive `True` values:

```python

# Age filtering logic inside mask_and_top_k

age_bucket = post_features[:, postAgeBucketSeq]
age_mask = age_bucket <= MAX_ALLOWED_BUCKET  # Keep only recent posts

combined = combined & age_mask

```

Posts exceeding the age threshold (e.g., older than 30 days) are marked `False` in the mask, effectively removing aged content from consideration before ranking occurs.

## Mask Combination and Score Application

After all filters generate their respective masks, the system combines them into a single boolean tensor and applies it to the raw similarity scores. The critical operation occurs at line 92 of `mask_and_top_k`, where `jnp.where` forces filtered scores to the minimum possible bfloat16 value:

```python

# Line 92 in mask_and_top_k

masked_scores = jnp.where(
    combined, 
    all_scores, 
    jnp.finfo(jnp.bfloat16).min  # Minimum value ensures exclusion

)

```

This operation guarantees that filtered candidates receive scores lower than any valid candidate, ensuring the subsequent `top_k_by_key` routine (called via `local_top_k` at line 49) never selects them.

## Summary

- **Pre-scoring filters** execute before final ranking by generating boolean masks that eliminate invalid candidates.
- **DropDuplicatesFilter** uses a Bloom filter to remove duplicate post IDs that appear across multiple shards, modifying the combined mask in `mask_and_top_k`.
- **AgeFilter** references `postAgeBucketSeq` from [`feature_config.py`](https://github.com/xai-org/x-algorithm/blob/main/feature_config.py) to exclude posts exceeding configured age thresholds.
- **Mask combination** occurs through logical AND operations (`combined = type_mask & user_eligible_mask & age_mask & duplicate_mask`).
- **Score masking** uses `jnp.where` to set disqualified candidate scores to `jnp.finfo(jnp.bfloat16).min`, ensuring they never appear in the final top-k results from `top_k_by_key`.

## Frequently Asked Questions

### What is the order of operations for pre-scoring filters in X Algorithm?

The filters execute after the model computes the initial similarity matrix in `compute_top_k` but before the `top_k_by_key` selection runs. Specifically, `mask_and_top_k` (lines 91-93) constructs and combines all masks—including duplicate and age filters—then applies them to the scores before returning the final candidates.

### How does the DropDuplicatesFilter handle posts across multiple shards?

The filter operates during the candidate gathering phase using a Bloom filter to track seen post IDs. When `mask_and_top_k` processes candidates from multiple shards, any post ID already present in the Bloom filter is marked `False` in the duplicate mask, ensuring only the first occurrence of each post advances to scoring.

### What happens to candidate scores that fail the age filter?

Candidates failing the age threshold receive `False` values in the age mask, which propagates to the combined mask. The system then uses `jnp.where(combined, all_scores, jnp.finfo(jnp.bfloat16).min)` to replace their similarity scores with the minimum bfloat16 value, effectively ranking them below all valid candidates and excluding them from the final output.

### Where are the age bucket thresholds defined in the codebase?

The age bucket feature sequence `postAgeBucketSeq` is declared in [`phoenix/xrex/data/recsys/feature_config.py`](https://github.com/xai-org/x-algorithm/blob/main/phoenix/xrex/data/recsys/feature_config.py) (lines 13-28). The specific maximum bucket threshold (`MAX_ALLOWED_BUCKET`) used by the AgeFilter is configured within the serving logic in [`phoenix/xrex/models/recsys_two_tower_serving_filters.py`](https://github.com/xai-org/x-algorithm/blob/main/phoenix/xrex/models/recsys_two_tower_serving_filters.py), where it compares against the extracted bucket values from the candidate features.