How the Out-of-Network Discount Factor Shapes Post Ranking in the X-Algorithm
The Out-of-Network discount factor applies a logarithmic position-based penalty to posts from accounts a viewer does not follow, reducing their contribution to the DCG score as they appear lower in the ranked feed.
In the xai-org/x-algorithm recommendation stack, the Out-of-Network discount factor serves as a critical mechanism within the Discounted Cumulative Gain (DCG) calculation to control the visibility of content from unfamiliar accounts. This factor ensures that posts originating from accounts a user does not follow receive diminishing influence based on their position in the ranked list, naturally prioritizing in-network content while maintaining diversity through position-aware weighting.
Understanding the Out-of-Network Discount Factor
The discount factor operates as a logarithmic position penalty that scales the contribution of each candidate post to the overall ranking metric. When the system evaluates relevance scores for posts from accounts a viewer does not follow, these candidates typically appear lower in the ranked list due to the model's inherent preference for known contacts.
The mathematical foundation relies on the standard DCG formulation adapted for the X-Algorithm's recommendation objectives. For every position p in the ranked list (starting at 0), the system calculates a discount value that increases logarithmically with depth, thereby compressing the impact of lower-ranked out-of-network content.
How the DCG Discount Is Computed in X-Algorithm
The core implementation resides in phoenix/xrex/models/recsys_model.py, where the discount calculation and DCG aggregation occur within the metric computation pipeline.
Specifically, lines 376-380 implement the discount mechanism using JAX operations:
discounts = jnp.log2(positions + 1) # ← phoenix/xrex/models/recsys_model.py
dcg = jnp.sum(sorted_y / discounts) # ← same file
idcg = jnp.sum(jnp.sort(valid_y, descending=True) / discounts) # ← same file
The computation follows this logic:
- Position mapping: For each position index in the ranked list, the discount factor calculates
log₂(p + 1), where earlier positions (smallp) yield discounts approaching 1.0, while later positions produce larger denominators - Score normalization: The sorted relevance scores (
sorted_y) are divided by their corresponding discount factors, reducing the contribution of items appearing at depths 2, 3, 4, and beyond - Ideal DCG calculation: The
idcgcomputation sorts valid labels in descending order to establish the optimal ranking benchmark against which actual performance is measured
Impact on Out-of-Network Content
The Out-of-Network discount factor creates a position-aware bias that aligns with typical user expectations while preventing over-optimization for distant candidates. When posts from accounts a viewer does not follow appear in the feed, they face a compounding disadvantage:
- Initial ranking penalty: Out-of-network posts typically receive lower raw similarity scores from the two-tower architecture, placing them at greater depth in the initial ranking
- DCG dampening: Once positioned lower in the list, the logarithmic discount further reduces their contribution to the loss and metric calculations
- Learning signal: During training, this dual penalty teaches the model to surface in-network content preferentially, as out-of-network candidates cannot overcome the position-based discount even with high raw similarity scores
Consequently, the system implicitly penalizes out-of-network posts that would otherwise dominate lower-rank slots, ensuring the model does not over-optimize for potentially irrelevant distant candidates.
Implementation in the Training Pipeline
The discount factor integrates into both training and evaluation phases through specific function calls and architectural components.
Training Phase Integration
During training, the compute_retrieval_loss function utilizes discounted scores via raw_batch_scores and raw_global_neg_scores to compute contrastive loss. The loss calculation incorporates the DCG-style discounts to weight the importance of positive versus negative examples based on their position in the candidate list.
The two-tower architecture defined in phoenix/xrex/models/recsys_two_tower_model.py feeds embeddings into this pipeline. The call chain progresses from compute_retrieval_loss through _sharded_full_matmul to metric aggregation, consistently applying the DCG-discounted scores throughout the gradient computation.
Evaluation Metrics
During evaluation, the _compute_retrieval_metrics function returns recall@k and DCG-based statistics using the same discounting logic. The attention mechanisms in phoenix/xrex/models/recsys_attention.py generate the raw scores that subsequently undergo discounting, while phoenix/xrex/utils/utils.py handles metric summarization and logging of the discount-affected results.
Practical Code Examples
The following implementations demonstrate how to compute and apply the DCG discount factor in JAX, mirroring the approach used in the X-Algorithm codebase.
Computing DCG Discounts for Batch Positions
# Example: computing the DCG discount for a batch of candidate positions
import jax.numpy as jnp
def dcg_discount(positions: jnp.ndarray) -> jnp.ndarray:
"""Return the per‑position discount used in DCG."""
return jnp.log2(positions + 1)
# Suppose we have candidate scores sorted by relevance:
sorted_scores = jnp.array([0.9, 0.7, 0.4, 0.2]) # highest‑score first
positions = jnp.arange(sorted_scores.shape[0]) # [0, 1, 2, 3]
discounts = dcg_discount(positions) # → [0., 1., 1.5849, 2.]
# Apply discount (add 1 to avoid division by zero at position 0)
dcg = jnp.sum(sorted_scores / (discounts + 1))
print(dcg) # ≈ 2.28
Integrating Discounts into Contrastive Loss
# Example: integrating the discount into a loss computation (simplified)
def contrastive_loss(user_vec, candidate_vecs, temperature=0.07):
# Compute raw similarity scores
logits = user_vec @ candidate_vecs.T # shape: (batch, candidates)
# Apply temperature scaling
logits = logits / temperature
# Apply DCG‑style discount based on candidate rank
ranks = jnp.argsort(-logits, axis=-1) # descending order
discounts = jnp.log2(jnp.arange(ranks.shape[-1]) + 2) # +2 → avoid log2(1)=0
discounted_logits = logits / discounts
# Standard softmax cross‑entropy (positive at index 0 after sorting)
loss = -jnp.mean(jnp.log_softmax(discounted_logits)[:, 0])
return loss
Summary
- The Out-of-Network discount factor implements a logarithmic position penalty (
log₂(p + 1)) within the DCG calculation atphoenix/xrex/models/recsys_model.py, specifically lines 376-380. - Posts from accounts a viewer does not follow receive compound penalties: lower initial ranking positions combined with DCG discounting that reduces their contribution to training loss and evaluation metrics.
- The mechanism integrates into the contrastive learning pipeline through
compute_retrieval_lossand_compute_retrieval_metrics, utilizing the two-tower architecture fromrecsys_two_tower_model.py. - By applying uniform logarithmic discounts, the system maintains a position-aware bias that prioritizes in-network content while preventing over-optimization for irrelevant out-of-network candidates.
Frequently Asked Questions
How does the Out-of-Network discount factor mathematical formula work?
The factor applies log₂(position + 1) to each rank in the list, creating a denominator that grows logarithmically with depth. Earlier positions (0, 1, 2) receive discounts of approximately 1.0, 1.0, and 1.58 respectively, while deeper positions face progressively larger discounts. When raw relevance scores are divided by these values in the DCG summation, lower-ranked out-of-network posts contribute disproportionately less to the final metric.
Why does the X-Algorithm use DCG-based discounting for out-of-network posts?
The X-Algorithm employs DCG discounting to enforce a position-aware bias that mirrors user behavior—users typically value content appearing at the top of their feed more than content buried deeper. By logarithmically compressing the influence of lower-ranked posts, the system ensures that out-of-network content must achieve significantly higher relevance scores to overcome positional penalties, maintaining feed quality while allowing serendipitous discovery.
Where in the codebase is the discount factor applied?
The primary implementation resides in phoenix/xrex/models/recsys_model.py at lines 376-380, where jnp.log2(positions + 1) calculates the discount array. This feeds into the broader recommendation pipeline through phoenix/xrex/models/recsys_two_tower_model.py during loss computation and phoenix/xrex/utils/utils.py during metric summarization. The raw scores that undergo discounting originate from attention mechanisms in phoenix/xrex/models/recsys_attention.py.
How does the discount factor affect model training versus inference?
During training, the discount factor shapes the gradient signals in compute_retrieval_loss by weighting contrastive examples according to their discounted DCG contributions, teaching the model to optimize for top-ranked positions. During inference and evaluation, _compute_retrieval_metrics applies the same discounts when calculating recall@k and DCG statistics, ensuring that evaluation metrics align with the position-aware objectives used during optimization.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →