How Phoenix Retrieval Finds Candidate Posts: Inside the X Algorithm Corpus System
Phoenix retrieval builds candidate post sets through a two-stage pipeline that first loads static post-author corpora via RetrievalDataset and then samples topically-aligned, popularity-weighted candidates using World.sample_candidate_posts.
The xai-org/x-algorithm repository powers X's recommendation infrastructure, including the Phoenix retrieval subsystem responsible for selecting candidate posts for users. Understanding how this system identifies potential content requires examining both the static corpus loading mechanism and the dynamic sampling logic that balances user interests with post popularity.
Loading the Candidate Corpus
The retrieval pipeline begins by materializing massive, static post indices from persistent storage.
RetrievalDataset Enum and Parquet Loading
The system defines distinct retrieval datasets—such as HOME, RELEVANT_ADS, and ACTIVE_ADS—as an Enum in phoenix/xrex/data/retrieval_dataset.py. At training initialization, the system invokes RetrievalDataset.load_datasets(...), which reads parquet files listed in each enum entry (with O₂/GCS fallbacks) and returns concatenated arrays containing post_id, author_id, type identifiers, and optional semantic-ID (SID) codes.
Specifically, the load and load_datasets methods handle the I/O logic, truncating or padding corpora to fixed sizes as configured. This implementation ensures the trainer has immediate access to millions of potential candidates without real-time database queries.
from xrex.data.retrieval_dataset import RetrievalDataset
# Load the HOME and RELEVANT_ADS corpora, with optional SID codes
post_ids, author_ids, types, post_sids = RetrievalDataset.load_datasets(
datasets=[RetrievalDataset.HOME, RetrievalDataset.RELEVANT_ADS],
max_posts=500_000, # truncate or pad to a fixed size
read_post_sid=True, # also load semantic IDs
sid_num_levels=6,
)
Sampling Candidate Posts for a User
Once the static corpus is loaded, the synthetic training world dynamically samples candidate posts tailored to individual user preferences.
Topic-Aligned Candidate Generation
The World.sample_candidate_posts method in phoenix/reference/world.py implements a sophisticated sampling strategy. For a given user, it first draws 2 × n + 8 raw topic IDs, where each draw aligns with the user's preferred topics with 0.7 probability; otherwise, it selects a random topic. This alignment step biases the candidate pool toward user interests while preserving diversity.
Popularity-Weighted Selection and Deduplication
These topic IDs feed into _sample_topic_pool, which uses the serve-weight cumulative distribution (self._serve_cum) to select posts proportional to their serve weight (post_popularity). The cumulative arrays are pre-computed once per world initialization in _build_topic_pools.
After initial selection, the system deduplicates results using np.unique and truncates to the requested n candidates. If the distinct set remains smaller than required, random post IDs augment the pool until reaching the target size.
import numpy as np
from phoenix.reference.world import generate_world
world = generate_world(seed=20260721)
rng = np.random.default_rng(123)
user_row = 0 # first user in the world
n_candidates = 64
candidates = world.sample_candidate_posts(rng, user_row, n_candidates)
print("candidate post IDs:", candidates)
Integration into Training Pipelines
The retrieval logic integrates directly into Phoenix's training infrastructure. Files such as phoenix/xrex/train/trainer_recsys.py and phoenix/xrex/train/trainer_sid_retrieval.py invoke World.sample_candidate_posts to construct candidate portions of each training batch. Similarly, phoenix/reference/dump_gen.py generates parquet training data by batching these samples into rows via _build_rows, assigning candidate arrays to cand_rows tensor columns.
# Inside _build_rows(...) in dump_gen.py
cand = world.sample_candidate_posts(rng, u, int(cand_lens[i]))
cols["cand_rows"][i, :n_cand] = cand
cols["n_cand"][i] = n_cand
Summary
- Two-stage architecture: Phoenix retrieval separates static corpus loading (
RetrievalDataset) from dynamic user-specific sampling (World.sample_candidate_posts). - Corpus sources: The system supports multiple dataset types (HOME, RELEVANT_ADS, ACTIVE_ADS) loaded from parquet files with O₂/GCS fallbacks.
- Topic alignment: Candidate generation uses a 0.7 probability alignment with user preferences mixed with random topics to balance relevance and diversity.
- Popularity weighting: Post selection uses cumulative serve-weight distributions based on
post_popularityscores. - Robust sampling: The pipeline deduplicates candidates using
np.uniqueand pads with random posts when necessary to guarantee fixed-size candidate sets.
Frequently Asked Questions
What file handles the initial loading of post corpora in Phoenix?
The phoenix/xrex/data/retrieval_dataset.py file defines the RetrievalDataset enum and implements the load_datasets method that reads parquet files and returns arrays of post metadata, author IDs, and optional semantic-ID codes.
How does Phoenix balance user interests with content diversity when sampling?
The system draws 2 × n + 8 topic IDs where each has a 70% probability of aligning with the user's preferred topics and a 30% probability of being random. This mixture ensures the candidate pool reflects user interests while introducing diverse content into the training distribution.
What happens if the deduplicated candidate set is smaller than requested?
If np.unique produces fewer distinct posts than the target number n, the system appends random post IDs to the candidate pool until reaching the required size. This guarantees fixed-size tensors for batch training regardless of topic pool overlap.
Where is the popularity weighting for posts calculated?
The cumulative serve-weight distribution (_serve_cum) is built once per world in _build_topic_pools within phoenix/reference/world.py. This distribution uses each post's post_popularity score to weight the sampling probability, ensuring higher-engagement content appears more frequently in candidate sets.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →