# How Phoenix Retrieval Finds Candidate Posts: Inside the X Algorithm Corpus System

> Discover how Phoenix retrieval finds candidate posts using a two-stage pipeline with RetrievalDataset and World.sample_candidate_posts for efficient post selection.

- Repository: [SpaceXAI Org/x-algorithm](https://github.com/xai-org/x-algorithm)
- Tags: deep-dive
- Published: 2026-09-12

---

**Phoenix retrieval builds candidate post sets through a two-stage pipeline that first loads static post-author corpora via `RetrievalDataset` and then samples topically-aligned, popularity-weighted candidates using `World.sample_candidate_posts`.**

The `xai-org/x-algorithm` repository powers X's recommendation infrastructure, including the Phoenix retrieval subsystem responsible for selecting candidate posts for users. Understanding how this system identifies potential content requires examining both the static corpus loading mechanism and the dynamic sampling logic that balances user interests with post popularity.

## Loading the Candidate Corpus

The retrieval pipeline begins by materializing massive, static post indices from persistent storage.

### RetrievalDataset Enum and Parquet Loading

The system defines distinct retrieval datasets—such as **HOME**, **RELEVANT_ADS**, and **ACTIVE_ADS**—as an `Enum` in [`phoenix/xrex/data/retrieval_dataset.py`](https://github.com/xai-org/x-algorithm/blob/main/phoenix/xrex/data/retrieval_dataset.py). At training initialization, the system invokes `RetrievalDataset.load_datasets(...)`, which reads parquet files listed in each enum entry (with O₂/GCS fallbacks) and returns concatenated arrays containing `post_id`, `author_id`, type identifiers, and optional semantic-ID (SID) codes.

Specifically, the `load` and `load_datasets` methods handle the I/O logic, truncating or padding corpora to fixed sizes as configured. This implementation ensures the trainer has immediate access to millions of potential candidates without real-time database queries.

```python
from xrex.data.retrieval_dataset import RetrievalDataset

# Load the HOME and RELEVANT_ADS corpora, with optional SID codes

post_ids, author_ids, types, post_sids = RetrievalDataset.load_datasets(
    datasets=[RetrievalDataset.HOME, RetrievalDataset.RELEVANT_ADS],
    max_posts=500_000,            # truncate or pad to a fixed size

    read_post_sid=True,          # also load semantic IDs

    sid_num_levels=6,
)

```

## Sampling Candidate Posts for a User

Once the static corpus is loaded, the synthetic training world dynamically samples candidate posts tailored to individual user preferences.

### Topic-Aligned Candidate Generation

The `World.sample_candidate_posts` method in [`phoenix/reference/world.py`](https://github.com/xai-org/x-algorithm/blob/main/phoenix/reference/world.py) implements a sophisticated sampling strategy. For a given user, it first draws **2 × n + 8** raw topic IDs, where each draw aligns with the user's preferred topics with **0.7 probability**; otherwise, it selects a random topic. This alignment step biases the candidate pool toward user interests while preserving diversity.

### Popularity-Weighted Selection and Deduplication

These topic IDs feed into `_sample_topic_pool`, which uses the **serve-weight cumulative distribution** (`self._serve_cum`) to select posts proportional to their **serve weight** (`post_popularity`). The cumulative arrays are pre-computed once per world initialization in `_build_topic_pools`.

After initial selection, the system deduplicates results using `np.unique` and truncates to the requested `n` candidates. If the distinct set remains smaller than required, random post IDs augment the pool until reaching the target size.

```python
import numpy as np
from phoenix.reference.world import generate_world

world = generate_world(seed=20260721)
rng = np.random.default_rng(123)

user_row = 0                # first user in the world

n_candidates = 64
candidates = world.sample_candidate_posts(rng, user_row, n_candidates)
print("candidate post IDs:", candidates)

```

## Integration into Training Pipelines

The retrieval logic integrates directly into Phoenix's training infrastructure. Files such as [`phoenix/xrex/train/trainer_recsys.py`](https://github.com/xai-org/x-algorithm/blob/main/phoenix/xrex/train/trainer_recsys.py) and [`phoenix/xrex/train/trainer_sid_retrieval.py`](https://github.com/xai-org/x-algorithm/blob/main/phoenix/xrex/train/trainer_sid_retrieval.py) invoke `World.sample_candidate_posts` to construct candidate portions of each training batch. Similarly, [`phoenix/reference/dump_gen.py`](https://github.com/xai-org/x-algorithm/blob/main/phoenix/reference/dump_gen.py) generates parquet training data by batching these samples into rows via `_build_rows`, assigning candidate arrays to `cand_rows` tensor columns.

```python

# Inside _build_rows(...) in dump_gen.py

cand = world.sample_candidate_posts(rng, u, int(cand_lens[i]))
cols["cand_rows"][i, :n_cand] = cand
cols["n_cand"][i] = n_cand

```

## Summary

- **Two-stage architecture**: Phoenix retrieval separates static corpus loading (`RetrievalDataset`) from dynamic user-specific sampling (`World.sample_candidate_posts`).
- **Corpus sources**: The system supports multiple dataset types (HOME, RELEVANT_ADS, ACTIVE_ADS) loaded from parquet files with O₂/GCS fallbacks.
- **Topic alignment**: Candidate generation uses a 0.7 probability alignment with user preferences mixed with random topics to balance relevance and diversity.
- **Popularity weighting**: Post selection uses cumulative serve-weight distributions based on `post_popularity` scores.
- **Robust sampling**: The pipeline deduplicates candidates using `np.unique` and pads with random posts when necessary to guarantee fixed-size candidate sets.

## Frequently Asked Questions

### What file handles the initial loading of post corpora in Phoenix?

The [`phoenix/xrex/data/retrieval_dataset.py`](https://github.com/xai-org/x-algorithm/blob/main/phoenix/xrex/data/retrieval_dataset.py) file defines the `RetrievalDataset` enum and implements the `load_datasets` method that reads parquet files and returns arrays of post metadata, author IDs, and optional semantic-ID codes.

### How does Phoenix balance user interests with content diversity when sampling?

The system draws **2 × n + 8** topic IDs where each has a **70% probability** of aligning with the user's preferred topics and a **30% probability** of being random. This mixture ensures the candidate pool reflects user interests while introducing diverse content into the training distribution.

### What happens if the deduplicated candidate set is smaller than requested?

If `np.unique` produces fewer distinct posts than the target number `n`, the system appends random post IDs to the candidate pool until reaching the required size. This guarantees fixed-size tensors for batch training regardless of topic pool overlap.

### Where is the popularity weighting for posts calculated?

The cumulative serve-weight distribution (`_serve_cum`) is built once per world in `_build_topic_pools` within [`phoenix/reference/world.py`](https://github.com/xai-org/x-algorithm/blob/main/phoenix/reference/world.py). This distribution uses each post's `post_popularity` score to weight the sampling probability, ensuring higher-engagement content appears more frequently in candidate sets.