Phoenix Scorer Query Context Hydration: How User Action Sequences Drive Relevance Scoring

The PhoenixScorer hydrates a comprehensive QueryContext object containing the user ID, current query string, historical query log, session metadata, item-level candidate features, and a chronological user action sequence—with each action recording action_type, item_id, timestamp, and optional metadata—enabling temporal pattern recognition during relevance scoring.

In the xai-org/x-algorithm repository, the PhoenixScorer serves as the central ranking component for Phoenix pipelines. Before computing relevance scores, it performs query context hydration, assembling a rich QueryContext (referenced as ctx in the source) that transforms raw dataset records into a structured representation of user intent and behavioral history.

What Is Query Context Hydration?

Query context hydration is the process of extracting and structuring data from the PhoenixDataset—defined in phoenix/xrex/data/parquet_recsys.py—into a QueryContext object usable by the scoring model. This hydration step occurs in the _build_query_context method (implemented in concrete scorers such as ReplyScorer at grox/flows/reply_spam/classifier_reply_ranking.py), which merges real-time request data with historical user behavior to create a complete scoring context.

Six Critical Components of the Hydrated Context

The scorer assembles six distinct categories of information into the final context object passed to the model.

User Identity and Current Query

At the foundation of the context lies user identification and immediate intent. The scorer extracts user_id and the raw query string directly from the PhoenixDataset record. These fields establish ownership of the session and capture the current search intent that triggered the ranking request.

Historical Query Log

To understand evolving user intent, the scorer retrieves prior interactions through self._get_user_history(user_id). This method queries the dataset to build a list of previous query strings submitted by the same user, allowing the model to recognize patterns such as query refinement or topic drift across the session.

User Action Sequences

The most behaviorally rich component is the user action sequence (user_action_seq), a chronologically ordered list of engagement events occurring in the current session. According to the PhoenixDataset implementation in phoenix/xrex/data/parquet_recsys.py, the scorer calls self._parse_actions(row) to yield Action objects from raw dataset rows.

Each action in the sequence contains:

  • action_type: The engagement category (e.g., click, impression, like, dwell_time, scroll)
  • item_id: The identifier of the ranked item receiving the action
  • timestamp: When the action occurred chronologically
  • metadata: Optional additional fields capturing supplementary engagement data

This sequence enables the model to leverage temporal patterns—for example, recognizing that a user recently clicked sports content when ranking subsequent queries.

Per-Item Candidate Features

For every candidate item under consideration, the scorer invokes self._load_item_features(item_id) to hydrate item-specific signals. This loader reads from the PhoenixDataset parquet columns—including embedding, tags, and pre-computed relevance signals—and attaches them to the candidate representation within the context.

Session Metadata and Device Context

Beyond behavioral signals, the context includes session-wide metadata extracted from PhoenixDataset fields such as device (device type), locale (geographic/language settings), and experiment_id (A/B test bucket assignments). These fields allow the scoring model to apply contextual rules based on platform capabilities or regional preferences.

Temporal Signals and Freshness

The scorer captures the request timestamp via request_timestamp = datetime.utcnow() and attaches it to the context. This temporal anchor enables time-aware scoring logic, such as decaying the influence of older actions or boosting recently trending content.

Implementation: How the Scorer Assembles Context

The concrete implementation in grox/flows/reply_spam/classifier_reply_ranking.py follows a structured hydration pipeline. Below is a simplified illustration of the _build_query_context method from phoenix/xrex/scorer.py:

def _build_query_context(self, request):
    # 1️⃣ Load the raw record from PhoenixDataset

    record = self.dataset.get_record(request.record_id)

    # 2️⃣ Extract basic fields

    user_id = record.user_id
    query = record.query
    timestamp = request.timestamp

    # 3️⃣ Assemble the user‑action sequence

    action_seq = []
    for act in record.actions:                # actions are stored as a list of dicts

        action_seq.append(
            UserAction(
                type=act["action_type"],
                item_id=act["item_id"],
                ts=act["timestamp"],
                meta=act.get("metadata", {})
            )
        )

    # 4️⃣ Pull historical queries

    history = self._get_user_history(user_id)

    # 5️⃣ Gather per‑item features for the candidates

    candidates = [
        self._load_item_features(item_id) for item_id in request.candidate_ids
    ]

    # 6️⃣ Create the final context object

    ctx = QueryContext(
        user_id=user_id,
        current_query=query,
        previous_queries=history,
        user_action_seq=action_seq,
        candidates=candidates,
        device=record.device,
        locale=record.locale,
        experiment_id=record.experiment_id,
        request_timestamp=timestamp,
    )
    return ctx

During inference, the entry point at phoenix/xrex/inference/serving_filters_runner.py creates the request object that flows into the scorer's score method:

def score(self, request):
    ctx = self._build_query_context(request)
    scores = self.model.predict(ctx)          # model consumes the hydrated context

    return scores

From Context to Prediction

Once hydrated, the QueryContext is passed to the scoring model defined in phoenix/xrex/models/recsys_two_tower_model.py. The model consumes the complete context—including the user action sequence and historical queries—to compute final relevance scores. By hydrating rich behavioral sequences rather than isolated query-item pairs, the PhoenixScorer enables the underlying model to capture complex session dynamics and temporal user intent patterns.

Summary

  • PhoenixScorer hydrates a QueryContext object before scoring, merging data from PhoenixDataset with real-time request information.
  • The user action sequence contains chronologically ordered events (clicks, impressions, likes) with action_type, item_id, timestamp, and metadata.
  • Historical queries are retrieved via _get_user_history() to provide temporal context beyond the current request.
  • Item features, session metadata (device, locale, experiment ID), and temporal signals complete the context.
  • The concrete implementation resides in grox/flows/reply_spam/classifier_reply_ranking.py, while the dataset wrapper is defined in phoenix/xrex/data/parquet_recsys.py.

Frequently Asked Questions

What specific fields does each user action contain in the PhoenixScorer context?

Each action in the user_action_seq contains four core fields: action_type (such as click, impression, like, or scroll), item_id (the identifier of the ranked item), timestamp (when the event occurred), and optional metadata for supplementary engagement data. These are parsed from PhoenixDataset records via the _parse_actions method.

How does PhoenixScorer retrieve a user's historical queries?

The scorer calls self._get_user_history(user_id), which queries the PhoenixDataset to extract prior query entries associated with the user ID. This historical log is appended to the QueryContext as previous_queries, enabling the model to recognize patterns across multiple requests in the session.

Which source file contains the concrete implementation of the context hydration logic?

The concrete implementation of _build_query_context and the scorer's hydration logic is found in grox/flows/reply_spam/classifier_reply_ranking.py within the ReplyScorer class. The underlying PhoenixDataset wrapper used during hydration is defined in phoenix/xrex/data/parquet_recsys.py.

Why does the scorer include timestamps in the query context?

The request_timestamp is captured using datetime.utcnow() and attached to the context to enable time-aware scoring. This allows the model to apply temporal decay to older actions, detect session boundaries, and weight recent user behavior more heavily than historical interactions when computing relevance scores.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →