How the DwellRegret Scoring Mode Adjusts Dwell Time Based on User Feedback
The DwellRegret scoring mode modulates raw dwell time using a sigmoid function of positive engagement predictions and an exponential penalty for negative feedback signals, producing a regret-aware relevance score.
The DwellRegret scoring mode is an alternative ranking strategy in the Home-Mixer pipeline of the xai-org/x-algorithm repository. Instead of relying on traditional weighted sums of feature scores, this mode treats dwell time as a base metric that is multiplicatively adjusted based on how much the model expects the user to engage positively or negatively with a candidate post.
Configuration with DwellRegretWeights
All hyper-parameters for the scoring mode are encapsulated in the DwellRegretWeights struct defined in home-mixer/scorers/ranking_scorer.rs at lines 59-73:
pub(crate) struct DwellRegretWeights {
alpha_favorite: f64,
alpha_reply: f64,
alpha_retweet: f64,
alpha_quote: f64,
alpha_share: f64,
alpha_share_via_dm: f64,
alpha_share_via_copy_link: f64,
neg_not_interested: f64,
neg_block_author: f64,
neg_mute_author: f64,
neg_report: f64,
temperature: f64,
dwell_floor: f64,
}
- Positive-feedback alphas (
alpha_*): Weight coefficients for favorite, reply, retweet, quote, and share prediction probabilities. - Negative-feedback weights (
neg_*): Scaling factors for explicit negative signals including not-interested, block, mute, and report actions. - Temperature: Controls the sharpness of the sigmoid response curve (lower values create steeper transitions).
- Dwell-floor: Enforces a minimum dwell time value to prevent extreme amplification of very short dwells.
These weights are populated from feature switches at query time, allowing dynamic adjustment without code deployment.
Core Computation in compute_dwell_regret_base_scores
When the query parameter rust_home_mixer_value_model_mode is set to dwell_regret_sigmoid, the RankingScorer::score method (lines 16-22 in ranking_scorer.rs) bypasses standard scoring and invokes compute_dwell_regret_base_scores around line 546. This function implements a three-stage modulation pipeline:
Centered Ratios for Positive Feedback
The algorithm first calculates the mean predicted probability for each positive action across the entire candidate set. For each post, it computes a centered ratio measuring deviation from this mean:
fn centered_ratio(p: f64, mean: f64) -> f64 {
if mean < 1e-12 { 0.0 } else { (p / mean) - 1.0 }
}
This helper (defined at lines 13-19) returns zero for negligible means, otherwise returning (p / mean) - 1. The result is positive when the candidate exceeds the average engagement probability and negative when it underperforms.
The positive term aggregates these centered ratios weighted by their respective alphas:
positive = α_fav * centered_ratio(fav_score, mean_fav)
+ α_reply * centered_ratio(reply_score, mean_reply)
+ ... // other positive actions
Negative Feedback Penalties
Explicit negative signals are combined linearly without centering:
negative = neg_not_interested * not_interested_score
+ neg_block_author * block_author_score
+ neg_mute_author * mute_author_score
+ neg_report * report_score
Unlike the positive term, the negative term is clamped to non-positive values (negative.min(0.0)) before entering the final modulation formula.
Sigmoid Modulation Formula
The final multiplier applied to dwell time combines both terms using a sigmoid and exponential transformation:
modulation = 2.0 *
sigmoid(positive / temperature) *
exp(negative.min(0.0) / temperature)
Where sigmoid(x) = 1 / (1 + e⁻ˣ). The raw dwell time is first clamped to the dwell_floor, then multiplied by this modulation factor:
dwell = ps.dwell_time.unwrap_or(0.0).max(w.dwell_floor).max(0.0);
adjusted_score = dwell * modulation
- High positive deviation: Increases the sigmoid output toward 1.0, raising the multiplier up to a maximum of 2.0.
- Negative feedback: Drives the exponential term below 1.0, reducing the final score proportionally.
Gating and Mode Selection
The DwellRegret mode is guarded by a learned linear classifier defined in home-mixer/scorers/value_model_gate.rs. The GateModel struct (lines 99-116) loads weights, bias, threshold, and hysteresis from feature switches. The method serve_new_scoring (lines 132-141) computes a linear score from query features; if the margin falls within a hysteresis band, a deterministic hash of the user ID breaks the tie.
In RankingScorer::score, the match statement on ValueModelMode determines the execution path:
match value_model_mode {
DwellRegretSigmoid => Self::compute_dwell_regret_base_scores(&dr_weights, candidates),
// ... fallback to standard weighted scoring
}
This architecture allows gradual rollout via feature switches—operators can enable the regret model for specific user cohorts while maintaining the legacy scorer as a fallback.
Enabling and Testing the DwellRegret Mode
To activate this scoring mode in a test environment or production query, override the relevant feature switches:
use xai_feature_switches::FeatureSwitches;
let mut query = ScoredPostsQuery::default();
let mut fs = FeatureSwitches::new(vec![]).unwrap();
let mut result = fs.match_recipient(
&xai_feature_switches::RecipientBuilder::new().build()
);
// Enable DwellRegret mode
result.override_fs(
"rust_home_mixer_value_model_mode".to_string(),
"dwell_regret_sigmoid".to_string(),
);
// Configure weights
result.override_fs(
"rust_home_mixer_dwell_regret_alpha_favorite".to_string(),
"0.5".to_string(),
);
result.override_fs(
"rust_home_mixer_dwell_regret_neg_not_interested".to_string(),
"-1.0".to_string(),
);
result.override_fs(
"rust_home_mixer_dwell_regret_temperature".to_string(),
"0.1".to_string(),
);
query.params = result.into();
When the gate model is active, the scorer automatically evaluates eligibility before applying dwell-regret logic. For direct unit testing, invoke the computation explicitly:
let dr_weights = DwellRegretWeights::from_params(&query.params);
let adjusted_scores = RankingScorer::compute_dwell_regret_base_scores(
&dr_weights,
&candidates,
);
// Returns Vec<f64> of regret-adjusted dwell times
Summary
- The DwellRegret scoring mode replaces standard weighted scoring with a dwell-time modulation approach in
home-mixer/scorers/ranking_scorer.rs. - Positive engagement predictions are centered against batch means and passed through a sigmoid function to create a multiplicative boost up to 2.0x.
- Negative feedback signals (not-interested, block, mute, report) exponentially decay the multiplier, shortening effective dwell time.
- The temperature parameter controls sigmoid sensitivity, while dwell_floor prevents inflation of very short dwells.
- A GateModel in
value_model_gate.rsenables graduated rollout via learned feature thresholds and hysteresis-based bucketing.
Frequently Asked Questions
What is the mathematical difference between DwellRegret and standard weighted scoring?
Standard weighted scoring computes a linear combination of feature values (e.g., w₁·fav + w₂·reply). DwellRegret instead uses the raw dwell time as a base and applies a non-linear transformation: dwell × 2 × sigmoid(positive/temp) × exp(negative/temp). This creates a multiplicative rather than additive relationship between predicted engagement and final ranking score.
How does the temperature parameter affect ranking outcomes?
The temperature field in DwellRegretWeights divides the positive and negative terms before applying the sigmoid and exponential functions. Lower temperatures (e.g., 0.01) create steep transitions where small deviations from the mean engagement probability cause large score swings. Higher temperatures (e.g., 1.0) produce smoother, more gradual adjustments to the dwell time multiplier.
Can the DwellRegret mode be enabled for specific users only?
Yes. The system uses a GateModel defined in home-mixer/scorers/value_model_gate.rs that evaluates a linear classifier on per-query features. By adjusting the gate weights and threshold via feature switches, operators can restrict the DwellRegret scoring mode to users who meet specific criteria (e.g., high activity levels or specific geo-locations) while serving the standard scorer to others.
Why is there a dwell floor in the calculation?
The dwell_floor prevents the modulation multiplier from inflating extremely short dwell times into high scores. Without this floor, a post viewed for 50 milliseconds with a high positive prediction could receive an artificially high regret-adjusted score. By clamping to a minimum threshold (e.g., 2.0 seconds), the system ensures that only substantively consumed content benefits from positive engagement predictions.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →