Configurable Weights in ScoringWeights: How X's Reply-Spam Detection Blends LLM Probabilities with Heuristic Signals
The reply-spam scoring system combines four configurable weights—weight_confidence, weight_follower, weight_content, and weight_history—with the language model's predicted spam probability in a weighted sum formula to produce the final classification score.
The xai-org/x-algorithm repository implements a hybrid spam detection pipeline that merges statistical machine learning with rule-based heuristics. These configurable weights, defined in the system's prompt generation and scoring logic, allow operators to tune detection sensitivity without deploying new code.
The Four Configurable Weights
The scoring architecture exposes four distinct weights that scale different signal components. These values are injected into the Jinja template at grox/flows/reply_spam/templates/reply_scoring_system.j2 via the prompt generator in grox/flows/reply_spam/prompts.py, then applied during final score computation in bdsm/runtime/score_results_sink_focal.py.
weight_confidence
weight_confidence directly multiplies the LLM's raw spam probability (model_predicted_spam_prob). Increasing this value shifts the system toward pure model-based decisions, reducing the influence of heuristic signals like follower count or content analysis. When set significantly higher than other weights, the final score tracks the model's confidence almost exclusively.
weight_follower
weight_follower scales a signal derived from the author's follower count relative to the large_account_follower_threshold parameter. This weight implements a "high-reach account protection" mechanism—when elevated, it requires the model to produce higher predicted probabilities before flagging accounts with substantial followings as spam. The threshold itself is passed into reply_scoring_system_prompt() at runtime.
weight_content
weight_content scales content-quality metrics calculated during preprocessing, such as the presence of disallowed phrases, suspicious URL patterns, or sentiment indicators. Raising this weight increases the penalty for spammy textual content regardless of the posting account's reputation or the model's initial assessment.
weight_history
weight_history incorporates behavioral data about the user's past interactions, including prior spam flags or enforcement actions. A higher value creates a recidivism penalty that makes the system more sensitive to repeat offenders, allowing historical bad behavior to override model uncertainty.
How Weights Interact with Predicted Probabilities
The interaction between these configurable weights and the LLM's predicted probability follows a linear weighted-sum formula implemented in the runtime scorer. The system retrieves the raw probability from the model inference, then combines it with heuristic signals using the current weight configuration:
# Core scoring logic from bdsm/runtime/score_results_sink_focal.py
final_score = (
weight_confidence * model_predicted_spam_prob
+ weight_follower * follower_signal
+ weight_content * content_quality_signal
+ weight_history * user_history_signal
)
Each signal is normalized or scaled appropriately before entering this calculation. The model_predicted_spam_prob term represents the LLM's direct output—a float between 0 and 1—while the heuristic signals derive from preprocessing pipelines that analyze account metadata and content structure.
The final boolean decision emerges from a threshold comparison (default 0.5):
is_spam = final_score > 0.5
Adjusting any weight modifies the decision boundary in probability space. For example, increasing weight_follower from 0.3 to 0.8 effectively raises the required model_predicted_spam_prob for high-follower accounts to trigger the > 0.5 condition, thereby reducing false positives on popular accounts.
Configuration and Calibration Strategies
Default weight values reside in phoenix/xrex/configs/xrecsys.py, where they are loaded at service startup via environment variables or YAML configuration. This separation of configuration from code allows production teams to A/B test sensitivity levels without deployment pipelines.
Practical tuning follows specific semantic patterns:
- Higher
weight_confidence→ Trust the LLM over heuristics; reduces dependency on potentially brittle rule-based signals. - Higher
weight_follower→ Protect influencer-tier accounts; increases the evidentiary bar for high-reach users. - Higher
weight_content→ Aggressive filtering of spammy keywords and patterns; prioritizes message content over sender reputation. - Higher
weight_history→ Strict enforcement on repeat violators; historical flags compound current assessment.
The following example demonstrates building a prompt with custom weights and applying the scoring formula:
from grox.flows.reply_spam.prompts import reply_scoring_system_prompt
# Configure weights to favor model confidence while protecting large accounts
prompt = reply_scoring_system_prompt(
large_account_follower_threshold=100_000,
weight_confidence=1.5, # Boost model influence
weight_follower=0.3, # Soften follower penalty
weight_content=1.0,
weight_history=0.8,
)
# Execute LLM inference
model_predicted_spam_prob = llm.infer(prompt) # Returns float in [0, 1]
# Calculate final score using identical weights
final_score = (
1.5 * model_predicted_spam_prob
+ 0.3 * follower_signal
+ 1.0 * content_quality_signal
+ 0.8 * user_history_signal
)
# Apply decision threshold
is_spam = final_score > 0.5
Summary
- Four configurable weights control the spam scoring system:
weight_confidence,weight_follower,weight_content, andweight_history. - Weighted-sum architecture in
bdsm/runtime/score_results_sink_focal.pyblends LLM probabilities (model_predicted_spam_prob) with heuristic signals. - Template injection occurs in
grox/flows/reply_spam/prompts.py, where weights populate the Jinja templatereply_scoring_system.j2. - Runtime configuration defaults are stored in
phoenix/xrex/configs/xrecsys.pyand modifiable via environment variables. - Decision threshold of
0.5applies to the weighted output, making weight adjustments directly impact false-positive and false-negative rates.
Frequently Asked Questions
What are the specific configurable weights available in the ScoringWeights system?
The system exposes four configurable weights: weight_confidence (scales the LLM's predicted probability), weight_follower (accounts for follower count relative to thresholds), weight_content (measures text quality and spam indicators), and weight_history (incorporates user behavioral history). These are defined in the configuration schema at phoenix/xrex/configs/xrecsys.py and injected into the scoring pipeline via grox/flows/reply_spam/prompts.py.
How does the weight_confidence parameter affect the final spam decision?
The weight_confidence parameter multiplies the raw model probability (model_predicted_spam_prob) in the weighted sum formula implemented in bdsm/runtime/score_results_sink_focal.py. Increasing this weight above the others causes the final score to track closely with the LLM's confidence, effectively reducing the impact of heuristic factors like follower count or content keywords on the classification outcome.
Can the configurable weights be changed without modifying source code?
Yes, all weights are configurable via external YAML configuration or environment variables loaded at service startup from phoenix/xrex/configs/xrecsys.py. The reply_scoring_system_prompt() function accepts these values as arguments and injects them into the Jinja template at runtime, allowing operators to tune spam detection sensitivity without redeploying the xai-org/x-algorithm codebase.
How does weight_follower protect high-reach accounts from false positives?
The weight_follower parameter scales a signal comparing the account's follower count to the large_account_follower_threshold. When this weight is high, the scoring formula adds a significant offset to the final score for large accounts, requiring the LLM to produce a substantially higher model_predicted_spam_prob to push the weighted sum above the 0.5 decision threshold. This mechanic reduces the likelihood of influential accounts being incorrectly flagged as spam due to ambiguous content.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →