How Negative Weights for User Actions Affect Post Scores in the X-Algorithm

Negative weights for user actions subtract value from post risk scores, making it harder for those actions to trigger enforcement thresholds.

In the xai-org/x-algorithm repository, the post-scoring pipeline treats each user interaction as a weighted feature that contributes to a post’s overall risk score. The magnitude and sign of these weights determine whether an action increases or decreases the likelihood of enforcement. Understanding how negative weights influence the final score is essential for interpreting policy configuration and model behavior.

Action Histogram Extraction

The scoring process begins by constructing a raw histogram of user actions. In bdsm/runtime/score_results_sink_focal.py, the _user_action_histogram helper (lines 86‑93) parses the action history for each user, filtering out padding tokens and zero-count entries.

def _user_action_histogram(action_histograms, i):
    if action_histograms and i < len(action_histograms):
        return [
            {"action_type": n, "cnt": int(c)}
            for n, c in action_histograms[i]
            if n != "PAD" and c > 0
        ]
    return None

Each dictionary in the returned list contains an action_type string and an integer cnt representing how many times that action occurred. This histogram serves as the foundation for all downstream score calculations.

Policy-Driven Weight Mapping

Before scores are computed, the system loads enforcement policy parameters from sink_policy.yaml. The _load_policy method in the same file (lines 69‑94) retrieves action-weight tables that map each action type to a floating-point coefficient.

These weights may be positive or negative. While the full weight table resides in internal configuration, the loader expects a mapping where negative values indicate actions that should mitigate rather than amplify risk. The policy also defines per-head thresholds (_tau) that serve as enforcement gates.

Score Aggregation and Negative Weight Impact

During feature construction (conceptually handled in bdsm/runtime/score_layout.py and consumed by the model), each action count is multiplied by its corresponding weight to produce the feature vector fed into the scoring heads. When a weight is negative, the product count × weight becomes a negative value.


# Conceptual feature pipeline showing weight application

for act in user_action_hist:
    weighted_value = act["cnt"] * ACTION_WEIGHTS[act["action_type"]]
    feature_vector.append(weighted_value)   # negative values lower the head score

This negative contribution reduces the raw score output by the model head. Because the final risk score is derived from these weighted features, actions with negative weights actively pull the post’s score downward rather than pushing it toward the enforcement threshold.

Enforcement Gate Logic

The enforcement decision occurs in _decide_hard_enforcement (lines 998‑1035). After computing the total_actions_gate, the system compares each head’s raw score (scores_by_name) against its policy-defined threshold _tau.


# Enforcement check that respects lowered scores

if scores_by_name.get(_h, 0) >= _tau and legit <= _lam:
    dec.enforcement_met = True
    dec.enforcement_head = _h

A sufficiently negative weight can depress the raw score below _tau, causing the condition at lines 1031‑1035 to fail. Consequently, even high-frequency actions will not trigger the enforcement_met flag if their weights are negative enough to keep the aggregated score below the gate.

Summary

  • Negative weights subtract value from the feature vector during score calculation, directly lowering the risk score.
  • The action histogram is built in _user_action_histogram (score_results_sink_focal.py), while weights are loaded from sink_policy.yaml via _load_policy.
  • Lower raw scores make it statistically harder to exceed per-head thresholds (_tau), reducing enforcement likelihood.
  • This mechanism allows policy administrators to designate specific user behaviors as benign signals rather than risk indicators.

Frequently Asked Questions

What happens when a user action has a negative weight in the X-Algorithm?

When an action carries a negative weight, its count is multiplied by a negative coefficient during feature construction. This produces a negative contribution to the model’s input vector, which reduces the raw score output by the scoring head and makes enforcement less likely.

Where are action weights configured in the repository?

Action weights are defined in the enforcement policy YAML file (sink_policy.yaml). The _load_policy method in bdsm/runtime/score_results_sink_focal.py (lines 69‑94) loads these mappings at runtime, supplying the weight table used during feature aggregation.

How does the enforcement decision use weighted scores?

The _decide_hard_enforcement function compares the aggregated score against a threshold _tau. Because negative weights lower the aggregated score, they effectively raise the bar for triggering enforcement, as the score must still exceed _tau despite the downward pressure.

Can negative weights prevent a post from being flagged entirely?

Yes. If the sum of weighted contributions keeps the final score below the enforcement threshold (_tau), the check at lines 1031‑1035 will not set enforcement_met to True, preventing the flag from being raised even if other risk factors are present.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →