# How to Debug Load Shedding Issues in the Recos Decider

> Effectively debug load shedding issues in the Recos Decider by verifying flag keys, inspecting fractional values, checking overlay precedence, and adding targeted logs. Learn essential techniques for developers.

- Repository: [X (fka Twitter)/the-algorithm](https://github.com/twitter/the-algorithm)
- Tags: debugging-guide
- Published: 2026-03-03

---

**Developers can debug load shedding in the Recos Decider by verifying the dynamically generated flag keys, inspecting live fractional values via the admin endpoint, checking overlay file precedence, and adding targeted logging around the `EndpointLoadShedder.apply` method.**

The Twitter recommendation algorithm uses a sophisticated throttling mechanism to protect the Recos (recommendations) stack from overload. Understanding how to debug load shedding issues is critical when requests are unexpectedly dropped or when testing new capacity limits. In the `twitter/the-algorithm` repository, the Recos Decider implements fractional flag-based load shedding through the `EndpointLoadShedder` class, which dynamically constructs flag names and randomly samples traffic based on YAML configuration values.

## Understanding the Load Shedding Architecture

The Recos stack controls request throttling through **Decider** flag keys defined in YAML configuration files. The architecture relies on several coordinated components to determine whether a specific request should be served or dropped.

### Decider Configuration and YAML Files

Load shedding behavior is governed by fractional flags defined in [`com/twitter/recos/config/decider.yml`](https://github.com/twitter/the-algorithm/blob/main/com/twitter/recos/config/decider.yml) and optional overlay files. These configurations specify percentage values (0.0 to 100.0) that determine what fraction of requests should be rejected. A value of **50.0** instructs the system to drop half of all requests to a specific endpoint by throwing `LoadSheddingException`.

The flag naming convention follows the pattern: `enable_loadshedding_<graphPrefix>_<endpoint>`. For example, the `relatedTweets` endpoint in the user-tweet graph uses the key `enable_loadshedding_user-tweet-graph_relatedTweets`.

### Core Components

Three primary Scala classes implement the decision logic:

- **`BaseDecider`** – Creates a `com.twitter.decider.Decider` instance from base and overlay configurations. Located in [`src/scala/com/twitter/recos/decider/BaseDecider.scala`](https://github.com/twitter/the-algorithm/blob/main/src/scala/com/twitter/recos/decider/BaseDecider.scala), this class provides the foundation for all decider operations.

- **`GraphDecider`** – A trait that adds a `graphNamePrefix` property used to build the full flag name. Specializations like `UserTweetGraphDecider` and `UserUserGraphDecider` provide specific prefixes such as `"user-tweet-graph"` or `"user-user-graph"`.

- **`EndpointLoadShedder`** – The throttling executor located in [`src/scala/com/twitter/recos/decider/EndpointLoadShedder.scala`](https://github.com/twitter/the-algorithm/blob/main/src/scala/com/twitter/recos/decider/EndpointLoadShedder.scala). Its `apply` method constructs the flag key from the prefix, graph name, and endpoint name, then queries the `GraphDecider` to determine if the request should be dropped.

## Why Load Shedding Debugging Is Challenging

Several factors make load shedding issues difficult to diagnose:

- **Fractional randomization** – The decision to drop a request depends on random sampling performed by the Decider, causing intermittent failures that are hard to reproduce consistently.
- **Dynamic key generation** – The flag key is assembled at runtime using string concatenation (`enable_loadshedding_${graphPrefix}_${endpoint}`), meaning typos or incorrect prefixes silently disable shedding without errors.
- **Overlay precedence** – Runtime overlay files can override base configuration values, making the effective load shedding percentage different from what appears in the local [`decider.yml`](https://github.com/twitter/the-algorithm/blob/main/decider.yml) file.

## Step-by-Step Debugging Strategies

Follow these targeted steps to isolate and resolve load shedding misconfigurations.

### Verify the Flag Key Construction

Confirm the exact key used for your endpoint by examining the key generation logic. In [`EndpointLoadShedder.scala`](https://github.com/twitter/the-algorithm/blob/main/EndpointLoadShedder.scala) (line 30), the key is built using:

```scala
val key = s"${keyPrefix}_${decider.graphNamePrefix}_${endpointName}"

```

For the `relatedTweets` endpoint using `UserTweetGraphDecider`, the resulting key is `enable_loadshedding_user-tweet-graph_relatedTweets`. Verify that your `graphNamePrefix` matches the configuration file exactly, including hyphens and case sensitivity.

### Inspect Live Flag Values

Query the current fractional value using the Decider admin endpoint to see the effective percentage in production:

```bash
curl http://<service>:9990/admin/decider/enable_loadshedding_user-tweet-graph_relatedTweets

```

This returns the active value, accounting for any runtime overlays that may differ from your local configuration.

### Check Overlay Precedence

Determine whether an overlay file is overriding the base value. Examine the overlay file referenced by the decider configuration, typically located at paths like [`/usr/local/config/overlays/recos/service/prod/atla/decider_overlay.yml`](https://github.com/twitter/the-algorithm/blob/main//usr/local/config/overlays/recos/service/prod/atla/decider_overlay.yml). Compare the flag value there against the base [`decider.yml`](https://github.com/twitter/the-algorithm/blob/main/decider.yml) to identify which configuration layer is active.

### Add Temporary Logging

Instrument the `EndpointLoadShedder.apply` method to log decision outcomes before throwing exceptions. This reveals whether the decider is actually being queried and what key is being used:

```scala
def apply[T](endpointName: String)(serve: => Future[T]): Future[T] = {
  val key = s"${keyPrefix}_${decider.graphNamePrefix}_${endpointName}"
  if (decider.isAvailable(key, Some(RandomRecipient))) {
    log.info(s"Load-shedding $key enabled – dropping request")
    Future.exception(LoadSheddingException)
  } else serve
}

```

### Enable Deterministic Testing

Force predictable behavior in test environments by creating a test-specific overlay file (e.g., [`decider_test.yml`](https://github.com/twitter/the-algorithm/blob/main/decider_test.yml)). Set the flag to **100.0** to always drop requests or **0.0** to never drop them. Point `BaseDecider.overlayConfig` to this file during testing to eliminate randomness.

### Simulate Failures Locally

Mock the `GraphDecider` interface to return deterministic values and verify your service handles `LoadSheddingException` correctly:

```scala
val mockDecider = mock[GraphDecider]
when(mockDecider.graphNamePrefix).thenReturn("user-tweet-graph")
when(mockDecider.isAvailable(any[String], any[Option[Recipient]]))
  .thenReturn(true) // Force shedding

val shedder = new EndpointLoadShedder(mockDecider)
val result = shedder("relatedTweets") { Future.Unit }

assert(result.isThrow) // Verify LoadSheddingException is thrown

```

### Validate Metrics

Ensure your service emits a counter metric when shedding occurs. If missing, add instrumentation inside the shedding block:

```scala
if (decider.isAvailable(key, Some(RandomRecipient))) {
  statsReceiver.counter("recos_load_shed").incr()
  Future.exception(LoadSheddingException)
} else serve

```

Search the codebase for `recos_load_shed` to verify existing metric collection.

### Reproduce in Staging

Apply the same overlay configuration to a staging cluster and generate traffic to the specific endpoint. Monitor success rates, latency, and logs to confirm the shedding behavior matches expectations before applying changes to production.

## Practical Code Examples

### Integrating EndpointLoadShedder

Wrap your service logic with the load shedder to enable per-endpoint throttling:

```scala
import com.twitter.recos.decider.{EndpointLoadShedder, UserTweetGraphDecider}
import com.twitter.util.Future

class RelatedTweetService(env: String, dc: String) {
  private val shedder = new EndpointLoadShedder(UserTweetGraphDecider(env, dc))

  def getRelated(req: Request): Future[Response] =
    shedder("relatedTweets") {
      // Normal graph computation logic
      computeRelated(req)
    }
}

```

### Unit Testing with Mocked Deciders

Create deterministic tests for load shedding behavior:

```scala
test("load shedding aborts request when flag is enabled") {
  val mockDecider = mock[GraphDecider]
  when(mockDecider.graphNamePrefix).thenReturn("user-tweet-graph")
  when(mockDecider.isAvailable(any[String], any[Option[Recipient]]))
    .thenReturn(true) // Simulate 100% shedding

  val shedder = new EndpointLoadShedder(mockDecider)

  val result = shedder("relatedTweets") { Future.Unit }
  assertThrows[EndpointLoadShedder.LoadSheddingException] {
    Await.result(result)
  }
}

```

### Adding Debug Instrumentation

For production debugging, add structured logging to capture the exact key and decision:

```scala
def apply[T](endpointName: String)(serve: => Future[T]): Future[T] = {
  val key = s"${keyPrefix}_${decider.graphNamePrefix}_${endpointName}"
  val shouldShed = decider.isAvailable(key, Some(RandomRecipient))
  
  log.debug(s"Load shedding check: key=$key, shouldShed=$shouldShed")
  
  if (shouldShed) Future.exception(LoadSheddingException)
  else serve
}

```

## Summary

- **Verify flag key construction** in [`EndpointLoadShedder.scala`](https://github.com/twitter/the-algorithm/blob/main/EndpointLoadShedder.scala) to ensure the graph prefix and endpoint name match your YAML configuration exactly.
- **Query live values** via the `/admin/decider/<flag>` endpoint to see runtime fractional settings including overlay overrides.
- **Check overlay files** in production paths like `/usr/local/config/overlays/` to identify configuration layers affecting your service.
- **Add temporary logging** around the `decider.isAvailable` call to capture the exact key being evaluated and the boolean result.
- **Mock GraphDecider** in unit tests to force deterministic shedding behavior (100% or 0%) and verify exception handling.
- **Emit metrics** using `statsReceiver.counter("recos_load_shed")` to track shedding frequency in production dashboards.

## Frequently Asked Questions

### How do I find the exact Decider flag key for a specific endpoint?

The flag key is dynamically constructed in `EndpointLoadShedder.apply` using the pattern `enable_loadshedding_${graphPrefix}_${endpoint}`. Check the `graphNamePrefix` value in your specific `GraphDecider` implementation (e.g., `UserTweetGraphDecider` uses `"user-tweet-graph"`), then combine it with your endpoint name. For the `relatedTweets` endpoint, the complete key is `enable_loadshedding_user-tweet-graph_relatedTweets`.

### Why does my local testing show different shedding behavior than production?

Production environments often use overlay configuration files that override the base [`decider.yml`](https://github.com/twitter/the-algorithm/blob/main/decider.yml) values. Check the overlay file referenced by your service's `BaseDecider.overlayConfig` path, typically located in deployment-specific directories like `/usr/local/config/overlays/`. These runtime overlays can set different fractional values than those committed to the repository.

### How can I temporarily disable load shedding for debugging?

Set the specific flag to **0.0** in a local overlay file and configure your test environment to use this overlay via `BaseDecider.overlayConfig`. Alternatively, mock the `GraphDecider.isAvailable` method to always return `false` in your test suite. Never disable shedding in production; instead, use the `/admin/decider` endpoint to verify current values without modifying code.