How to Debug Load Shedding Issues in the Recos Decider
Developers can debug load shedding in the Recos Decider by verifying the dynamically generated flag keys, inspecting live fractional values via the admin endpoint, checking overlay file precedence, and adding targeted logging around the EndpointLoadShedder.apply method.
The Twitter recommendation algorithm uses a sophisticated throttling mechanism to protect the Recos (recommendations) stack from overload. Understanding how to debug load shedding issues is critical when requests are unexpectedly dropped or when testing new capacity limits. In the twitter/the-algorithm repository, the Recos Decider implements fractional flag-based load shedding through the EndpointLoadShedder class, which dynamically constructs flag names and randomly samples traffic based on YAML configuration values.
Understanding the Load Shedding Architecture
The Recos stack controls request throttling through Decider flag keys defined in YAML configuration files. The architecture relies on several coordinated components to determine whether a specific request should be served or dropped.
Decider Configuration and YAML Files
Load shedding behavior is governed by fractional flags defined in com/twitter/recos/config/decider.yml and optional overlay files. These configurations specify percentage values (0.0 to 100.0) that determine what fraction of requests should be rejected. A value of 50.0 instructs the system to drop half of all requests to a specific endpoint by throwing LoadSheddingException.
The flag naming convention follows the pattern: enable_loadshedding_<graphPrefix>_<endpoint>. For example, the relatedTweets endpoint in the user-tweet graph uses the key enable_loadshedding_user-tweet-graph_relatedTweets.
Core Components
Three primary Scala classes implement the decision logic:
-
BaseDecider– Creates acom.twitter.decider.Deciderinstance from base and overlay configurations. Located insrc/scala/com/twitter/recos/decider/BaseDecider.scala, this class provides the foundation for all decider operations. -
GraphDecider– A trait that adds agraphNamePrefixproperty used to build the full flag name. Specializations likeUserTweetGraphDeciderandUserUserGraphDeciderprovide specific prefixes such as"user-tweet-graph"or"user-user-graph". -
EndpointLoadShedder– The throttling executor located insrc/scala/com/twitter/recos/decider/EndpointLoadShedder.scala. Itsapplymethod constructs the flag key from the prefix, graph name, and endpoint name, then queries theGraphDeciderto determine if the request should be dropped.
Why Load Shedding Debugging Is Challenging
Several factors make load shedding issues difficult to diagnose:
- Fractional randomization – The decision to drop a request depends on random sampling performed by the Decider, causing intermittent failures that are hard to reproduce consistently.
- Dynamic key generation – The flag key is assembled at runtime using string concatenation (
enable_loadshedding_${graphPrefix}_${endpoint}), meaning typos or incorrect prefixes silently disable shedding without errors. - Overlay precedence – Runtime overlay files can override base configuration values, making the effective load shedding percentage different from what appears in the local
decider.ymlfile.
Step-by-Step Debugging Strategies
Follow these targeted steps to isolate and resolve load shedding misconfigurations.
Verify the Flag Key Construction
Confirm the exact key used for your endpoint by examining the key generation logic. In EndpointLoadShedder.scala (line 30), the key is built using:
val key = s"${keyPrefix}_${decider.graphNamePrefix}_${endpointName}"
For the relatedTweets endpoint using UserTweetGraphDecider, the resulting key is enable_loadshedding_user-tweet-graph_relatedTweets. Verify that your graphNamePrefix matches the configuration file exactly, including hyphens and case sensitivity.
Inspect Live Flag Values
Query the current fractional value using the Decider admin endpoint to see the effective percentage in production:
curl http://<service>:9990/admin/decider/enable_loadshedding_user-tweet-graph_relatedTweets
This returns the active value, accounting for any runtime overlays that may differ from your local configuration.
Check Overlay Precedence
Determine whether an overlay file is overriding the base value. Examine the overlay file referenced by the decider configuration, typically located at paths like /usr/local/config/overlays/recos/service/prod/atla/decider_overlay.yml. Compare the flag value there against the base decider.yml to identify which configuration layer is active.
Add Temporary Logging
Instrument the EndpointLoadShedder.apply method to log decision outcomes before throwing exceptions. This reveals whether the decider is actually being queried and what key is being used:
def apply[T](endpointName: String)(serve: => Future[T]): Future[T] = {
val key = s"${keyPrefix}_${decider.graphNamePrefix}_${endpointName}"
if (decider.isAvailable(key, Some(RandomRecipient))) {
log.info(s"Load-shedding $key enabled – dropping request")
Future.exception(LoadSheddingException)
} else serve
}
Enable Deterministic Testing
Force predictable behavior in test environments by creating a test-specific overlay file (e.g., decider_test.yml). Set the flag to 100.0 to always drop requests or 0.0 to never drop them. Point BaseDecider.overlayConfig to this file during testing to eliminate randomness.
Simulate Failures Locally
Mock the GraphDecider interface to return deterministic values and verify your service handles LoadSheddingException correctly:
val mockDecider = mock[GraphDecider]
when(mockDecider.graphNamePrefix).thenReturn("user-tweet-graph")
when(mockDecider.isAvailable(any[String], any[Option[Recipient]]))
.thenReturn(true) // Force shedding
val shedder = new EndpointLoadShedder(mockDecider)
val result = shedder("relatedTweets") { Future.Unit }
assert(result.isThrow) // Verify LoadSheddingException is thrown
Validate Metrics
Ensure your service emits a counter metric when shedding occurs. If missing, add instrumentation inside the shedding block:
if (decider.isAvailable(key, Some(RandomRecipient))) {
statsReceiver.counter("recos_load_shed").incr()
Future.exception(LoadSheddingException)
} else serve
Search the codebase for recos_load_shed to verify existing metric collection.
Reproduce in Staging
Apply the same overlay configuration to a staging cluster and generate traffic to the specific endpoint. Monitor success rates, latency, and logs to confirm the shedding behavior matches expectations before applying changes to production.
Practical Code Examples
Integrating EndpointLoadShedder
Wrap your service logic with the load shedder to enable per-endpoint throttling:
import com.twitter.recos.decider.{EndpointLoadShedder, UserTweetGraphDecider}
import com.twitter.util.Future
class RelatedTweetService(env: String, dc: String) {
private val shedder = new EndpointLoadShedder(UserTweetGraphDecider(env, dc))
def getRelated(req: Request): Future[Response] =
shedder("relatedTweets") {
// Normal graph computation logic
computeRelated(req)
}
}
Unit Testing with Mocked Deciders
Create deterministic tests for load shedding behavior:
test("load shedding aborts request when flag is enabled") {
val mockDecider = mock[GraphDecider]
when(mockDecider.graphNamePrefix).thenReturn("user-tweet-graph")
when(mockDecider.isAvailable(any[String], any[Option[Recipient]]))
.thenReturn(true) // Simulate 100% shedding
val shedder = new EndpointLoadShedder(mockDecider)
val result = shedder("relatedTweets") { Future.Unit }
assertThrows[EndpointLoadShedder.LoadSheddingException] {
Await.result(result)
}
}
Adding Debug Instrumentation
For production debugging, add structured logging to capture the exact key and decision:
def apply[T](endpointName: String)(serve: => Future[T]): Future[T] = {
val key = s"${keyPrefix}_${decider.graphNamePrefix}_${endpointName}"
val shouldShed = decider.isAvailable(key, Some(RandomRecipient))
log.debug(s"Load shedding check: key=$key, shouldShed=$shouldShed")
if (shouldShed) Future.exception(LoadSheddingException)
else serve
}
Summary
- Verify flag key construction in
EndpointLoadShedder.scalato ensure the graph prefix and endpoint name match your YAML configuration exactly. - Query live values via the
/admin/decider/<flag>endpoint to see runtime fractional settings including overlay overrides. - Check overlay files in production paths like
/usr/local/config/overlays/to identify configuration layers affecting your service. - Add temporary logging around the
decider.isAvailablecall to capture the exact key being evaluated and the boolean result. - Mock GraphDecider in unit tests to force deterministic shedding behavior (100% or 0%) and verify exception handling.
- Emit metrics using
statsReceiver.counter("recos_load_shed")to track shedding frequency in production dashboards.
Frequently Asked Questions
How do I find the exact Decider flag key for a specific endpoint?
The flag key is dynamically constructed in EndpointLoadShedder.apply using the pattern enable_loadshedding_${graphPrefix}_${endpoint}. Check the graphNamePrefix value in your specific GraphDecider implementation (e.g., UserTweetGraphDecider uses "user-tweet-graph"), then combine it with your endpoint name. For the relatedTweets endpoint, the complete key is enable_loadshedding_user-tweet-graph_relatedTweets.
Why does my local testing show different shedding behavior than production?
Production environments often use overlay configuration files that override the base decider.yml values. Check the overlay file referenced by your service's BaseDecider.overlayConfig path, typically located in deployment-specific directories like /usr/local/config/overlays/. These runtime overlays can set different fractional values than those committed to the repository.
How can I temporarily disable load shedding for debugging?
Set the specific flag to 0.0 in a local overlay file and configure your test environment to use this overlay via BaseDecider.overlayConfig. Alternatively, mock the GraphDecider.isAvailable method to always return false in your test suite. Never disable shedding in production; instead, use the /admin/decider endpoint to verify current values without modifying code.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →