Architectural Considerations for the User Video Graph (UVG) in Delivering Video Recommendations

The User Video Graph (UVG) is a high-performance, in-memory graph service built on GraphJet that powers video recommendations through sub-millisecond traversals of user-video engagement edges, exposing results via Finagle Thrift APIs to the TweetMixer and CR-Mixer services.

The User Video Graph (UVG) serves as a critical component within Twitter's open-source recommendation stack, enabling real-time video tweet suggestions by modeling temporal user engagement patterns. According to the twitter/the-algorithm repository, UVG maintains a sliding window of user-video interactions in RAM to deliver low-latency candidate generation for the Home timeline. This article examines the architectural considerations for the User Video Graph (UVG) that enable it to serve relevant video content while maintaining strict latency SLAs.

Service Design and Graph Engine Architecture

UVG is implemented as a Finagle Thrift service built atop GraphJet, an in-memory graph processing library designed for sub-millisecond traversals. The graph stores user-video engagement edges exclusively for the most recent 24-48 hours, with older events evicted via garbage collection to keep memory usage bounded.

This temporal restriction captures the freshest signals—such as likes, video views, and retweets—that correlate with current user interests while preventing unbounded memory growth. As documented in src/scala/com/twitter/recos/user_video_graph/README.md, GraphJet can hold millions of edges in RAM, enabling the high-throughput lookups required for real-time recommendation serving.

Real-Time Ingestion Pipeline

The UVG service consumes engagement events through a Kafka stream generated by Recos-Injector, ensuring near-real-time graph updates without blocking the serving path. This event-driven ingestion decouples event collection from graph mutation, providing resilience against traffic spikes.

The pipeline guarantees that new user-video interactions propagate to the recommendation graph within seconds, though this introduces eventual consistency—UVG may lag slightly behind the latest interactions during high-volume periods.

Thrift Interface and Client Configuration

Clients communicate with UVG using the UserVideoGraph Thrift API, which provides strongly-typed methods such as tweetBasedRelatedTweets and consumersBasedRelatedTweets. The client configuration in UserVideoGraphClientModule.scala defines aggressive latency constraints:

// From UserVideoGraphClientModule.scala
TimeoutConfig.thriftUserVideoGraphClientTimeout = 50.milliseconds

The client operates with a 50 ms request timeout and no retry budget, implementing a fire-and-forget pattern to prevent request pile-up under load. This design prioritizes low latency over request success rates, ensuring that slow UVG queries do not cascade through the recommendation stack.

Tweet-Based Recommendation Pipeline

The UserVideoGraphTweetBasedCandidateSource implements the primary UVG integration for TweetMixer, fetching related videos based on a seed tweet ID. The implementation follows a three-stage pattern:

  1. Key Generation: Builds a deterministic Memcached key from request parameters (tweet ID, score thresholds, maximum results)
  2. Thrift Invocation: Calls uvgClient.tweetBasedRelatedTweets(key) to retrieve scored candidate tuples
  3. Post-Processing: Converts raw (tweetId, score) pairs into TweetMixerCandidate objects

Results are cached in Memcached with a 10-minute TTL (Utils.randomizedTTL(600)) to reduce load on the UVG service while preserving freshness. The randomized jitter prevents cache stampede during popular content surges.

Source: tweet-mixer/server/src/main/scala/com/twitter/tweet_mixer/candidate_source/UVG/UserVideoGraphTweetBasedCandidateSource.scala

Consumer-Based Recommendation Pipeline

The consumer-based variant (UserVideoGraphConsumerBasedCandidateSource) operates similarly but accepts a seed set of user IDs rather than a tweet ID. It invokes uvgClient.consumersBasedRelatedTweets to discover videos consumed by similar users, enabling collaborative filtering-style recommendations.

This path is particularly valuable for coverage expansion scenarios where tweet-based signals are insufficient.

Integration with TweetMixer and CR-Mixer

UVG pipelines are registered as first-class candidate sources within the HomeRecommendedTweetsRecommendationPipelineConfig:

private val uvgPipelineId = CandidatePipelineIdentifier(identifierPrefix + UVGTweetBased)
private val uvgExpansionPipelineId = CandidatePipelineIdentifier(identifierPrefix + UVGExpansionTweetBased)

// Pipeline registration
uvgTweetBasedCandidatePipelineConfigFactory.build(identifierPrefix, signalsFn = tweetBasedSignalFn),
uvgExpansionTweetBasedCandidatePipelineConfigFactory.build(identifierPrefix, signalFn = highQualityTweetBasedSignalFn)

The CR-Mixer service also queries UVG directly via TweetBasedUserVideoGraphSimilarityEngine, which implements coverage expansion logic: when processing older tweets, the engine fetches recent engaged users from tweetEngagedUsersStore and seeds a secondary consumer-based UVG query to improve recall.

Downstream filtering uses IsVideoTweetFilter to ensure only video candidates proceed to ranking, using pipeline identifiers to distinguish UVG-sourced candidates.

Caching Strategy and Performance Optimization

UVG employs a multi-layered caching strategy to meet latency SLAs:

  • Service-level caching: 10-minute Memcached TTL with randomized jitter
  • Client-side timeouts: 50ms hard limit with no retries
  • Graph bounds: 24-48 hour sliding window to limit active memory footprint

These constraints ensure that UVG can serve millions of requests per second while maintaining sub-millisecond graph traversal performance.

Architectural Trade-offs and Design Decisions

The UVG architecture reflects deliberate trade-offs between freshness, latency, and resource utilization:

Decision Benefit Trade-off
In-memory graph (GraphJet) Sub-millisecond latency, high throughput Requires careful GC tuning; memory costs scale with active user base
24-48 hour sliding window Captures short-term interests; bounds graph size Long-tail engagement signals are ignored; may miss rare but high-value historical signals
Kafka-driven ingestion Decouples event collection from serving; provides backpressure resilience Eventual consistency; UVG may lag seconds behind real-time interactions
Fire-and-forget RPC Prevents cascading failures under load Lower success rate for individual requests; requires downstream resilience
Coverage expansion Improves recall for older content by leveraging consumer signals Additional query latency; increased computational load on consumer-based paths

Summary

  • UVG uses GraphJet for sub-millisecond in-memory graph traversals of user-video engagement edges
  • Maintains a 24-48 hour sliding window to bound memory usage while capturing fresh user interest signals
  • Ingests via Kafka from Recos-Injector for decoupled, real-time updates with eventual consistency guarantees
  • Exposes tweet-based and consumer-based APIs via Finagle Thrift with aggressive 50ms timeouts and no retries
  • Employs 10-minute Memcached caching with randomized TTL to reduce service load and prevent stampede effects
  • Integrates with TweetMixer and CR-Mixer through configurable pipeline factories and coverage expansion logic
  • Filters through IsVideoTweetFilter to ensure downstream components receive only video candidates

Frequently Asked Questions

What is the User Video Graph (UVG) in Twitter's recommendation system?

The User Video Graph (UVG) is an in-memory graph service that models relationships between users and video content based on recent engagement signals. Built on the GraphJet library, it stores user-video edges for 24-48 hours in RAM to enable sub-millisecond traversal queries that power video recommendations in the Home timeline. The service exposes these capabilities through Finagle Thrift interfaces consumed by TweetMixer and CR-Mixer.

Why does UVG use an in-memory graph instead of persistent storage?

UVG uses GraphJet's in-memory architecture to achieve sub-millisecond query latency required for real-time recommendation serving. Persistent storage would introduce network and disk I/O latency incompatible with Twitter's strict SLA requirements. The trade-off is memory management complexity—UVG must carefully control graph size through 24-48 hour sliding windows and garbage collection to prevent heap exhaustion.

How does UVG handle real-time engagement updates?

UVG processes real-time updates through a Kafka ingestion pipeline fed by the Recos-Injector service. As users engage with videos (likes, views, retweets), these events stream into UVG and update the graph edges within seconds. This architecture provides near-real-time freshness while decoupling the serving path from write operations, though it introduces eventual consistency where the graph may lag slightly behind the very latest interactions.

What distinguishes tweet-based from consumer-based UVG queries?

Tweet-based queries (tweetBasedRelatedTweets) accept a seed tweet ID and return related video content based on co-occurrence patterns in the engagement graph. Consumer-based queries (consumersBasedRelatedTweets) accept a set of user IDs and return videos consumed by similar users, enabling collaborative filtering. TweetMixer uses both paths directly, while CR-Mixer employs consumer-based queries specifically for coverage expansion when processing older tweets with stale signals.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →