How the User Tweet Graph (UTG) Uses Collaborative Filtering and Random Walks for Advanced Tweet Recommendations

The User Tweet Graph (UTG) generates tweet recommendations by combining collaborative filtering signals—such as co-occurrence counting and similarity algorithms—with personalized random walks that traverse user-tweet engagement edges to score candidate content.

The twitter/the-algorithm repository implements UTG as a stateful GraphJet-based service that powers the TweetMixer pipeline. By analyzing recent user interactions (likes, retweets, replies) as edges in a bipartite graph, UTG identifies relevant tweets through both immediate co-occurrence patterns and multi-hop graph traversals.

UTG Architecture and Graph Storage

UTG maintains an in-memory bipartite graph connecting users to tweets through engagement edges. As implemented in the GraphJet backend, the service stores recent interactions to enable real-time recommendation queries.

The graph structure supports two distinct retrieval modes: collaborative filtering through edge co-occurrence analysis, and personalized random walks that simulate user navigation through the engagement network. These modes can operate independently or combine to form hybrid recommendation scores.

Collaborative Filtering Mechanisms

Co-Occurrence Counting with MinCoOccurrenceParam

UTG employs co-occurrence counting to identify tweets that frequently appear together in user engagement histories. When a user interacts with multiple tweets, UTG increments shared counters between those tweet nodes, creating a "users who engaged with X also engaged with Y" signal.

The MinCoOccurrenceParam—exposed in the TweetMixer parameter stack—filters candidate tweets by requiring a minimum threshold of shared engagements with seed tweets. This parameter directly implements classic collaborative filtering by ensuring that only statistically significant co-occurrence patterns surface as recommendations.

Similarity Algorithms (LogCosine and Jaccard)

Beyond simple counting, UTG applies similarity algorithms to weight co-occurrence relationships. The system supports multiple similarity metrics including LogCosine and Jaccard, configured through the SimilarityAlgorithm parameter in the Thrift request.

These algorithms adjust the strength of collaborative filtering signals based on the distribution of engagements across the user base. LogCosine downweights universally popular tweets to surface niche content with specific audience overlap, while Jaccard provides a straightforward intersection-over-union measure of shared user sets.

Personalized Random Walk Implementation

Random Walk Configuration and Parameters

UTG executes random walks starting from seed tweets to discover latent relationships in the engagement graph. According to the user_tweet_graph.thrift definition, the Thrift request accepts:

  • A map of seed tweet IDs to weights (seeds)
  • The number of walks to execute (numRandomWalks)
  • The maximum traversal hops (maxRandomWalkLength)

Each walk follows outgoing engagement edges (tweet → user → tweet) and can restart at the seed set, producing a visit-frequency distribution across candidate tweets. This distribution becomes the score field in the TweetBasedRelatedTweetResponse.

Capturing Multi-Hop Relationships

Unlike direct co-occurrence, random walks capture higher-order relationships through 2-hop, 3-hop, or longer traversals. This allows UTG to surface tweets that share no direct engaged users with the seed set but reside in the same "neighborhood" of the graph.

The degreeExponent parameter optionally weights nodes by their degree during the walk, while minScore thresholds filter out low-probability candidates. This approach discovers latent interests that explicit collaborative filtering might miss.

End-to-End Flow in TweetMixer

The TweetMixer service orchestrates UTG queries through a four-stage pipeline:

  1. Query Transformation – UTGTweetBasedQueryTransformer.scala reads pipeline parameters and builds UTGTweetBasedRequest, mapping MinCoOccurrenceParam and SimilarityAlgorithm into the request structure.
  2. Key Generation – UserTweetGraphTweetBasedCandidateSource.scala creates TweetBasedRelatedTweetRequest instances for each seed tweet, embedding parameters like minScore and degreeExponent.
  3. Service Invocation – The Thrift client calls tweetBasedRelatedTweets on UserTweetGraph.MethodPerEndpoint, executing random walks and applying collaborative filtering thresholds on the in-memory graph.
  4. Post-Processing – Results convert to TweetMixerCandidate objects with preserved score and seedId metadata for downstream ranking.

Code Implementation Examples

Building a UTG Request in TweetMixer

import com.twitter.recos.user_tweet_graph.thriftscala.RelatedTweetSimilarityAlgorithm
import com.twitter.tweet_mixer.candidate_source.UTG.UTGTweetBasedRequest

val utgRequest = UTGTweetBasedRequest(
  seedTweetIds        = Seq(123456789L, 987654321L),
  maxResults          = Some(50),
  minCooccurrence     = Some(2),
  minScore            = Some(0.5),
  maxTweetAgeInHours  = Some(24),
  similarityAlgorithm = Some(RelatedTweetSimilarityAlgorithm.LogCosine),
  enableCache         = true,
  degreeExponent      = Some(0.75)
)

Source: tweet-mixer/server/src/main/scala/com/twitter/tweet_mixer/candidate_source/UTG/UTGTweetBasedRequest.scala

Invoking the Candidate Source

import com.twitter.tweet_mixer.candidate_source.UTG.UserTweetGraphTweetBasedCandidateSource
import com.twitter.stitch.Stitch

val candidateSource: UserTweetGraphTweetBasedCandidateSource = ...

val candidatesStitch: Stitch[Seq[TweetMixerCandidate]] = candidateSource.get(utgRequest)

candidatesStitch.map { candidates =>
  candidates.take(10).foreach { c =>
    println(s"Tweet ${c.id} (score=${c.score}) from seed ${c.seedId}")
  }
}

Source: tweet-mixer/server/src/main/scala/com/twitter/tweet_mixer/candidate_source/UTG/UserTweetGraphTweetBasedCandidateSource.scala

Query Transformer Logic

val request = UTGTweetBasedRequest(
  filteredPosts,
  maxResults = Some(params(MaxCandidateNumPerSourceKeyParam)),
  minCooccurrence = Some(params(MinCoOccurrenceParam)),
  minScore = Some(params(minScoreParam)),
  maxTweetAgeInHours = Some(params(MaxTweetAgeHoursParam).inHours),
  similarityAlgorithm =
    Some(SimilarityAlgorithmEnum.enumToSimilarityAlgorithmMap(params(SimilarityAlgorithm))),
  enableCache = inputQuery.params(EnableUTGCacheParam),
  degreeExponent = Some(params(degreeExponent))
)

Source: tweet-mixer/server/src/main/scala/com/twitter/tweet_mixer/functional_component/transformer/UTGTweetBasedQueryTransformer.scala

Summary

  • User Tweet Graph (UTG) operates as a stateful GraphJet service storing user-tweet engagement edges for real-time recommendation queries.
  • Collaborative filtering relies on MinCoOccurrenceParam thresholds and similarity algorithms (LogCosine, Jaccard) to filter candidates based on shared engagement patterns.
  • Random walks utilize numRandomWalks, maxRandomWalkLength, and seed tweet weights to generate personalized visit-frequency scores capturing multi-hop relationships.
  • The TweetMixer pipeline integrates UTG through UTGTweetBasedQueryTransformer and UserTweetGraphTweetBasedCandidateSource, which convert pipeline parameters into Thrift requests for the tweetBasedRelatedTweets endpoint.
  • Results combine into TweetMixerCandidate objects with preserved score and seedId metadata for downstream ranking.

Frequently Asked Questions

What is the User Tweet Graph (UTG) in Twitter's recommendation system?

UTG is a stateful recommendation service built on GraphJet that stores recent engagement edges between users and tweets. It powers the TweetMixer component of the Twitter algorithm by providing candidate tweets based on existing engagement patterns, operating as a hybrid engine that blends collaborative filtering with graph-based random walks.

How does UTG combine collaborative filtering with random walks?

UTG applies collaborative filtering through co-occurrence counting and similarity algorithms to identify tweets with overlapping user engagement. Simultaneously, it executes personalized random walks from seed tweets to discover latent connections through multi-hop traversals. The system can filter candidates using the MinCoOccurrenceParam while ranking them via random-walk visit frequencies, creating a hybrid scoring mechanism.

What parameters control the random walk behavior in UTG?

The random walk behavior is configured through the Thrift request parameters defined in user_tweet_graph.thrift: seeds (a map of seed tweet IDs to weights), numRandomWalks (total walks to execute), and maxRandomWalkLength (maximum hops per walk). Additional tuning parameters include degreeExponent for node weighting and minScore for filtering low-probability results.

How does MinCoOccurrenceParam affect recommendation quality?

MinCoOccurrenceParam establishes a threshold for the minimum number of shared engagements required between a candidate tweet and the seed set. Higher values restrict recommendations to tweets with statistically significant co-occurrence patterns, reducing noise but potentially limiting discovery. Lower values increase candidate volume but may surface less relevant content with weak collaborative filtering signals.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →