How the User Tweet Graph (UTG) Uses Collaborative Filtering and Random Walks for Advanced Tweet Recommendations
The User Tweet Graph (UTG) generates tweet recommendations by combining collaborative filtering signals—such as co-occurrence counting and similarity algorithms—with personalized random walks that traverse user-tweet engagement edges to score candidate content.
The twitter/the-algorithm repository implements UTG as a stateful GraphJet-based service that powers the TweetMixer pipeline. By analyzing recent user interactions (likes, retweets, replies) as edges in a bipartite graph, UTG identifies relevant tweets through both immediate co-occurrence patterns and multi-hop graph traversals.
UTG Architecture and Graph Storage
UTG maintains an in-memory bipartite graph connecting users to tweets through engagement edges. As implemented in the GraphJet backend, the service stores recent interactions to enable real-time recommendation queries.
The graph structure supports two distinct retrieval modes: collaborative filtering through edge co-occurrence analysis, and personalized random walks that simulate user navigation through the engagement network. These modes can operate independently or combine to form hybrid recommendation scores.
Collaborative Filtering Mechanisms
Co-Occurrence Counting with MinCoOccurrenceParam
UTG employs co-occurrence counting to identify tweets that frequently appear together in user engagement histories. When a user interacts with multiple tweets, UTG increments shared counters between those tweet nodes, creating a "users who engaged with X also engaged with Y" signal.
The MinCoOccurrenceParam—exposed in the TweetMixer parameter stack—filters candidate tweets by requiring a minimum threshold of shared engagements with seed tweets. This parameter directly implements classic collaborative filtering by ensuring that only statistically significant co-occurrence patterns surface as recommendations.
Similarity Algorithms (LogCosine and Jaccard)
Beyond simple counting, UTG applies similarity algorithms to weight co-occurrence relationships. The system supports multiple similarity metrics including LogCosine and Jaccard, configured through the SimilarityAlgorithm parameter in the Thrift request.
These algorithms adjust the strength of collaborative filtering signals based on the distribution of engagements across the user base. LogCosine downweights universally popular tweets to surface niche content with specific audience overlap, while Jaccard provides a straightforward intersection-over-union measure of shared user sets.
Personalized Random Walk Implementation
Random Walk Configuration and Parameters
UTG executes random walks starting from seed tweets to discover latent relationships in the engagement graph. According to the user_tweet_graph.thrift definition, the Thrift request accepts:
- A map of seed tweet IDs to weights (
seeds) - The number of walks to execute (
numRandomWalks) - The maximum traversal hops (
maxRandomWalkLength)
Each walk follows outgoing engagement edges (tweet → user → tweet) and can restart at the seed set, producing a visit-frequency distribution across candidate tweets. This distribution becomes the score field in the TweetBasedRelatedTweetResponse.
Capturing Multi-Hop Relationships
Unlike direct co-occurrence, random walks capture higher-order relationships through 2-hop, 3-hop, or longer traversals. This allows UTG to surface tweets that share no direct engaged users with the seed set but reside in the same "neighborhood" of the graph.
The degreeExponent parameter optionally weights nodes by their degree during the walk, while minScore thresholds filter out low-probability candidates. This approach discovers latent interests that explicit collaborative filtering might miss.
End-to-End Flow in TweetMixer
The TweetMixer service orchestrates UTG queries through a four-stage pipeline:
- Query Transformation –
UTGTweetBasedQueryTransformer.scalareads pipeline parameters and buildsUTGTweetBasedRequest, mappingMinCoOccurrenceParamandSimilarityAlgorithminto the request structure. - Key Generation –
UserTweetGraphTweetBasedCandidateSource.scalacreatesTweetBasedRelatedTweetRequestinstances for each seed tweet, embedding parameters likeminScoreanddegreeExponent. - Service Invocation – The Thrift client calls
tweetBasedRelatedTweetsonUserTweetGraph.MethodPerEndpoint, executing random walks and applying collaborative filtering thresholds on the in-memory graph. - Post-Processing – Results convert to
TweetMixerCandidateobjects with preservedscoreandseedIdmetadata for downstream ranking.
Code Implementation Examples
Building a UTG Request in TweetMixer
import com.twitter.recos.user_tweet_graph.thriftscala.RelatedTweetSimilarityAlgorithm
import com.twitter.tweet_mixer.candidate_source.UTG.UTGTweetBasedRequest
val utgRequest = UTGTweetBasedRequest(
seedTweetIds = Seq(123456789L, 987654321L),
maxResults = Some(50),
minCooccurrence = Some(2),
minScore = Some(0.5),
maxTweetAgeInHours = Some(24),
similarityAlgorithm = Some(RelatedTweetSimilarityAlgorithm.LogCosine),
enableCache = true,
degreeExponent = Some(0.75)
)
Invoking the Candidate Source
import com.twitter.tweet_mixer.candidate_source.UTG.UserTweetGraphTweetBasedCandidateSource
import com.twitter.stitch.Stitch
val candidateSource: UserTweetGraphTweetBasedCandidateSource = ...
val candidatesStitch: Stitch[Seq[TweetMixerCandidate]] = candidateSource.get(utgRequest)
candidatesStitch.map { candidates =>
candidates.take(10).foreach { c =>
println(s"Tweet ${c.id} (score=${c.score}) from seed ${c.seedId}")
}
}
Query Transformer Logic
val request = UTGTweetBasedRequest(
filteredPosts,
maxResults = Some(params(MaxCandidateNumPerSourceKeyParam)),
minCooccurrence = Some(params(MinCoOccurrenceParam)),
minScore = Some(params(minScoreParam)),
maxTweetAgeInHours = Some(params(MaxTweetAgeHoursParam).inHours),
similarityAlgorithm =
Some(SimilarityAlgorithmEnum.enumToSimilarityAlgorithmMap(params(SimilarityAlgorithm))),
enableCache = inputQuery.params(EnableUTGCacheParam),
degreeExponent = Some(params(degreeExponent))
)
Summary
- User Tweet Graph (UTG) operates as a stateful GraphJet service storing user-tweet engagement edges for real-time recommendation queries.
- Collaborative filtering relies on
MinCoOccurrenceParamthresholds and similarity algorithms (LogCosine, Jaccard) to filter candidates based on shared engagement patterns. - Random walks utilize
numRandomWalks,maxRandomWalkLength, and seed tweet weights to generate personalized visit-frequency scores capturing multi-hop relationships. - The TweetMixer pipeline integrates UTG through
UTGTweetBasedQueryTransformerandUserTweetGraphTweetBasedCandidateSource, which convert pipeline parameters into Thrift requests for thetweetBasedRelatedTweetsendpoint. - Results combine into
TweetMixerCandidateobjects with preservedscoreandseedIdmetadata for downstream ranking.
Frequently Asked Questions
What is the User Tweet Graph (UTG) in Twitter's recommendation system?
UTG is a stateful recommendation service built on GraphJet that stores recent engagement edges between users and tweets. It powers the TweetMixer component of the Twitter algorithm by providing candidate tweets based on existing engagement patterns, operating as a hybrid engine that blends collaborative filtering with graph-based random walks.
How does UTG combine collaborative filtering with random walks?
UTG applies collaborative filtering through co-occurrence counting and similarity algorithms to identify tweets with overlapping user engagement. Simultaneously, it executes personalized random walks from seed tweets to discover latent connections through multi-hop traversals. The system can filter candidates using the MinCoOccurrenceParam while ranking them via random-walk visit frequencies, creating a hybrid scoring mechanism.
What parameters control the random walk behavior in UTG?
The random walk behavior is configured through the Thrift request parameters defined in user_tweet_graph.thrift: seeds (a map of seed tweet IDs to weights), numRandomWalks (total walks to execute), and maxRandomWalkLength (maximum hops per walk). Additional tuning parameters include degreeExponent for node weighting and minScore for filtering low-probability results.
How does MinCoOccurrenceParam affect recommendation quality?
MinCoOccurrenceParam establishes a threshold for the minimum number of shared engagements required between a candidate tweet and the seed set. Higher values restrict recommendations to tweets with statistically significant co-occurrence patterns, reducing noise but potentially limiting discovery. Lower values increase candidate volume but may surface less relevant content with weak collaborative filtering signals.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →