How SimClusters Groups Accounts and Posts by Engagement for Candidate Discovery
SimClusters groups accounts and posts by encoding engagement signals into sparse cluster embeddings, then retrieves candidates by computing approximate cosine similarity between source embeddings and pre-indexed cluster-tweet associations.
SimClusters is the core recommendation engine in the xai-org/x-algorithm repository that transforms user-post interaction data into actionable candidate recommendations. By grouping accounts and posts by engagement patterns into dense vector representations, the system enables efficient large-scale candidate discovery for downstream ranking pipelines.
The Core Mechanism: From Engagement to Embeddings
The foundation of SimClusters is the SimClustersEmbedding—a sparse vector that maps ClusterId values to engagement scores. This embedding captures how a specific account or post relates to distinct communities (clusters) of users with similar interaction patterns.
Building Engagement-Based Embeddings
Engagement signals such as likes, retweets, and favorites are aggregated to construct separate embedding stores for users and tweets. In UserInterestedInReadableStore.scala and TweetEmbeddingGenerationJob, raw engagement data is compressed into a map structure where each dimension represents a cluster of accounts that share similar engagement behaviors.
These embeddings are stored in read-only stores like simClustersEmbeddingStore and retrieved during candidate generation to represent the "interests" of a source account or the "audience" of a source tweet.
Truncation and Top-N Cluster Selection
When processing a candidate request, SimClusters does not scan all clusters. Instead, it truncates the source embedding to the top-N clusters based on engagement magnitude. In SimClustersANNCandidateSource.scala (lines 69-71), the code calls sourceEmbedding.truncate(config.maxScanClusters).getClusterIds() to isolate the most relevant clusters for the current query.
This truncation step is critical for performance, limiting the search space to config.maxScanClusters (typically 200) rather than the full embedding dimension.
The Candidate Discovery Pipeline
The retrieval process follows a multi-stage pipeline that transforms a source embedding into a ranked list of tweet candidates.
Step 1: Source Embedding Lookup
Candidate generation begins in SimClustersANNCandidateSource.get (lines 34-36), where the system fetches the source embedding using a SimClustersEmbeddingId. This lookup retrieves the sparse vector representation that encodes the source account's or tweet's engagement history across all known clusters.
If no embedding exists for the source, the pipeline returns an empty result immediately.
Step 2: Per-Cluster Tweet Retrieval
For each cluster ID selected during truncation, the system queries clusterTweetCandidatesStore to retrieve pre-indexed tweets. As implemented in SimClustersANNCandidateSource.scala (lines 75-80), this store returns tuples of (TweetId, score, favSquaredL2Norm) representing tweets that have received significant engagement from users within that specific cluster.
These raw scores reflect the intrinsic popularity of tweets within their respective engagement communities.
Step 3: Approximate Cosine Similarity Scoring
The heavy lifting occurs in ApproximateCosineSimilarity.apply (lines 84-85), which aggregates per-cluster tweet scores and weights them by the source's cluster engagement scores. This effectively computes a dot-product approximation of cosine similarity between the source embedding and candidate tweets.
The implementation also enforces strict engagement-based filters:
- Temporal constraints: Tweets must fall within
earliestTweetIdandlatestTweetIdwindows - Minimum score threshold: Candidates must exceed
config.minScore - Fav-L2-norm requirements: For algorithms like
CosineSimilarityFavL2NormBasedWithExploration, tweets must satisfy a minimum favorite-based L2 norm threshold (config.engagementThreshold)
Step 4: Final Ranking and Filtering
After aggregation, scores undergo normalization according to the selected ScoringAlgorithm—options include LogCosineSimilarity or CosineSimilarityFavL2NormBasedWithExploration. The final candidate list is filtered by config.minScore, sorted in descending order, and capped at config.maxNumResults before returning as Seq[SimClustersANNTweetCandidate].
Implementation Example
To invoke the candidate source in production code, you construct a query specifying the embedding type, model version, and configuration constraints:
// Build a query targeting a specific source tweet
val query = SimClustersANNCandidateSource.Query(
sourceEmbeddingId = SimClustersEmbeddingId(
internalId = InternalId.TweetId(sourceTweetId),
modelVersion = ModelVersion.Model20m145k2020,
embeddingType = EmbeddingType.FavBasedProducer
),
config = SimClustersANNConfig(
maxScanClusters = 200,
maxTopTweetsPerCluster = 50,
minScore = 0.2,
annAlgorithm = ScoringAlgorithm.CosineSimilarityFavL2NormBasedWithExploration,
maxTweetCandidateAgeHours = 24
// Additional knobs: candidateEmbeddingType, engagementThreshold, etc.
)
)
// Execute the candidate retrieval
val futureCandidates: Future[Option[Seq[SimClustersANNTweetCandidate]]] =
simClustersANNCandidateSource.get(query)
// Process results
futureCandidates.map {
case Some(candidates) =>
candidates.foreach { c =>
println(s"Tweet ${c.tweetId} with score ${c.score}")
}
case None =>
println("No embedding found for the source")
}
Key Source Files and Architecture
According to the xai-org/x-algorithm source code, the following components implement the engagement-driven clustering system:
simclusters/simclustersann/candidate_source/SimClustersANNCandidateSource.scala– Orchestrates embedding lookup, cluster truncation, and candidate assembly.simclusters/simclustersann/candidate_source/ApproximateCosineSimilarity.scala– Performs weighted aggregation of per-cluster scores and applies engagement filters.simclusters/simclusters_v2/summingbird/stores/UserInterestedInReadableStore.scala– Defines storage for user-centric engagement embeddings.simclusters/simclustersann/modules/SimClustersANNCandidateSourceModule.scala– Guice module providing dependency injection for stores and configuration.simclusters/simclustersann/modules/EmbeddingStoreModule.scala– Constructs the read-only stores backing both user and tweet embeddings.
Summary
- SimClusters embeddings represent accounts and posts as sparse vectors mapped to engagement-based clusters, stored in
UserInterestedInReadableStoreand related tweet stores. - Top-N truncation limits candidate discovery to the most relevant clusters (controlled by
maxScanClusters), ensuring efficient retrieval. - Approximate cosine similarity in
ApproximateCosineSimilarity.scalaweights candidates by co-engagement patterns and applies strict quality filters including fav-L2-norm thresholds. - Configurable scoring algorithms allow the system to balance exploration and exploitation when surfacing candidates.
Frequently Asked Questions
What is a SimClusters embedding and how is it structured?
A SimClusters embedding is a sparse vector that maps integer ClusterId values to floating-point engagement scores. Each cluster represents a community of accounts with similar interaction patterns. The embedding captures how strongly a given account or post relates to each community, enabling mathematical similarity comparisons between arbitrary sources and candidates.
How does the system select which clusters to scan for candidates?
The system employs truncation logic in SimClustersANNCandidateSource that selects only the top-N clusters (typically 200) with the highest engagement scores from the source embedding. This is implemented via sourceEmbedding.truncate(config.maxScanClusters).getClusterIds(), which dramatically reduces computational overhead while preserving the strongest semantic signals.
What role does the fav-L2-norm threshold play in candidate filtering?
The fav-L2-norm threshold acts as a quality gate that favors tweets with sustained engagement across multiple users rather than viral outliers. When using CosineSimilarityFavL2NormBasedWithExploration or similar algorithms, the ApproximateCosineSimilarity component filters candidates that fall below config.engagementThreshold, ensuring only content with robust community endorsement enters the final candidate pool.
How does SimClusters handle real-time candidate generation at scale?
SimClusters achieves scalability through pre-computation and aggressive truncation. Engagement embeddings and cluster-tweet mappings are computed in batch jobs (such as TweetEmbeddingGenerationJob) and stored in read-optimized stores. At query time, the system scans only the top clusters (maxScanClusters) and limits results per cluster (maxTopTweetsPerCluster), enabling sub-second candidate retrieval even with billions of potential tweets.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →