# How SimClusters Groups Accounts and Posts by Engagement for Candidate Discovery

> Discover candidates efficiently with SimClusters. Learn how it groups accounts and posts by engagement using sparse cluster embeddings and cosine similarity for effective candidate discovery.

- Repository: [SpaceXAI Org/x-algorithm](https://github.com/xai-org/x-algorithm)
- Tags: how-to-guide
- Published: 2026-09-10

---

**SimClusters groups accounts and posts by encoding engagement signals into sparse cluster embeddings, then retrieves candidates by computing approximate cosine similarity between source embeddings and pre-indexed cluster-tweet associations.**

SimClusters is the core recommendation engine in the `xai-org/x-algorithm` repository that transforms user-post interaction data into actionable candidate recommendations. By grouping accounts and posts by engagement patterns into dense vector representations, the system enables efficient large-scale candidate discovery for downstream ranking pipelines.

## The Core Mechanism: From Engagement to Embeddings

The foundation of SimClusters is the **SimClustersEmbedding**—a sparse vector that maps `ClusterId` values to engagement scores. This embedding captures how a specific account or post relates to distinct communities (clusters) of users with similar interaction patterns.

### Building Engagement-Based Embeddings

Engagement signals such as likes, retweets, and favorites are aggregated to construct separate embedding stores for users and tweets. In [`UserInterestedInReadableStore.scala`](https://github.com/xai-org/x-algorithm/blob/main/UserInterestedInReadableStore.scala) and `TweetEmbeddingGenerationJob`, raw engagement data is compressed into a map structure where each dimension represents a cluster of accounts that share similar engagement behaviors.

These embeddings are stored in read-only stores like `simClustersEmbeddingStore` and retrieved during candidate generation to represent the "interests" of a source account or the "audience" of a source tweet.

### Truncation and Top-N Cluster Selection

When processing a candidate request, SimClusters does not scan all clusters. Instead, it truncates the source embedding to the top-N clusters based on engagement magnitude. In [`SimClustersANNCandidateSource.scala`](https://github.com/xai-org/x-algorithm/blob/main/SimClustersANNCandidateSource.scala) (lines 69-71), the code calls `sourceEmbedding.truncate(config.maxScanClusters).getClusterIds()` to isolate the most relevant clusters for the current query.

This truncation step is critical for performance, limiting the search space to `config.maxScanClusters` (typically 200) rather than the full embedding dimension.

## The Candidate Discovery Pipeline

The retrieval process follows a multi-stage pipeline that transforms a source embedding into a ranked list of tweet candidates.

### Step 1: Source Embedding Lookup

Candidate generation begins in `SimClustersANNCandidateSource.get` (lines 34-36), where the system fetches the source embedding using a `SimClustersEmbeddingId`. This lookup retrieves the sparse vector representation that encodes the source account's or tweet's engagement history across all known clusters.

If no embedding exists for the source, the pipeline returns an empty result immediately.

### Step 2: Per-Cluster Tweet Retrieval

For each cluster ID selected during truncation, the system queries `clusterTweetCandidatesStore` to retrieve pre-indexed tweets. As implemented in [`SimClustersANNCandidateSource.scala`](https://github.com/xai-org/x-algorithm/blob/main/SimClustersANNCandidateSource.scala) (lines 75-80), this store returns tuples of `(TweetId, score, favSquaredL2Norm)` representing tweets that have received significant engagement from users within that specific cluster.

These raw scores reflect the intrinsic popularity of tweets within their respective engagement communities.

### Step 3: Approximate Cosine Similarity Scoring

The heavy lifting occurs in `ApproximateCosineSimilarity.apply` (lines 84-85), which aggregates per-cluster tweet scores and weights them by the source's cluster engagement scores. This effectively computes a dot-product approximation of cosine similarity between the source embedding and candidate tweets.

The implementation also enforces strict engagement-based filters:
- **Temporal constraints**: Tweets must fall within `earliestTweetId` and `latestTweetId` windows
- **Minimum score threshold**: Candidates must exceed `config.minScore`
- **Fav-L2-norm requirements**: For algorithms like `CosineSimilarityFavL2NormBasedWithExploration`, tweets must satisfy a minimum favorite-based L2 norm threshold (`config.engagementThreshold`)

### Step 4: Final Ranking and Filtering

After aggregation, scores undergo normalization according to the selected `ScoringAlgorithm`—options include `LogCosineSimilarity` or `CosineSimilarityFavL2NormBasedWithExploration`. The final candidate list is filtered by `config.minScore`, sorted in descending order, and capped at `config.maxNumResults` before returning as `Seq[SimClustersANNTweetCandidate]`.

## Implementation Example

To invoke the candidate source in production code, you construct a query specifying the embedding type, model version, and configuration constraints:

```scala
// Build a query targeting a specific source tweet
val query = SimClustersANNCandidateSource.Query(
  sourceEmbeddingId = SimClustersEmbeddingId(
    internalId = InternalId.TweetId(sourceTweetId),
    modelVersion = ModelVersion.Model20m145k2020,
    embeddingType = EmbeddingType.FavBasedProducer
  ),
  config = SimClustersANNConfig(
    maxScanClusters = 200,
    maxTopTweetsPerCluster = 50,
    minScore = 0.2,
    annAlgorithm = ScoringAlgorithm.CosineSimilarityFavL2NormBasedWithExploration,
    maxTweetCandidateAgeHours = 24
    // Additional knobs: candidateEmbeddingType, engagementThreshold, etc.
  )
)

// Execute the candidate retrieval
val futureCandidates: Future[Option[Seq[SimClustersANNTweetCandidate]]] =
  simClustersANNCandidateSource.get(query)

// Process results
futureCandidates.map {
  case Some(candidates) =>
    candidates.foreach { c =>
      println(s"Tweet ${c.tweetId} with score ${c.score}")
    }
  case None =>
    println("No embedding found for the source")
}

```

## Key Source Files and Architecture

According to the `xai-org/x-algorithm` source code, the following components implement the engagement-driven clustering system:

- **[`simclusters/simclustersann/candidate_source/SimClustersANNCandidateSource.scala`](https://github.com/xai-org/x-algorithm/blob/main/simclusters/simclustersann/candidate_source/SimClustersANNCandidateSource.scala)** – Orchestrates embedding lookup, cluster truncation, and candidate assembly.
- **[`simclusters/simclustersann/candidate_source/ApproximateCosineSimilarity.scala`](https://github.com/xai-org/x-algorithm/blob/main/simclusters/simclustersann/candidate_source/ApproximateCosineSimilarity.scala)** – Performs weighted aggregation of per-cluster scores and applies engagement filters.
- **[`simclusters/simclusters_v2/summingbird/stores/UserInterestedInReadableStore.scala`](https://github.com/xai-org/x-algorithm/blob/main/simclusters/simclusters_v2/summingbird/stores/UserInterestedInReadableStore.scala)** – Defines storage for user-centric engagement embeddings.
- **[`simclusters/simclustersann/modules/SimClustersANNCandidateSourceModule.scala`](https://github.com/xai-org/x-algorithm/blob/main/simclusters/simclustersann/modules/SimClustersANNCandidateSourceModule.scala)** – Guice module providing dependency injection for stores and configuration.
- **[`simclusters/simclustersann/modules/EmbeddingStoreModule.scala`](https://github.com/xai-org/x-algorithm/blob/main/simclusters/simclustersann/modules/EmbeddingStoreModule.scala)** – Constructs the read-only stores backing both user and tweet embeddings.

## Summary

- **SimClusters embeddings** represent accounts and posts as sparse vectors mapped to engagement-based clusters, stored in `UserInterestedInReadableStore` and related tweet stores.
- **Top-N truncation** limits candidate discovery to the most relevant clusters (controlled by `maxScanClusters`), ensuring efficient retrieval.
- **Approximate cosine similarity** in [`ApproximateCosineSimilarity.scala`](https://github.com/xai-org/x-algorithm/blob/main/ApproximateCosineSimilarity.scala) weights candidates by co-engagement patterns and applies strict quality filters including fav-L2-norm thresholds.
- **Configurable scoring algorithms** allow the system to balance exploration and exploitation when surfacing candidates.

## Frequently Asked Questions

### What is a SimClusters embedding and how is it structured?

A SimClusters embedding is a sparse vector that maps integer `ClusterId` values to floating-point engagement scores. Each cluster represents a community of accounts with similar interaction patterns. The embedding captures how strongly a given account or post relates to each community, enabling mathematical similarity comparisons between arbitrary sources and candidates.

### How does the system select which clusters to scan for candidates?

The system employs truncation logic in `SimClustersANNCandidateSource` that selects only the top-N clusters (typically 200) with the highest engagement scores from the source embedding. This is implemented via `sourceEmbedding.truncate(config.maxScanClusters).getClusterIds()`, which dramatically reduces computational overhead while preserving the strongest semantic signals.

### What role does the fav-L2-norm threshold play in candidate filtering?

The fav-L2-norm threshold acts as a quality gate that favors tweets with sustained engagement across multiple users rather than viral outliers. When using `CosineSimilarityFavL2NormBasedWithExploration` or similar algorithms, the `ApproximateCosineSimilarity` component filters candidates that fall below `config.engagementThreshold`, ensuring only content with robust community endorsement enters the final candidate pool.

### How does SimClusters handle real-time candidate generation at scale?

SimClusters achieves scalability through pre-computation and aggressive truncation. Engagement embeddings and cluster-tweet mappings are computed in batch jobs (such as `TweetEmbeddingGenerationJob`) and stored in read-optimized stores. At query time, the system scans only the top clusters (`maxScanClusters`) and limits results per cluster (`maxTopTweetsPerCluster`), enabling sub-second candidate retrieval even with billions of potential tweets.