# How SimClusters Finds Content from Accounts the Viewer Does Not Follow

> Discover how SimClusters surfaces tweets from unfollowed accounts. Learn about its ANN cosine similarity approach that bypasses the follow graph for relevant content discovery.

- Repository: [SpaceXAI Org/x-algorithm](https://github.com/xai-org/x-algorithm)
- Tags: deep-dive
- Published: 2026-09-11

---

**SimClusters surfaces tweets from unfollowed accounts by computing Approximate Nearest Neighbour (ANN) cosine similarity between a viewer’s interest embedding and pre‑computed cluster‑tweet vectors, completely bypassing the social follow graph.**

The `xai-org/x-algorithm` repository implements the SimClusters recommendation system used for large‑scale content discovery. Unlike follower‑based feeds that traverse the social graph, SimClusters identifies relevant tweets by matching viewer interest embeddings against clustered content indexes in dense vector space. This article examines the specific retrieval pipeline in `SimClustersANNCandidateSource` that enables discovery of content from accounts the viewer does not follow through embedding‑based approximate similarity.

## The SimClustersANN Retrieval Pipeline

The system discovers content through three decoupled layers that never inspect the viewer’s follow relationships:

- **Embedding Store** – Persists dense interest vectors derived from implicit signals such as favorites, retweets, and log‑fav events.
- **Similarity Engine** – Executes approximate cosine similarity via configurable algorithm implementations.
- **Cluster‑Tweet Index** – Maps cluster identifiers to tweet candidates with pre‑computed similarity scores.

This architecture ensures that candidate retrieval depends solely on vector proximity rather than social connections.

## How the Candidate Source Bypasses Follow Relationships

### Embedding Retrieval via ReadableStore

The pipeline begins in `SimClustersANNCandidateSource` by fetching the viewer’s embedding from a `ReadableStore[SimClustersEmbeddingId, SimClustersEmbedding]`. The store, provided by [`EmbeddingStoreModule.scala`](https://github.com/xai-org/x-algorithm/blob/main/EmbeddingStoreModule.scala), returns a dense vector representing the user’s interests without referencing who they follow.

The embedding identifier is constructed as:

```scala
val embeddingId = SimClustersEmbeddingId(
  userId = viewerId,
  modelVersion = Model20m145k2020
)

```

### Approximate Cosine Similarity Computation

Rather than filtering by follow status, the source computes similarity using the algorithm specified by the `approximate_cosine_similarity` flag. Available implementations include `ApproximateCosineSimilarity`, `OptimizedApproximateCosineSimilarity`, and `ExperimentalApproximateCosineSimilarity`, each trading recall for latency characteristics.

The similarity step compares the viewer’s vector against cluster centroids to identify the nearest neighbor clusters in the high‑dimensional embedding space.

### Cluster‑Tweet Candidate Resolution

For each similar cluster returned, the system queries a `ReadableStore[ClusterId, Seq[(TweetId, Double, Double)]]` provided by [`ClusterTweetIndexProviderModule.scala`](https://github.com/xai-org/x-algorithm/blob/main/ClusterTweetIndexProviderModule.scala). This mapping contains tweet IDs paired with their similarity scores and timestamps, aggregated from global interaction patterns rather than per‑user follow lists.

### Candidate Generation

The final step assembles `SimClustersANNTweetCandidate` objects from the retrieved tweet IDs. Because the selection relies purely on vector proximity within the embedding space, the resulting candidates may originate from any account in the index, including those the viewer has never followed.

## Implementation in the Source Code

The wiring occurs in [`SimClustersANNCandidateSourceModule.scala`](https://github.com/xai-org/x-algorithm/blob/main/SimClustersANNCandidateSourceModule.scala), which injects the embedding store, cluster‑tweet store, and similarity implementation into `SimClustersANNCandidateSource`. The controller layer in [`SimClustersANNController.scala`](https://github.com/xai-org/x-algorithm/blob/main/SimClustersANNController.scala) exposes this functionality through Thrift and HTTP endpoints.

**Key Source Files**

- **[[`SimClustersANNCandidateSourceModule.scala`](https://github.com/xai-org/x-algorithm/blob/main/SimClustersANNCandidateSourceModule.scala)](https://github.com/xai-org/x-algorithm/blob/main/simclusters/simclustersann/modules/SimClustersANNCandidateSourceModule.scala)** – Composes the candidate source with embedding and cluster stores.
- **[[`SimClustersANNCandidateSource.scala`](https://github.com/xai-org/x-algorithm/blob/main/SimClustersANNCandidateSource.scala)](https://github.com/xai-org/x-algorithm/blob/main/simclusters/simclustersann/candidate_source/SimClustersANNCandidateSource.scala)** – Executes approximate cosine similarity and returns tweet candidates.
- **[[`SimClustersANNController.scala`](https://github.com/xai-org/x-algorithm/blob/main/SimClustersANNController.scala)](https://github.com/xai-org/x-algorithm/blob/main/simclusters/simclustersann/controllers/SimClustersANNController.scala)** – Handles `getTweetCandidates` requests and forwards them to the source.
- **[[`ClusterTweetIndexProviderModule.scala`](https://github.com/xai-org/x-algorithm/blob/main/ClusterTweetIndexProviderModule.scala)](https://github.com/xai-org/x-algorithm/blob/main/simclusters/simclustersann/modules/ClusterTweetIndexProviderModule.scala)** – Provides the cluster‑to‑tweet mapping store.
- **[[`EmbeddingStoreModule.scala`](https://github.com/xai-org/x-algorithm/blob/main/EmbeddingStoreModule.scala)](https://github.com/xai-org/x-algorithm/blob/main/simclusters/simclustersann/modules/EmbeddingStoreModule.scala)** – Instantiates the user embedding readable store.

## Practical Query Examples

**Thrift Service Call (Scala)**

```scala
import com.twitter.simclustersann.thriftscala.{
  SimClustersANNService, 
  SimClustersANNConfig, 
  ScoringAlgorithm
}

val embeddingId = SimClustersEmbeddingId(
  userId = viewerId, 
  modelVersion = Model20m145k2020
)

val request = SimClustersANNService.GetTweetCandidates.Args(
  sourceEmbeddingId = embeddingId,
  config = SimClustersANNConfig(
    numResults = 20,
    annAlgorithm = ScoringAlgorithm.CosineSimilarity
  )
)

val candidatesFuture = simClustersAnnService.getTweetCandidates(request)

```

**HTTP Endpoint Usage**

```bash
curl -X POST https://api.simclusters-ann/getTweetCandidates \
  -H "Content-Type: application/json" \
  -d '{
    "sourceEmbeddingId": {
      "userId": 12345,
      "modelVersion": "Model20m145k2020"
    },
    "config": {
      "numResults": 20,
      "annAlgorithm": "CosineSimilarity"
    }
  }'

```

## Summary

- **SimClustersANN** retrieves content using vector similarity in embedding space rather than traversing follow relationships.
- The **embedding store** (`ReadableStore[SimClustersEmbeddingId, SimClustersEmbedding]`) provides interest vectors derived from implicit user signals.
- **Approximate cosine similarity** algorithms match viewers to content clusters without social graph constraints.
- The **cluster‑tweet index** (`ReadableStore[ClusterId, Seq[(TweetId, Double, Double)]]`) supplies candidate tweets from global clusters.
- **No follow filtering** occurs in the retrieval pipeline; candidates may originate from any account represented in the embedding space.

## Frequently Asked Questions

### What data sources populate the SimClusters embedding store?

The embeddings are computed from implicit user interactions including favorites, retweets, and log‑fav events. These signals generate dense vectors in [`EmbeddingStoreModule.scala`](https://github.com/xai-org/x-algorithm/blob/main/EmbeddingStoreModule.scala) that represent interest affinities rather than social connections.

### How does SimClustersANN differ from graph-based recommendation systems?

Traditional systems traverse the follow graph to find candidates from followed accounts. SimClustersANN uses **approximate cosine similarity** on pre‑computed embeddings to find geometrically similar clusters, enabling discovery from accounts the viewer does not follow.

### Which similarity algorithms are available in the SimClusters implementation?

The system supports three implementations selected via the `approximate_cosine_similarity` flag: `ApproximateCosineSimilarity`, `OptimizedApproximateCosineSimilarity`, and `ExperimentalApproximateCosineSimilarity`. Each offers different latency‑recall trade‑offs for the nearest neighbor search.

### Can the system filter candidates by follow status after retrieval?

While the core `SimClustersANNCandidateSource` does not filter by follows, downstream consumers of the `SimClustersANNTweetCandidate` stream may apply additional filtering. However, the retrieval mechanism itself is intentionally agnostic to follow relationships to maximize content discovery.