How SimClusters Finds Content from Accounts the Viewer Does Not Follow

SimClusters surfaces tweets from unfollowed accounts by computing Approximate Nearest Neighbour (ANN) cosine similarity between a viewer’s interest embedding and pre‑computed cluster‑tweet vectors, completely bypassing the social follow graph.

The xai-org/x-algorithm repository implements the SimClusters recommendation system used for large‑scale content discovery. Unlike follower‑based feeds that traverse the social graph, SimClusters identifies relevant tweets by matching viewer interest embeddings against clustered content indexes in dense vector space. This article examines the specific retrieval pipeline in SimClustersANNCandidateSource that enables discovery of content from accounts the viewer does not follow through embedding‑based approximate similarity.

The SimClustersANN Retrieval Pipeline

The system discovers content through three decoupled layers that never inspect the viewer’s follow relationships:

  • Embedding Store – Persists dense interest vectors derived from implicit signals such as favorites, retweets, and log‑fav events.
  • Similarity Engine – Executes approximate cosine similarity via configurable algorithm implementations.
  • Cluster‑Tweet Index – Maps cluster identifiers to tweet candidates with pre‑computed similarity scores.

This architecture ensures that candidate retrieval depends solely on vector proximity rather than social connections.

How the Candidate Source Bypasses Follow Relationships

Embedding Retrieval via ReadableStore

The pipeline begins in SimClustersANNCandidateSource by fetching the viewer’s embedding from a ReadableStore[SimClustersEmbeddingId, SimClustersEmbedding]. The store, provided by EmbeddingStoreModule.scala, returns a dense vector representing the user’s interests without referencing who they follow.

The embedding identifier is constructed as:

val embeddingId = SimClustersEmbeddingId(
  userId = viewerId,
  modelVersion = Model20m145k2020
)

Approximate Cosine Similarity Computation

Rather than filtering by follow status, the source computes similarity using the algorithm specified by the approximate_cosine_similarity flag. Available implementations include ApproximateCosineSimilarity, OptimizedApproximateCosineSimilarity, and ExperimentalApproximateCosineSimilarity, each trading recall for latency characteristics.

The similarity step compares the viewer’s vector against cluster centroids to identify the nearest neighbor clusters in the high‑dimensional embedding space.

Cluster‑Tweet Candidate Resolution

For each similar cluster returned, the system queries a ReadableStore[ClusterId, Seq[(TweetId, Double, Double)]] provided by ClusterTweetIndexProviderModule.scala. This mapping contains tweet IDs paired with their similarity scores and timestamps, aggregated from global interaction patterns rather than per‑user follow lists.

Candidate Generation

The final step assembles SimClustersANNTweetCandidate objects from the retrieved tweet IDs. Because the selection relies purely on vector proximity within the embedding space, the resulting candidates may originate from any account in the index, including those the viewer has never followed.

Implementation in the Source Code

The wiring occurs in SimClustersANNCandidateSourceModule.scala, which injects the embedding store, cluster‑tweet store, and similarity implementation into SimClustersANNCandidateSource. The controller layer in SimClustersANNController.scala exposes this functionality through Thrift and HTTP endpoints.

Key Source Files

Practical Query Examples

Thrift Service Call (Scala)

import com.twitter.simclustersann.thriftscala.{
  SimClustersANNService, 
  SimClustersANNConfig, 
  ScoringAlgorithm
}

val embeddingId = SimClustersEmbeddingId(
  userId = viewerId, 
  modelVersion = Model20m145k2020
)

val request = SimClustersANNService.GetTweetCandidates.Args(
  sourceEmbeddingId = embeddingId,
  config = SimClustersANNConfig(
    numResults = 20,
    annAlgorithm = ScoringAlgorithm.CosineSimilarity
  )
)

val candidatesFuture = simClustersAnnService.getTweetCandidates(request)

HTTP Endpoint Usage

curl -X POST https://api.simclusters-ann/getTweetCandidates \
  -H "Content-Type: application/json" \
  -d '{
    "sourceEmbeddingId": {
      "userId": 12345,
      "modelVersion": "Model20m145k2020"
    },
    "config": {
      "numResults": 20,
      "annAlgorithm": "CosineSimilarity"
    }
  }'

Summary

  • SimClustersANN retrieves content using vector similarity in embedding space rather than traversing follow relationships.
  • The embedding store (ReadableStore[SimClustersEmbeddingId, SimClustersEmbedding]) provides interest vectors derived from implicit user signals.
  • Approximate cosine similarity algorithms match viewers to content clusters without social graph constraints.
  • The cluster‑tweet index (ReadableStore[ClusterId, Seq[(TweetId, Double, Double)]]) supplies candidate tweets from global clusters.
  • No follow filtering occurs in the retrieval pipeline; candidates may originate from any account represented in the embedding space.

Frequently Asked Questions

What data sources populate the SimClusters embedding store?

The embeddings are computed from implicit user interactions including favorites, retweets, and log‑fav events. These signals generate dense vectors in EmbeddingStoreModule.scala that represent interest affinities rather than social connections.

How does SimClustersANN differ from graph-based recommendation systems?

Traditional systems traverse the follow graph to find candidates from followed accounts. SimClustersANN uses approximate cosine similarity on pre‑computed embeddings to find geometrically similar clusters, enabling discovery from accounts the viewer does not follow.

Which similarity algorithms are available in the SimClusters implementation?

The system supports three implementations selected via the approximate_cosine_similarity flag: ApproximateCosineSimilarity, OptimizedApproximateCosineSimilarity, and ExperimentalApproximateCosineSimilarity. Each offers different latency‑recall trade‑offs for the nearest neighbor search.

Can the system filter candidates by follow status after retrieval?

While the core SimClustersANNCandidateSource does not filter by follows, downstream consumers of the SimClustersANNTweetCandidate stream may apply additional filtering. However, the retrieval mechanism itself is intentionally agnostic to follow relationships to maximize content discovery.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →