# Edge Type Masks for Different Engagement Types in Graphs: Best Practices from Twitter's Algorithm

> Learn best practices for edge type masks at Twitter. Encode engagement types like follows and retweets into integer masks for efficient graph storage and retrieval. Optimize your algorithm today.

- Repository: [X (fka Twitter)/the-algorithm](https://github.com/twitter/the-algorithm)
- Tags: best-practices
- Published: 2026-03-03

---

**Edge type masks encode engagement types—such as follow, favorite, retweet, or video playback—into the high-order bits of a 32-bit integer, allowing GraphJet to store both node identifiers and relationship metadata in a single cache-friendly value.**

In Twitter's open-source recommendation system (the-algorithm), edge type masks are the canonical mechanism for distinguishing between different user interactions within bipartite graphs. By packing type information into the integer itself, the system avoids separate metadata lookups and maintains high-throughput edge storage. The following patterns, derived from the production Scala codebase, define how to implement, validate, and query these masks correctly.

## Why Edge Type Masks Matter

GraphJet stores up to **128 million edges per segment**, which requires 27 bits for node addressing. The remaining 5 bits of a 32-bit integer are available for metadata. The Twitter algorithm reserves the **top four bits** for the edge type, leaving one bit free for future extensions. This design yields two critical benefits:

- **Cache efficiency**: Reading or writing an edge requires only one integer operation, eliminating the indirection of parallel arrays or hash maps.
- **Type safety**: The mask enforces a hard limit of 16 distinct edge types per graph, preventing accidental proliferation of engagement categories.

## Defining Edge Type Masks in Scala

### Mask Constants and Enumeration

Each graph module defines its own mask object extending Scala’s `Enumeration`. The canonical implementation appears in [`UserVideoEdgeTypeMask.scala`](https://github.com/twitter/the-algorithm/blob/main/UserVideoEdgeTypeMask.scala):

```scala
object UserVideoEdgeTypeMask extends Enumeration {
  type UserVideoEdgeTypeMask = Value

  // Only 16 values allowed (4 high-order bits)
  val VideoPlayback50: UserVideoEdgeTypeMask = Value(1)

  // Isolates the low 27 bits for node ID extraction
  val MASK: Int = Integer.parseInt("00001111111111111111111111111111", 2)
  
  // Total number of defined edge types
  val SIZE: Int = values.size
}

```

The `MASK` constant is the bitwise complement of the edge-type region. Applying `node & MASK` strips the high bits, restoring the original node identifier.

### Encoding and Decoding Logic

The `EdgeTypeMask` trait requires three operations: `encode`, `edgeType`, and `restore`. The `UserVideoEdgeTypeMask` class implements these using bit-shifting:

```scala
class UserVideoEdgeTypeMask extends EdgeTypeMask {
  import UserVideoEdgeTypeMask._

  // Pack node ID and edge type into a single Int
  override def encode(node: Int, edgeType: Byte): Int = {
    require(edgeType >= 0 && edgeType <= SIZE,
      s"encode: Illegal edge type $edgeType")
    node | (edgeType << 28)                // Type occupies bits 28-31
  }

  // Extract the edge type byte
  override def edgeType(node: Int): Byte = (node >>> 28).toByte

  // Remove type bits, returning pure node ID
  override def restore(node: Int): Int = node & MASK
}

```

Key implementation details:
- **Left shift (`<< 28`)**: Aligns the 4-bit type with the high-order bits of the integer.
- **Unsigned right shift (`>>>`)**: Prevents sign-extension when extracting the type byte.
- **Validation**: The `require` clause ensures that only defined enum values are encoded, catching mapping errors at the source.

## Mapping Actions to Edge Types

### Centralized Conversion Methods

Graph writers receive Thrift `action` bytes from the ingestion pipeline. Each mask provides a single conversion method to translate these external actions into internal edge-type bytes. In [`UserVideoGraphWriter.scala`](https://github.com/twitter/the-algorithm/blob/main/UserVideoGraphWriter.scala), the mapping looks like this:

```scala
def actionTypeToEdgeType(actionByte: Byte): Byte = {
  Action(actionByte) match {
    case Action.VideoPlayback50 => UserVideoEdgeTypeMask.VideoPlayback50.id
    case _ => throw new IllegalArgumentException(
           s"getEdgeType: Illegal edge type $actionByte")
  }
}

```

This pattern appears consistently across [`UserTweetEdgeTypeMask.scala`](https://github.com/twitter/the-algorithm/blob/main/UserTweetEdgeTypeMask.scala), [`UserEdgeTypeMask.scala`](https://github.com/twitter/the-algorithm/blob/main/UserEdgeTypeMask.scala), and [`ActionEdgeTypeMask.scala`](https://github.com/twitter/the-algorithm/blob/main/ActionEdgeTypeMask.scala). Centralizing the mapping prevents drift between the ingestion schema and the graph storage format.

### Fast Encoding with Pre-computed Arrays

When the set of edge types is small and fixed, the mask can store a pre-computed `EDGEARRAY` to eliminate runtime bit-shifting. [`UserEdgeTypeMask.scala`](https://github.com/twitter/the-algorithm/blob/main/UserEdgeTypeMask.scala) demonstrates this optimization:

```scala
val EDGEARRAY: Array[Int] = Array(
  0,            // unused slot
  1 << 28,      // FOLLOW
  2 << 28,      // MENTION
  3 << 28       // MEDIATAG
  // ...
)

```

The `encode` method then becomes a simple bitwise OR:

```scala
def encode(node: Int, edgeType: Byte): Int = {
  node | EDGEARRAY(edgeType)
}

```

This approach trades a small amount of static memory for faster edge insertion, which is critical for high-throughput user-user graphs.

## Filtering and Querying by Edge Type

Masks expose helper methods to convert client-requested **social-proof types** into the raw byte arrays required by the graph query engine. For example, `UserEdgeTypeMask.getUserUserGraphSocialProofTypes` handles optional filtering:

```scala
def getUserUserGraphSocialProofTypes(
  socialProofTypes: Option[Seq[UserSocialProofType]]
): Array[Byte] = {
  socialProofTypes match {
    case Some(types) => types.map(t => actionTypeToEdgeType(t.id.toByte)).toArray
    case None => (0 until SIZE).map(_.toByte).toArray  // return all types
  }
}

```

If the client supplies no preference, the method returns the full range `(0 until SIZE)`, ensuring backward compatibility while allowing precise query optimization when specific engagement types are requested.

## Summary

- **Reserve four high-order bits** in 32-bit integers for edge type encoding, leaving room for 16 distinct engagement categories and one spare bit for future expansion.
- **Validate aggressively** in `encode` methods to prevent illegal edge type values from corrupting graph segments.
- **Centralize action mapping** via `actionTypeToEdgeType` methods to isolate Thrift schema changes from graph storage logic.
- **Pre-compute EDGEARRAY** for fixed-type graphs to eliminate runtime bit-shift overhead during high-volume edge insertion.
- **Expose restore and edgeType helpers** to ensure graph readers can always recover original node IDs and filter by engagement category.
- **Use social-proof type helpers** to translate client requests into raw byte arrays for efficient graph traversal.

## Frequently Asked Questions

### Why encode edge types in the high bits instead of using separate metadata arrays?

Storing the edge type inside the 32-bit integer allows GraphJet to read or write a complete edge with a single cache line access. Parallel arrays or external hash maps would require additional memory lookups and increase the memory footprint per edge. By packing the type into the unused high bits of the node ID, the system maintains **cache locality** and **atomicity** without sacrificing the 128 million edges-per-segment capacity.

### How many distinct edge types can the mask support?

The current implementation reserves **four bits** for the edge type, which allows for **16 distinct values** (0–15). The codebase deliberately leaves one bit unused (out of the five available high bits) to ensure future extensibility without requiring a migration of existing graph segments. If more than 16 types are needed, the mask width would need to expand, reducing the maximum node ID space accordingly.

### What happens if an invalid edge type is passed to the encoder?

The `encode` method in every mask implementation includes a `require` assertion that validates the edge type byte against the defined `SIZE` of the enumeration. If the value is negative or exceeds the maximum allowed index, the encoder throws an `IllegalArgumentException` immediately. This **fail-fast** behavior prevents corrupted integers from entering the graph, where they could cause silent data loss or incorrect traversal results during recommendation queries.

### How do I add a new engagement type to an existing mask?

First, add the new engagement value to the mask’s `Enumeration` object in the corresponding `*EdgeTypeMask.scala` file, assigning it the next available integer ID. Next, update the `actionTypeToEdgeType` method to map the incoming Thrift `Action` byte to your new edge type ID. If the mask uses a pre-computed `EDGEARRAY` (as in [`UserEdgeTypeMask.scala`](https://github.com/twitter/the-algorithm/blob/main/UserEdgeTypeMask.scala)), append the shifted value `(newId << 28)` to the array. Finally, increment any relevant `SIZE` constants and add unit tests to verify the encode/decode round-trip and the action mapping logic.