# How Biohub Platform Safety Filters Detect Controlled Pathogen Sequences

> Discover how Biohub Platform safety filters detect controlled pathogen sequences using curated blacklists of high-risk entries and regulated keywords to flag potential concerns automatically.

- Repository: [Biohub/esm](https://github.com/Biohub/esm)
- Tags: deep-dive
- Published: 2026-05-30

---

**Biohub Platform safety filters detect controlled pathogen sequences by cross-referencing submitted protein annotations against curated blacklists of 29,026 high-risk InterPro entries and 58,641 regulated keywords, automatically flagging matches via the `potential_sequence_of_concern` boolean field.**

The Biohub Platform implements a data-driven guardrail system within the `Biohub/esm` repository to prevent the processing of hazardous biological sequences. This architecture combines static curated databases with real-time tokenization pipelines to screen every submission for pathogen signatures before inference execution.

## The Three-Component Guardrail System

The safety mechanism operates through three integrated components that load at initialization and execute during the tokenization phase.

### Safety-Filtered InterPro Entry List

The platform maintains a curated Tab-Separated Values (TSV) file containing **29,026 InterPro IDs** associated with pathogens, toxins, and other regulated biological functions. This list provides a whitelist of high-risk functional families and domains that trigger immediate flagging when detected in submitted sequences.

In [`esm/utils/constants/esm3.py`](https://github.com/Biohub/esm/blob/main/esm/utils/constants/esm3.py), the constant `INTERPRO_ENTRY` points to the data file:

```python
INTERPRO_ENTRY = (
    "data/entry_list_safety_29026.list"
)

```

The physical data file resides at `esm/data/entry_list_safety_29026.list` within the repository structure.

### Safety-Filtered Keyword Vocabulary

Beyond structural domains, the system monitors textual functional annotations using a vocabulary of **58,641 safety-filtered keywords** (e.g., "controlled", "toxin", "viral"). Each term ships with pre-computed Inverse Document Frequency (IDF) values to enable TF-IDF weighting during vectorization.

The vocabulary and IDF arrays are defined in [`esm/utils/constants/esm3.py`](https://github.com/Biohub/esm/blob/main/esm/utils/constants/esm3.py):

- `KEYWORDS_VOCABULARY` — maps to [`esm/data/keyword_vocabulary_safety_filtered_58641.txt`](https://github.com/Biohub/esm/blob/main/esm/data/keyword_vocabulary_safety_filtered_58641.txt)
- `KEYWORDS_IDF` — maps to `esm/data/keyword_idf_safety_filtered_58641.npy`

### InterProQuantizedTokenizer and FunctionTokenDecoder

The **InterProQuantizedTokenizer** converts a protein’s InterPro IDs and functional keywords into compact token streams using Locally-Sensitive Hashing (L-SH) of TF-IDF vectors. The tokenizer loads safety-filtered lists during initialization at lines 61-73 of [`esm/tokenization/function_tokenizer.py`](https://github.com/Biohub/esm/blob/main/esm/tokenization/function_tokenizer.py).

The **FunctionTokenDecoder** complements this by reading the InterPro safety list at model initialization (lines 58-66 of [`esm/models/function_decoder.py`](https://github.com/Biohub/esm/blob/main/esm/models/function_decoder.py)) to predict the presence of high-risk annotations during server-side inference.

## The Detection Pipeline

The guardrail executes automatically through the following request lifecycle:

1. **Client Submission**: The user submits a protein via `client.generate()` or `client.encode()`, creating an `ESMProtein` or `ESMProteinTensor` object.

2. **Flag Initialization**: Every request object includes the boolean field `potential_sequence_of_concern` (default `False`), defined in [`esm/sdk/api.py`](https://github.com/Biohub/esm/blob/main/esm/sdk/api.py) at lines 50-55.

3. **Tokenization**: The `InterProQuantizedTokenizer` encodes functional annotations by reading `C.INTERPRO_ENTRY` and `C.KEYWORDS_VOCABULARY` from the safety data paths.

4. **Server-Side Screening**: The cloud "forge" service parses the tokenized representation and checks against the curated InterPro list and keyword vocabulary.

5. **Flag and Response**: If any high-risk InterPro ID or keyword matches, the service sets `potential_sequence_of_concern` to `True` and rejects or warns on the request. The response handling in [`esm/sdk/forge.py`](https://github.com/Biohub/esm/blob/main/esm/sdk/forge.py) (lines 406-417) extracts this flag from the returned payload.

## Code Example: Detecting Controlled Sequences

The following Python implementation demonstrates how the SDK automatically surfaces safety violations:

```python
from esm.sdk import ForgeClient, ProteinInput, GenerationConfig

# Initialize the client (public endpoint)

client = ForgeClient()

# Example sequence mapping to a controlled InterPro entry (e.g., IPR018103)

seq = "MKTIIALSYIFCLVFADYKDDDDK"

# Build request; SDK automatically runs the tokenizer

protein = ProteinInput(id="test", sequence=seq)

# Request embedding generation

config = GenerationConfig(num_steps=10)
result = client.generate([protein], config=config)

# Check safety flag returned from server

if result.potential_sequence_of_concern:
    print("⚠️  Input matches controlled pathogen entry — request blocked.")
else:
    print("✅  Input passed safety filters.")

```

Under the hood, `client.generate()` reads `esm/data/entry_list_safety_29026.list` via the tokenizer, encodes the sequence using L-SH hashed TF-IDF vectors, and receives the boolean safety determination from the server.

## Key Implementation Files

- **[`esm/utils/constants/esm3.py`](https://github.com/Biohub/esm/blob/main/esm/utils/constants/esm3.py)** — Declares paths to safety-filtered InterPro lists and keyword vocabularies.
- **[`esm/tokenization/function_tokenizer.py`](https://github.com/Biohub/esm/blob/main/esm/tokenization/function_tokenizer.py)** — Loads safety data and tokenizes functional annotations (lines 61-73).
- **[`esm/models/function_decoder.py`](https://github.com/Biohub/esm/blob/main/esm/models/function_decoder.py)** — Reads InterPro safety lists at model initialization (lines 58-66).
- **[`esm/sdk/api.py`](https://github.com/Biohub/esm/blob/main/esm/sdk/api.py)** — Defines the `potential_sequence_of_concern` field on request objects (lines 50-55).
- **[`esm/sdk/forge.py`](https://github.com/Biohub/esm/blob/main/esm/sdk/forge.py)** — Handles server communication and extracts safety flags from responses (lines 406-417).
- **`esm/data/entry_list_safety_29026.list`** — Curated TSV of 29,026 high-risk InterPro entries.
- **[`esm/data/keyword_vocabulary_safety_filtered_58641.txt`](https://github.com/Biohub/esm/blob/main/esm/data/keyword_vocabulary_safety_filtered_58641.txt)** — Text file containing 58,641 regulated keywords.
- **`esm/data/keyword_idf_safety_filtered_58641.npy`** — NumPy array of TF-IDF IDF weights for keyword scoring.

## Summary

- Biohub Platform safety filters rely on **curated static lists** of 29,026 InterPro entries and 58,641 keywords, not runtime heuristics.
- The **`InterProQuantizedTokenizer`** automatically loads safety data from [`esm/utils/constants/esm3.py`](https://github.com/Biohub/esm/blob/main/esm/utils/constants/esm3.py) during request initialization.
- Every API request carries a **`potential_sequence_of_concern`** field that defaults to `False` and flips to `True` upon detection of regulated sequences.
- Detection occurs through **L-SH hashed TF-IDF vector comparison** against pre-computed safety databases.
- Flagged requests are blocked or warned at the server level before biological inference execution.

## Frequently Asked Questions

### What triggers the potential_sequence_of_concern flag?

The flag triggers when a submitted protein sequence contains InterPro IDs listed in `entry_list_safety_29026.list` or functional keywords present in [`keyword_vocabulary_safety_filtered_58641.txt`](https://github.com/Biohub/esm/blob/main/keyword_vocabulary_safety_filtered_58641.txt). The system checks both structural domain annotations and textual functional descriptions using TF-IDF weighted matching.

### Where does the Biohub Platform store its curated safety lists?

Static safety lists reside in the `esm/data/` directory as TSV and text files, while their paths are referenced through constants in [`esm/utils/constants/esm3.py`](https://github.com/Biohub/esm/blob/main/esm/utils/constants/esm3.py). The InterPro list contains 29,026 entries, and the keyword vocabulary contains 58,641 terms with corresponding IDF values in NumPy format.

### Can developers bypass the safety filters when using the Biohub SDK?

No. The safety check executes server-side within the cloud "forge" service after the client submits tokenized representations. While the SDK initializes the `potential_sequence_of_concern` field as `False` locally, the server overwrites this value based on its independent screening against curated safety databases, making client-side bypass impossible.

### How does the TF-IDF weighting improve pathogen detection accuracy?

The platform uses pre-computed IDF values stored in `keyword_idf_safety_filtered_58641.npy` to weight keywords by their specificity within the biological literature. Rare terms associated with toxins or viral functions receive higher weights, reducing false positives from common protein descriptors while maintaining sensitivity to genuine controlled pathogen signatures.