How Biohub Platform Safety Filters Detect Controlled Pathogen Sequences

Biohub Platform safety filters detect controlled pathogen sequences by cross-referencing submitted protein annotations against curated blacklists of 29,026 high-risk InterPro entries and 58,641 regulated keywords, automatically flagging matches via the potential_sequence_of_concern boolean field.

The Biohub Platform implements a data-driven guardrail system within the Biohub/esm repository to prevent the processing of hazardous biological sequences. This architecture combines static curated databases with real-time tokenization pipelines to screen every submission for pathogen signatures before inference execution.

The Three-Component Guardrail System

The safety mechanism operates through three integrated components that load at initialization and execute during the tokenization phase.

Safety-Filtered InterPro Entry List

The platform maintains a curated Tab-Separated Values (TSV) file containing 29,026 InterPro IDs associated with pathogens, toxins, and other regulated biological functions. This list provides a whitelist of high-risk functional families and domains that trigger immediate flagging when detected in submitted sequences.

In esm/utils/constants/esm3.py, the constant INTERPRO_ENTRY points to the data file:

INTERPRO_ENTRY = (
    "data/entry_list_safety_29026.list"
)

The physical data file resides at esm/data/entry_list_safety_29026.list within the repository structure.

Safety-Filtered Keyword Vocabulary

Beyond structural domains, the system monitors textual functional annotations using a vocabulary of 58,641 safety-filtered keywords (e.g., "controlled", "toxin", "viral"). Each term ships with pre-computed Inverse Document Frequency (IDF) values to enable TF-IDF weighting during vectorization.

The vocabulary and IDF arrays are defined in esm/utils/constants/esm3.py:

InterProQuantizedTokenizer and FunctionTokenDecoder

The InterProQuantizedTokenizer converts a protein’s InterPro IDs and functional keywords into compact token streams using Locally-Sensitive Hashing (L-SH) of TF-IDF vectors. The tokenizer loads safety-filtered lists during initialization at lines 61-73 of esm/tokenization/function_tokenizer.py.

The FunctionTokenDecoder complements this by reading the InterPro safety list at model initialization (lines 58-66 of esm/models/function_decoder.py) to predict the presence of high-risk annotations during server-side inference.

The Detection Pipeline

The guardrail executes automatically through the following request lifecycle:

  1. Client Submission: The user submits a protein via client.generate() or client.encode(), creating an ESMProtein or ESMProteinTensor object.

  2. Flag Initialization: Every request object includes the boolean field potential_sequence_of_concern (default False), defined in esm/sdk/api.py at lines 50-55.

  3. Tokenization: The InterProQuantizedTokenizer encodes functional annotations by reading C.INTERPRO_ENTRY and C.KEYWORDS_VOCABULARY from the safety data paths.

  4. Server-Side Screening: The cloud "forge" service parses the tokenized representation and checks against the curated InterPro list and keyword vocabulary.

  5. Flag and Response: If any high-risk InterPro ID or keyword matches, the service sets potential_sequence_of_concern to True and rejects or warns on the request. The response handling in esm/sdk/forge.py (lines 406-417) extracts this flag from the returned payload.

Code Example: Detecting Controlled Sequences

The following Python implementation demonstrates how the SDK automatically surfaces safety violations:

from esm.sdk import ForgeClient, ProteinInput, GenerationConfig

# Initialize the client (public endpoint)

client = ForgeClient()

# Example sequence mapping to a controlled InterPro entry (e.g., IPR018103)

seq = "MKTIIALSYIFCLVFADYKDDDDK"

# Build request; SDK automatically runs the tokenizer

protein = ProteinInput(id="test", sequence=seq)

# Request embedding generation

config = GenerationConfig(num_steps=10)
result = client.generate([protein], config=config)

# Check safety flag returned from server

if result.potential_sequence_of_concern:
    print("⚠️  Input matches controlled pathogen entry — request blocked.")
else:
    print("✅  Input passed safety filters.")

Under the hood, client.generate() reads esm/data/entry_list_safety_29026.list via the tokenizer, encodes the sequence using L-SH hashed TF-IDF vectors, and receives the boolean safety determination from the server.

Key Implementation Files

Summary

  • Biohub Platform safety filters rely on curated static lists of 29,026 InterPro entries and 58,641 keywords, not runtime heuristics.
  • The InterProQuantizedTokenizer automatically loads safety data from esm/utils/constants/esm3.py during request initialization.
  • Every API request carries a potential_sequence_of_concern field that defaults to False and flips to True upon detection of regulated sequences.
  • Detection occurs through L-SH hashed TF-IDF vector comparison against pre-computed safety databases.
  • Flagged requests are blocked or warned at the server level before biological inference execution.

Frequently Asked Questions

What triggers the potential_sequence_of_concern flag?

The flag triggers when a submitted protein sequence contains InterPro IDs listed in entry_list_safety_29026.list or functional keywords present in keyword_vocabulary_safety_filtered_58641.txt. The system checks both structural domain annotations and textual functional descriptions using TF-IDF weighted matching.

Where does the Biohub Platform store its curated safety lists?

Static safety lists reside in the esm/data/ directory as TSV and text files, while their paths are referenced through constants in esm/utils/constants/esm3.py. The InterPro list contains 29,026 entries, and the keyword vocabulary contains 58,641 terms with corresponding IDF values in NumPy format.

Can developers bypass the safety filters when using the Biohub SDK?

No. The safety check executes server-side within the cloud "forge" service after the client submits tokenized representations. While the SDK initializes the potential_sequence_of_concern field as False locally, the server overwrites this value based on its independent screening against curated safety databases, making client-side bypass impossible.

How does the TF-IDF weighting improve pathogen detection accuracy?

The platform uses pre-computed IDF values stored in keyword_idf_safety_filtered_58641.npy to weight keywords by their specificity within the biological literature. Rare terms associated with toxins or viral functions receive higher weights, reducing false positives from common protein descriptors while maintaining sensitivity to genuine controlled pathogen signatures.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →