How Biohub Platform Safety Filters Detect Controlled Pathogen Sequences
Biohub Platform safety filters detect controlled pathogen sequences by cross-referencing submitted protein annotations against curated blacklists of 29,026 high-risk InterPro entries and 58,641 regulated keywords, automatically flagging matches via the potential_sequence_of_concern boolean field.
The Biohub Platform implements a data-driven guardrail system within the Biohub/esm repository to prevent the processing of hazardous biological sequences. This architecture combines static curated databases with real-time tokenization pipelines to screen every submission for pathogen signatures before inference execution.
The Three-Component Guardrail System
The safety mechanism operates through three integrated components that load at initialization and execute during the tokenization phase.
Safety-Filtered InterPro Entry List
The platform maintains a curated Tab-Separated Values (TSV) file containing 29,026 InterPro IDs associated with pathogens, toxins, and other regulated biological functions. This list provides a whitelist of high-risk functional families and domains that trigger immediate flagging when detected in submitted sequences.
In esm/utils/constants/esm3.py, the constant INTERPRO_ENTRY points to the data file:
INTERPRO_ENTRY = (
"data/entry_list_safety_29026.list"
)
The physical data file resides at esm/data/entry_list_safety_29026.list within the repository structure.
Safety-Filtered Keyword Vocabulary
Beyond structural domains, the system monitors textual functional annotations using a vocabulary of 58,641 safety-filtered keywords (e.g., "controlled", "toxin", "viral"). Each term ships with pre-computed Inverse Document Frequency (IDF) values to enable TF-IDF weighting during vectorization.
The vocabulary and IDF arrays are defined in esm/utils/constants/esm3.py:
KEYWORDS_VOCABULARY— maps toesm/data/keyword_vocabulary_safety_filtered_58641.txtKEYWORDS_IDF— maps toesm/data/keyword_idf_safety_filtered_58641.npy
InterProQuantizedTokenizer and FunctionTokenDecoder
The InterProQuantizedTokenizer converts a protein’s InterPro IDs and functional keywords into compact token streams using Locally-Sensitive Hashing (L-SH) of TF-IDF vectors. The tokenizer loads safety-filtered lists during initialization at lines 61-73 of esm/tokenization/function_tokenizer.py.
The FunctionTokenDecoder complements this by reading the InterPro safety list at model initialization (lines 58-66 of esm/models/function_decoder.py) to predict the presence of high-risk annotations during server-side inference.
The Detection Pipeline
The guardrail executes automatically through the following request lifecycle:
-
Client Submission: The user submits a protein via
client.generate()orclient.encode(), creating anESMProteinorESMProteinTensorobject. -
Flag Initialization: Every request object includes the boolean field
potential_sequence_of_concern(defaultFalse), defined inesm/sdk/api.pyat lines 50-55. -
Tokenization: The
InterProQuantizedTokenizerencodes functional annotations by readingC.INTERPRO_ENTRYandC.KEYWORDS_VOCABULARYfrom the safety data paths. -
Server-Side Screening: The cloud "forge" service parses the tokenized representation and checks against the curated InterPro list and keyword vocabulary.
-
Flag and Response: If any high-risk InterPro ID or keyword matches, the service sets
potential_sequence_of_concerntoTrueand rejects or warns on the request. The response handling inesm/sdk/forge.py(lines 406-417) extracts this flag from the returned payload.
Code Example: Detecting Controlled Sequences
The following Python implementation demonstrates how the SDK automatically surfaces safety violations:
from esm.sdk import ForgeClient, ProteinInput, GenerationConfig
# Initialize the client (public endpoint)
client = ForgeClient()
# Example sequence mapping to a controlled InterPro entry (e.g., IPR018103)
seq = "MKTIIALSYIFCLVFADYKDDDDK"
# Build request; SDK automatically runs the tokenizer
protein = ProteinInput(id="test", sequence=seq)
# Request embedding generation
config = GenerationConfig(num_steps=10)
result = client.generate([protein], config=config)
# Check safety flag returned from server
if result.potential_sequence_of_concern:
print("⚠️ Input matches controlled pathogen entry — request blocked.")
else:
print("✅ Input passed safety filters.")
Under the hood, client.generate() reads esm/data/entry_list_safety_29026.list via the tokenizer, encodes the sequence using L-SH hashed TF-IDF vectors, and receives the boolean safety determination from the server.
Key Implementation Files
esm/utils/constants/esm3.py— Declares paths to safety-filtered InterPro lists and keyword vocabularies.esm/tokenization/function_tokenizer.py— Loads safety data and tokenizes functional annotations (lines 61-73).esm/models/function_decoder.py— Reads InterPro safety lists at model initialization (lines 58-66).esm/sdk/api.py— Defines thepotential_sequence_of_concernfield on request objects (lines 50-55).esm/sdk/forge.py— Handles server communication and extracts safety flags from responses (lines 406-417).esm/data/entry_list_safety_29026.list— Curated TSV of 29,026 high-risk InterPro entries.esm/data/keyword_vocabulary_safety_filtered_58641.txt— Text file containing 58,641 regulated keywords.esm/data/keyword_idf_safety_filtered_58641.npy— NumPy array of TF-IDF IDF weights for keyword scoring.
Summary
- Biohub Platform safety filters rely on curated static lists of 29,026 InterPro entries and 58,641 keywords, not runtime heuristics.
- The
InterProQuantizedTokenizerautomatically loads safety data fromesm/utils/constants/esm3.pyduring request initialization. - Every API request carries a
potential_sequence_of_concernfield that defaults toFalseand flips toTrueupon detection of regulated sequences. - Detection occurs through L-SH hashed TF-IDF vector comparison against pre-computed safety databases.
- Flagged requests are blocked or warned at the server level before biological inference execution.
Frequently Asked Questions
What triggers the potential_sequence_of_concern flag?
The flag triggers when a submitted protein sequence contains InterPro IDs listed in entry_list_safety_29026.list or functional keywords present in keyword_vocabulary_safety_filtered_58641.txt. The system checks both structural domain annotations and textual functional descriptions using TF-IDF weighted matching.
Where does the Biohub Platform store its curated safety lists?
Static safety lists reside in the esm/data/ directory as TSV and text files, while their paths are referenced through constants in esm/utils/constants/esm3.py. The InterPro list contains 29,026 entries, and the keyword vocabulary contains 58,641 terms with corresponding IDF values in NumPy format.
Can developers bypass the safety filters when using the Biohub SDK?
No. The safety check executes server-side within the cloud "forge" service after the client submits tokenized representations. While the SDK initializes the potential_sequence_of_concern field as False locally, the server overwrites this value based on its independent screening against curated safety databases, making client-side bypass impossible.
How does the TF-IDF weighting improve pathogen detection accuracy?
The platform uses pre-computed IDF values stored in keyword_idf_safety_filtered_58641.npy to weight keywords by their specificity within the biological literature. Rare terms associated with toxins or viral functions receive higher weights, reducing false positives from common protein descriptors while maintaining sensitivity to genuine controlled pathogen signatures.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →