How the Bionty Plugin Integrates Biological Ontologies into LaminDB
The Bionty plugin acts as a registry layer that imports public biological ontologies into LaminDB, providing standardized validation and annotation methods for genes, cell types, tissues, and diseases through a unified Python API.
LaminDB does not store biological vocabularies internally; instead, it relies on the Bionty plugin to provide curated ontologies for scientific data management. According to the K-Dense-AI/scientific-agent-skills repository, this integration allows researchers to validate, standardize, and link metadata from authoritative sources like Ensembl, UniProt, and the Cell Ontology directly to LaminDB artifacts.
Architecture of the Bionty Plugin Integration
Installation and Module Registration
The Bionty plugin is bundled with LaminDB and can be initialized during instance creation. As documented in scientific-skills/lamindb/references/setup-deployment.md, you install the package and register the module using either the CLI or Python API.
# pip install lamindb # pulls in bionty automatically
import lamindb as ln
import bionty as bt
# Initialize with Bionty module enabled
ln.init(storage="./mydata", modules=["bionty"])
When initialized, Bionty registers a collection of registries—including CellType, Gene, Tissue, and Disease—that expose a uniform API for ontology management. These registries are listed in scientific-skills/lamindb/SKILL.md as core components for biological data workflows.
Registry Structure
Each registry in Bionty corresponds to a specific biological domain. The plugin structures these as first-class Python objects that mirror the underlying public ontologies, enabling type-safe operations across genes, proteins, cell types, and anatomical structures.
Importing and Standardizing Biological Ontologies
Importing Public Ontology Sources
Bionty pulls public ontology files from authoritative databases into a local cache via the import_source() method. This step, detailed in scientific-skills/lamindb/references/ontologies.md, downloads controlled vocabularies and makes them available for local validation.
# Run once per project to populate local registries
bt.CellType.import_source()
bt.Gene.import_source(organism="human")
bt.Tissue.import_source()
bt.Disease.import_source(source="mondo") # Monarch Disease Ontology
The method accepts parameters like organism and source to specify which specific ontology build to download (e.g., Ensembl for genes, CL for cell types, or MONDO for diseases).
Validation and Standardization Workflow
Once imported, Bionty provides three core methods for data cleaning, as implemented in scientific-skills/lamindb/references/ontologies.md#standardizing-and-validating-data:
validate: Returns a boolean array indicating which terms exist in the ontologystandardize: Maps synonyms and variations to canonical ontology IDsfrom_values: Converts validated strings into ontology record objects
# Validate free-text annotations against the ontology
cell_terms = ["T cell", "fat cell", "invalid_term"]
valid = bt.CellType.validate(cell_terms)
# Returns: [True, True, False]
# Standardize to canonical identifiers
standardised = bt.CellType.standardize(cell_terms)
# Returns: ['t_cell', 'adipocyte', 'invalid_term']
# Convert valid terms to ontology records
cell_records = bt.CellType.from_values(
[t for t, v in zip(cell_terms, valid) if v]
)
Annotating LaminDB Artifacts with Ontology Metadata
After standardization, ontology records attach to LaminDB artifacts via the feature_sets.add_ontology() method. This linkage ensures that datasets carry validated, queryable metadata rather than free-text annotations, as described in scientific-skills/lamindb/references/ontologies.md#annotating-datasets.
import pandas as pd
# Create sample data
df = pd.DataFrame({
"cell_type": ["T cell", "B cell", "NK cell"],
"tissue": ["blood", "spleen", "liver"]
})
# Standardize columns using Bionty
df["cell_type"] = bt.CellType.standardize(df["cell_type"])
df["tissue"] = bt.Tissue.standardize(df["tissue"])
# Create artifact
artifact = ln.Artifact.from_dataframe(
df,
key="metadata/sample_info.parquet",
description="Curated sample metadata with ontology validation"
).save()
# Link ontology records to the artifact
artifact.feature_sets.add_ontology(
bt.CellType.from_values(df["cell_type"])
)
artifact.feature_sets.add_ontology(
bt.Tissue.from_values(df["tissue"])
)
Querying Ontology-Annotated Data
Because ontology records become first-class feature sets, downstream queries can leverage biological hierarchies without manual joins. The scientific-skills/lamindb/references/ontologies.md#querying-ontology-annotated-data section documents how to traverse parent-child relationships and filter artifacts by ontological descendants.
# Retrieve T cell record and query its subtypes
t_cell = bt.CellType.get(name="T cell")
t_subtypes = t_cell.query_children() # Includes all descendants
# Filter artifacts containing any T-cell subtype
results = ln.Artifact.filter(
feature_sets__cell_types__in=t_subtypes
).to_dataframe()
Summary
- Bionty serves as the ontology layer: LaminDB delegates biological vocabulary management to the Bionty plugin, which harvests from public sources like Ensembl and Cell Ontology.
- Standardization ensures data quality: The
validate,standardize, andfrom_valuesmethods map free-text terms to canonical IDs before storage. - Artifacts carry linked metadata: Ontology records attach to datasets via
feature_sets.add_ontology(), creating searchable, hierarchical annotations. - Hierarchical queries are native: Because ontologies are first-class registries, you can filter artifacts by parent classes (e.g., all T-cell subtypes) using
query_children()andparentsrelationships.
Frequently Asked Questions
What biological ontologies does the Bionty plugin support?
Bionty supports major biomedical ontologies including Cell Ontology (CL) for cell types, Ensembl for genes, UniProt for proteins, Uberon for tissues, and MONDO for diseases. The specific sources are configurable via the source parameter in import_source() methods, as documented in scientific-skills/lamindb/references/ontologies.md.
How does Bionty handle ontology versioning?
When calling import_source(), Bionty downloads specific builds of public ontologies (e.g., Ensembl Release 109) and caches them locally. You can specify versions using parameters like organism="human" or source="mondo" to ensure reproducibility across computational workflows.
Can I use custom ontologies with the Bionty plugin?
The Bionty plugin is designed around standardized public ontologies, but you can extend registries by adding custom records to the local cache after initial import. However, custom vocabularies without public identifiers lose the benefit of automated hierarchical querying through query_children() and parents relationships.
What is the difference between validate and standardize in Bionty?
validate performs a strict lookup that returns boolean flags indicating exact matches in the ontology. standardize attempts to correct synonyms and common variations (e.g., mapping "fat cell" to "adipocyte") before returning the canonical identifier. Use validate to check data quality, and standardize to clean messy input data before creating ontology records.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →