# How the Bionty Plugin Integrates Biological Ontologies into LaminDB

> Learn how the Bionty plugin integrates biological ontologies into LaminDB for standardized gene, cell type, tissue, and disease validation and annotation via a unified Python API.

- Repository: [K-Dense/scientific-agent-skills](https://github.com/K-Dense-AI/scientific-agent-skills)
- Tags: how-to-guide
- Published: 2026-05-14

---

**The Bionty plugin acts as a registry layer that imports public biological ontologies into LaminDB, providing standardized validation and annotation methods for genes, cell types, tissues, and diseases through a unified Python API.**

LaminDB does not store biological vocabularies internally; instead, it relies on the **Bionty** plugin to provide curated ontologies for scientific data management. According to the K-Dense-AI/scientific-agent-skills repository, this integration allows researchers to validate, standardize, and link metadata from authoritative sources like Ensembl, UniProt, and the Cell Ontology directly to LaminDB artifacts.

## Architecture of the Bionty Plugin Integration

### Installation and Module Registration

The Bionty plugin is bundled with LaminDB and can be initialized during instance creation. As documented in [`scientific-skills/lamindb/references/setup-deployment.md`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/scientific-skills/lamindb/references/setup-deployment.md), you install the package and register the module using either the CLI or Python API.

```python

# pip install lamindb  # pulls in bionty automatically

import lamindb as ln
import bionty as bt

# Initialize with Bionty module enabled

ln.init(storage="./mydata", modules=["bionty"])

```

When initialized, Bionty registers a collection of **registries**—including `CellType`, `Gene`, `Tissue`, and `Disease`—that expose a uniform API for ontology management. These registries are listed in [`scientific-skills/lamindb/SKILL.md`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/scientific-skills/lamindb/SKILL.md) as core components for biological data workflows.

### Registry Structure

Each registry in Bionty corresponds to a specific biological domain. The plugin structures these as first-class Python objects that mirror the underlying public ontologies, enabling type-safe operations across genes, proteins, cell types, and anatomical structures.

## Importing and Standardizing Biological Ontologies

### Importing Public Ontology Sources

Bionty pulls public ontology files from authoritative databases into a local cache via the `import_source()` method. This step, detailed in [`scientific-skills/lamindb/references/ontologies.md`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/scientific-skills/lamindb/references/ontologies.md), downloads controlled vocabularies and makes them available for local validation.

```python

# Run once per project to populate local registries

bt.CellType.import_source()
bt.Gene.import_source(organism="human")
bt.Tissue.import_source()
bt.Disease.import_source(source="mondo")  # Monarch Disease Ontology

```

The method accepts parameters like `organism` and `source` to specify which specific ontology build to download (e.g., Ensembl for genes, CL for cell types, or MONDO for diseases).

### Validation and Standardization Workflow

Once imported, Bionty provides three core methods for data cleaning, as implemented in `scientific-skills/lamindb/references/ontologies.md#standardizing-and-validating-data`:

- **`validate`**: Returns a boolean array indicating which terms exist in the ontology
- **`standardize`**: Maps synonyms and variations to canonical ontology IDs
- **`from_values`**: Converts validated strings into ontology record objects

```python

# Validate free-text annotations against the ontology

cell_terms = ["T cell", "fat cell", "invalid_term"]
valid = bt.CellType.validate(cell_terms)

# Returns: [True, True, False]

# Standardize to canonical identifiers

standardised = bt.CellType.standardize(cell_terms)

# Returns: ['t_cell', 'adipocyte', 'invalid_term']

# Convert valid terms to ontology records

cell_records = bt.CellType.from_values(
    [t for t, v in zip(cell_terms, valid) if v]
)

```

## Annotating LaminDB Artifacts with Ontology Metadata

After standardization, ontology records attach to **LaminDB artifacts** via the `feature_sets.add_ontology()` method. This linkage ensures that datasets carry validated, queryable metadata rather than free-text annotations, as described in `scientific-skills/lamindb/references/ontologies.md#annotating-datasets`.

```python
import pandas as pd

# Create sample data

df = pd.DataFrame({
    "cell_type": ["T cell", "B cell", "NK cell"],
    "tissue": ["blood", "spleen", "liver"]
})

# Standardize columns using Bionty

df["cell_type"] = bt.CellType.standardize(df["cell_type"])
df["tissue"] = bt.Tissue.standardize(df["tissue"])

# Create artifact

artifact = ln.Artifact.from_dataframe(
    df,
    key="metadata/sample_info.parquet",
    description="Curated sample metadata with ontology validation"
).save()

# Link ontology records to the artifact

artifact.feature_sets.add_ontology(
    bt.CellType.from_values(df["cell_type"])
)
artifact.feature_sets.add_ontology(
    bt.Tissue.from_values(df["tissue"])
)

```

## Querying Ontology-Annotated Data

Because ontology records become first-class feature sets, downstream queries can leverage biological hierarchies without manual joins. The `scientific-skills/lamindb/references/ontologies.md#querying-ontology-annotated-data` section documents how to traverse parent-child relationships and filter artifacts by ontological descendants.

```python

# Retrieve T cell record and query its subtypes

t_cell = bt.CellType.get(name="T cell")
t_subtypes = t_cell.query_children()  # Includes all descendants

# Filter artifacts containing any T-cell subtype

results = ln.Artifact.filter(
    feature_sets__cell_types__in=t_subtypes
).to_dataframe()

```

## Summary

- **Bionty serves as the ontology layer**: LaminDB delegates biological vocabulary management to the Bionty plugin, which harvests from public sources like Ensembl and Cell Ontology.
- **Standardization ensures data quality**: The `validate`, `standardize`, and `from_values` methods map free-text terms to canonical IDs before storage.
- **Artifacts carry linked metadata**: Ontology records attach to datasets via `feature_sets.add_ontology()`, creating searchable, hierarchical annotations.
- **Hierarchical queries are native**: Because ontologies are first-class registries, you can filter artifacts by parent classes (e.g., all T-cell subtypes) using `query_children()` and `parents` relationships.

## Frequently Asked Questions

### What biological ontologies does the Bionty plugin support?

Bionty supports major biomedical ontologies including **Cell Ontology (CL)** for cell types, **Ensembl** for genes, **UniProt** for proteins, **Uberon** for tissues, and **MONDO** for diseases. The specific sources are configurable via the `source` parameter in `import_source()` methods, as documented in [`scientific-skills/lamindb/references/ontologies.md`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/scientific-skills/lamindb/references/ontologies.md).

### How does Bionty handle ontology versioning?

When calling `import_source()`, Bionty downloads specific builds of public ontologies (e.g., Ensembl Release 109) and caches them locally. You can specify versions using parameters like `organism="human"` or `source="mondo"` to ensure reproducibility across computational workflows.

### Can I use custom ontologies with the Bionty plugin?

The Bionty plugin is designed around standardized public ontologies, but you can extend registries by adding custom records to the local cache after initial import. However, custom vocabularies without public identifiers lose the benefit of automated hierarchical querying through `query_children()` and `parents` relationships.

### What is the difference between `validate` and `standardize` in Bionty?

**`validate`** performs a strict lookup that returns boolean flags indicating exact matches in the ontology. **`standardize`** attempts to correct synonyms and common variations (e.g., mapping "fat cell" to "adipocyte") before returning the canonical identifier. Use `validate` to check data quality, and `standardize` to clean messy input data before creating ontology records.