# How LaminDB Artifact Versioning and Lineage Tracking Works: A Technical Guide

> Discover how LaminDB artifact versioning and lineage tracking ensures full reproducibility. Learn how it links datasets, code, and environments for complete tracking.

- Repository: [K-Dense/scientific-agent-skills](https://github.com/K-Dense-AI/scientific-agent-skills)
- Tags: how-to-guide
- Published: 2026-05-14

---

**LaminDB automatically creates immutable artifact versions and constructs provenance graphs linking every output dataset to its specific input data, transformation code, and execution environment, enabling complete reproducibility tracking in SQLite or PostgreSQL backends.**

The K-Dense-AI/scientific-agent-skills repository provides the definitive reference implementation for **LaminDB artifact versioning and lineage tracking**, an open-source data framework engineered for biological research workflows. According to the architectural overview in [`scientific-skills/lamindb/SKILL.md`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/scientific-skills/lamindb/SKILL.md) and the detailed specifications in [`scientific-skills/lamindb/references/core-concepts.md`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/scientific-skills/lamindb/references/core-concepts.md), LaminDB structures data management around four primary entities—**Artifacts**, **Runs**, **Transforms**, and **Features**—that collectively enforce FAIR principles (Findable, Accessible, Interoperable, Reusable) through automated provenance capture.

## Core Concepts: Artifacts, Runs, and Transforms

LaminDB organizes data management around four fundamental building blocks. **Artifacts** represent any dataset format—including DataFrames, AnnData objects, SpatialData, Parquet, or Zarr stores—and serve as immutable units of data storage. Each artifact links to a **Run**, which captures the specific execution context including timestamps, environment variables, and Python package versions. **Transforms** define the computational operations applied to data, while **Features** annotate the biological and experimental metadata.

When you invoke `ldb.save()` to persist data, LaminDB registers a new entry in the underlying database table, automatically assigning a version identifier and linking the artifact to the current execution context. This design ensures that every dataset exists as a permanent, referenceable object within the provenance graph rather than a mutable file path.

## Automatic Versioning Mechanics

LaminDB implements **automatic versioning** that triggers whenever source data or transformation logic changes. If you modify the input DataFrame or alter the Python code within a transformation block, the framework detects the change through content hashing and creates a new artifact version rather than overwriting the existing record. Previous versions remain immutable, preserving the exact state of data at specific analytical milestones.

This immutability guarantee means that `raw_counts.v1` will always point to the identical bytes and metadata as when first created, even after `raw_counts.v2` exists. The versioning logic specified in [`scientific-skills/lamindb/references/core-concepts.md`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/scientific-skills/lamindb/references/core-concepts.md) dictates that version increments occur automatically upon detecting hash mismatches in input data or code signature changes, eliminating manual version management.

## Provenance Graph and Lineage Tracking

The **lineage tracking** system stores relationships as a directed acyclic graph (DAG) in SQLite or PostgreSQL, mapping how data flows through your pipeline. Each execution of a transform creates a **Run** node that connects parent input artifacts to child output artifacts. This provenance graph enables bidirectional traversal: you can query upstream to identify which raw data and code version produced a specific result, or query downstream to discover all analyses dependent on a particular source dataset.

In [`scientific-skills/lamindb/references/core-concepts.md`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/scientific-skills/lamindb/references/core-concepts.md), the lineage system is described as a complete "data-science DAG" where edges represent transformations and nodes represent artifact versions. The graph persists in the underlying relational store, allowing you to audit pipelines using standard SQL queries without relying on external metadata files.

## Capturing Execution Context

Beyond data relationships, LaminDB captures comprehensive **execution context** to ensure reproducibility. When you initialize a run using `ldb.run()`, the framework automatically records the Python version, installed package versions, environment variables, and code hashes. This environmental snapshot attaches to every artifact created within the run block, enabling you to recreate the exact computational environment months later.

The execution context storage integrates with workflow managers like Snakemake, Nextflow, and Redun, as documented in [`scientific-skills/lamindb/references/integrations.md`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/scientific-skills/lamindb/references/integrations.md). This integration ensures that pipeline orchestration metadata augments the native LaminDB provenance capture, creating a unified audit trail across distributed computing environments.

## Practical Implementation: Versioning and Lineage Workflow

The following implementation demonstrates how to register datasets, execute transformations, and query lineage using the LaminDB Python API. These examples assume installation via `pip install lamin-db[bionty]`.

### Initializing the Database Connection

First, establish a connection to your LaminDB instance. By default, this creates a local SQLite database file named `lamindb.sqlite`.

```python
from lamin import db

# Connect to database - creates lamindb.sqlite if it doesn't exist

ldb = db.connect()

```

### Registering Raw Data Artifacts

Convert a pandas DataFrame into a tracked artifact with automatic initial versioning.

```python
import pandas as pd

# Create sample biological data

df_raw = pd.DataFrame({"gene": ["A", "B"], "count": [10, 20]})

# Save as version 1 artifact

raw_art = ldb.save(df_raw, name="raw_counts", description="Raw RNA-seq counts")
print(f"Artifact ID: {raw_art.id}, version: {raw_art.version}")

```

### Executing Transformations with Run Context

Perform data processing inside a run context to capture provenance automatically.

```python
import numpy as np

# Initialize run context to track execution

with ldb.run(name="log_norm", description="Log-normalise counts") as run:
    # Load input artifact within run context

    input_art = ldb.get_artifact(raw_art.id)
    df = input_art.load()
    
    # Apply transformation

    df["log_count"] = (df["count"] + 1).apply(np.log)
    
    # Save creates version 2 with automatic parent linkage

    norm_art = ldb.save(df, name="norm_counts", run=run)

print(f"New artifact version: {norm_art.version}")

```

### Querying Lineage and Dependencies

Inspect the provenance graph and discover relationships between artifacts.

```python

# Retrieve complete lineage for an artifact

lineage = ldb.get_lineage(norm_art.id)
print("Lineage graph:")
for node in lineage.nodes():
    print(f"  {node.type}: {node.id} – {node.name}")

# Discover all downstream artifacts depending on raw_counts

downstream = ldb.query_downstream(raw_art.id)
print("Downstream artifacts:")
for artifact in downstream:
    print(f"- {artifact.name} (v{artifact.version})")

```

## Biological Ontologies and Workflow Integrations

LaminDB extends its versioning capabilities through **Bionty**, a plugin for biological ontology management documented in [`scientific-skills/lamindb/references/ontologies.md`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/scientific-skills/lamindb/references/ontologies.md). This integration allows you to attach standardized biological annotations to artifacts while maintaining version control over both data and metadata, ensuring that gene identifiers, cell types, and experimental conditions remain semantically consistent across versions.

The framework also provides hooks for workflow managers and MLOps platforms. As detailed in [`scientific-skills/lamindb/references/integrations.md`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/scientific-skills/lamindb/references/integrations.md), LaminDB synchronizes with Snakemake pipelines, Nextflow workflows, and experiment tracking systems like Weights & Biases and MLflow. These integrations ensure that external pipeline executions automatically generate proper LaminDB **Run** contexts and artifact versions, eliminating the gap between pipeline orchestration and data provenance.

## Summary

- **Artifacts** serve as immutable dataset representations with automatic version increments stored in SQLite or PostgreSQL backends.
- **Runs** capture complete execution context including code versions, Python environments, and timestamps when using `ldb.run()`.
- **Lineage tracking** creates a queryable provenance graph linking parent and child artifacts through transformation operations.
- **Bidirectional queries** via `ldb.get_lineage()` and `ldb.query_downstream()` enable full audit trails and impact analysis.
- **Bionty integration** provides ontology-aware metadata management while maintaining strict version control over biological annotations.

## Frequently Asked Questions

### How does LaminDB detect when to create a new artifact version?

LaminDB monitors the hash of input data and the signature of transformation code during `ldb.save()` operations. When either the source data changes or the Python code within the run context modifies, the framework automatically increments the version number and creates a new immutable record while preserving the previous version, as implemented in the artifact management logic of [`scientific-skills/lamindb/references/core-concepts.md`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/scientific-skills/lamindb/references/core-concepts.md).

### Can I query lineage backward to find which code produced a specific dataset?

Yes. Using `ldb.get_lineage()` with an artifact ID traverses the provenance graph backward to identify the specific **Run** object and **Transform** that generated the dataset, including the exact code version, environment variables, and parent input artifacts stored in the database.

### What database backends support LaminDB versioning and lineage tracking?

LaminDB supports both SQLite for local development and PostgreSQL for production deployments. The provenance graph and artifact metadata persist in standard relational tables within these databases, enabling SQL-based queries of lineage relationships without requiring additional graph database infrastructure.

### How does LaminDB integrate with existing workflow managers like Snakemake or Nextflow?

According to [`scientific-skills/lamindb/references/integrations.md`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/scientific-skills/lamindb/references/integrations.md), LaminDB provides hooks that automatically create **Run** contexts and register output artifacts when integrated with Snakemake, Nextflow, or Redun. For MLOps platforms like Weights & Biases and MLflow, LaminDB synchronizes experiment tracking metadata with its native artifact versioning system, ensuring consistency between pipeline orchestration and data provenance.