How LaminDB Artifact Versioning and Lineage Tracking Works: A Technical Guide
LaminDB automatically creates immutable artifact versions and constructs provenance graphs linking every output dataset to its specific input data, transformation code, and execution environment, enabling complete reproducibility tracking in SQLite or PostgreSQL backends.
The K-Dense-AI/scientific-agent-skills repository provides the definitive reference implementation for LaminDB artifact versioning and lineage tracking, an open-source data framework engineered for biological research workflows. According to the architectural overview in scientific-skills/lamindb/SKILL.md and the detailed specifications in scientific-skills/lamindb/references/core-concepts.md, LaminDB structures data management around four primary entities—Artifacts, Runs, Transforms, and Features—that collectively enforce FAIR principles (Findable, Accessible, Interoperable, Reusable) through automated provenance capture.
Core Concepts: Artifacts, Runs, and Transforms
LaminDB organizes data management around four fundamental building blocks. Artifacts represent any dataset format—including DataFrames, AnnData objects, SpatialData, Parquet, or Zarr stores—and serve as immutable units of data storage. Each artifact links to a Run, which captures the specific execution context including timestamps, environment variables, and Python package versions. Transforms define the computational operations applied to data, while Features annotate the biological and experimental metadata.
When you invoke ldb.save() to persist data, LaminDB registers a new entry in the underlying database table, automatically assigning a version identifier and linking the artifact to the current execution context. This design ensures that every dataset exists as a permanent, referenceable object within the provenance graph rather than a mutable file path.
Automatic Versioning Mechanics
LaminDB implements automatic versioning that triggers whenever source data or transformation logic changes. If you modify the input DataFrame or alter the Python code within a transformation block, the framework detects the change through content hashing and creates a new artifact version rather than overwriting the existing record. Previous versions remain immutable, preserving the exact state of data at specific analytical milestones.
This immutability guarantee means that raw_counts.v1 will always point to the identical bytes and metadata as when first created, even after raw_counts.v2 exists. The versioning logic specified in scientific-skills/lamindb/references/core-concepts.md dictates that version increments occur automatically upon detecting hash mismatches in input data or code signature changes, eliminating manual version management.
Provenance Graph and Lineage Tracking
The lineage tracking system stores relationships as a directed acyclic graph (DAG) in SQLite or PostgreSQL, mapping how data flows through your pipeline. Each execution of a transform creates a Run node that connects parent input artifacts to child output artifacts. This provenance graph enables bidirectional traversal: you can query upstream to identify which raw data and code version produced a specific result, or query downstream to discover all analyses dependent on a particular source dataset.
In scientific-skills/lamindb/references/core-concepts.md, the lineage system is described as a complete "data-science DAG" where edges represent transformations and nodes represent artifact versions. The graph persists in the underlying relational store, allowing you to audit pipelines using standard SQL queries without relying on external metadata files.
Capturing Execution Context
Beyond data relationships, LaminDB captures comprehensive execution context to ensure reproducibility. When you initialize a run using ldb.run(), the framework automatically records the Python version, installed package versions, environment variables, and code hashes. This environmental snapshot attaches to every artifact created within the run block, enabling you to recreate the exact computational environment months later.
The execution context storage integrates with workflow managers like Snakemake, Nextflow, and Redun, as documented in scientific-skills/lamindb/references/integrations.md. This integration ensures that pipeline orchestration metadata augments the native LaminDB provenance capture, creating a unified audit trail across distributed computing environments.
Practical Implementation: Versioning and Lineage Workflow
The following implementation demonstrates how to register datasets, execute transformations, and query lineage using the LaminDB Python API. These examples assume installation via pip install lamin-db[bionty].
Initializing the Database Connection
First, establish a connection to your LaminDB instance. By default, this creates a local SQLite database file named lamindb.sqlite.
from lamin import db
# Connect to database - creates lamindb.sqlite if it doesn't exist
ldb = db.connect()
Registering Raw Data Artifacts
Convert a pandas DataFrame into a tracked artifact with automatic initial versioning.
import pandas as pd
# Create sample biological data
df_raw = pd.DataFrame({"gene": ["A", "B"], "count": [10, 20]})
# Save as version 1 artifact
raw_art = ldb.save(df_raw, name="raw_counts", description="Raw RNA-seq counts")
print(f"Artifact ID: {raw_art.id}, version: {raw_art.version}")
Executing Transformations with Run Context
Perform data processing inside a run context to capture provenance automatically.
import numpy as np
# Initialize run context to track execution
with ldb.run(name="log_norm", description="Log-normalise counts") as run:
# Load input artifact within run context
input_art = ldb.get_artifact(raw_art.id)
df = input_art.load()
# Apply transformation
df["log_count"] = (df["count"] + 1).apply(np.log)
# Save creates version 2 with automatic parent linkage
norm_art = ldb.save(df, name="norm_counts", run=run)
print(f"New artifact version: {norm_art.version}")
Querying Lineage and Dependencies
Inspect the provenance graph and discover relationships between artifacts.
# Retrieve complete lineage for an artifact
lineage = ldb.get_lineage(norm_art.id)
print("Lineage graph:")
for node in lineage.nodes():
print(f" {node.type}: {node.id} – {node.name}")
# Discover all downstream artifacts depending on raw_counts
downstream = ldb.query_downstream(raw_art.id)
print("Downstream artifacts:")
for artifact in downstream:
print(f"- {artifact.name} (v{artifact.version})")
Biological Ontologies and Workflow Integrations
LaminDB extends its versioning capabilities through Bionty, a plugin for biological ontology management documented in scientific-skills/lamindb/references/ontologies.md. This integration allows you to attach standardized biological annotations to artifacts while maintaining version control over both data and metadata, ensuring that gene identifiers, cell types, and experimental conditions remain semantically consistent across versions.
The framework also provides hooks for workflow managers and MLOps platforms. As detailed in scientific-skills/lamindb/references/integrations.md, LaminDB synchronizes with Snakemake pipelines, Nextflow workflows, and experiment tracking systems like Weights & Biases and MLflow. These integrations ensure that external pipeline executions automatically generate proper LaminDB Run contexts and artifact versions, eliminating the gap between pipeline orchestration and data provenance.
Summary
- Artifacts serve as immutable dataset representations with automatic version increments stored in SQLite or PostgreSQL backends.
- Runs capture complete execution context including code versions, Python environments, and timestamps when using
ldb.run(). - Lineage tracking creates a queryable provenance graph linking parent and child artifacts through transformation operations.
- Bidirectional queries via
ldb.get_lineage()andldb.query_downstream()enable full audit trails and impact analysis. - Bionty integration provides ontology-aware metadata management while maintaining strict version control over biological annotations.
Frequently Asked Questions
How does LaminDB detect when to create a new artifact version?
LaminDB monitors the hash of input data and the signature of transformation code during ldb.save() operations. When either the source data changes or the Python code within the run context modifies, the framework automatically increments the version number and creates a new immutable record while preserving the previous version, as implemented in the artifact management logic of scientific-skills/lamindb/references/core-concepts.md.
Can I query lineage backward to find which code produced a specific dataset?
Yes. Using ldb.get_lineage() with an artifact ID traverses the provenance graph backward to identify the specific Run object and Transform that generated the dataset, including the exact code version, environment variables, and parent input artifacts stored in the database.
What database backends support LaminDB versioning and lineage tracking?
LaminDB supports both SQLite for local development and PostgreSQL for production deployments. The provenance graph and artifact metadata persist in standard relational tables within these databases, enabling SQL-based queries of lineage relationships without requiring additional graph database infrastructure.
How does LaminDB integrate with existing workflow managers like Snakemake or Nextflow?
According to scientific-skills/lamindb/references/integrations.md, LaminDB provides hooks that automatically create Run contexts and register output artifacts when integrated with Snakemake, Nextflow, or Redun. For MLOps platforms like Weights & Biases and MLflow, LaminDB synchronizes experiment tracking metadata with its native artifact versioning system, ensuring consistency between pipeline orchestration and data provenance.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →