How to Integrate LaminDB with Nextflow and Snakemake Workflows
Integrate LaminDB into Nextflow and Snakemake pipelines by wrapping your analysis logic with ln.track() at the start and ln.finish() at the end, using ln.Artifact.get() to load inputs and ln.Artifact(...).save() to persist outputs with automatic lineage tracking.
The lamindb library provides a lightweight data-management layer designed specifically for computational biology pipelines. According to the K-Dense-AI/scientific-agent-skills repository, you can integrate LaminDB with Nextflow and Snakemake workflows by embedding Python tracking calls directly inside process scripts and rules. This approach captures complete provenance—including input dependencies, output artifacts, and environment parameters—without modifying your underlying workflow engine configuration.
The Five-Step Integration Pattern
Every pipeline step follows a consistent pattern documented in scientific-skills/lamindb/references/integrations.md. Implement these five stages to ensure full data lineage and provenance tracking:
- Initialize tracking: Call
ln.track()to start a new LaminDB run, record environment details, and prepare a cache for artifacts. - Load inputs: Retrieve previously stored artifacts using
ln.Artifact.get(key=...)and access their data with the.load()method. - Process data: Execute your analysis logic (pandas transformations, statistical modeling, etc.).
- Persist outputs: Save results back to LaminDB using
ln.Artifact.from_dataframe(...).save()orln.Artifact(path, key=...).save(). - Finalize: Close the run with
ln.finish()to write provenance information including git hashes, timestamps, and parameters.
Because LaminDB stores metadata and lineage (input → output), downstream steps can query artifacts using their unique keys without re-computing intermediate results.
Nextflow Integration
Embedding LaminDB in a Nextflow Process
In Nextflow, embed LaminDB calls within the script: block of a process definition. The Python code executes inside the process environment, requiring only that lamindb is installed in the container or conda environment.
process ANALYZE {
input:
val input_key
output:
path "result.csv"
script:
"""
#!/usr/bin/env python
import lamindb as ln
# Start tracking for this process
ln.track()
# Load the input artifact from LaminDB
artifact = ln.Artifact.get(key="${input_key}")
data = artifact.load()
# ---- Your analysis logic ----
# Example: simple pandas transformation
import pandas as pd
df = pd.read_csv(data)
result = df.groupby("category").sum().reset_index()
result.to_csv("result.csv", index=False)
# Register the output artifact
ln.Artifact("result.csv", key="outputs/result.csv").save()
# Finish tracking – writes provenance to the DB
ln.finish()
"""
}
The ln.track() call creates a LaminDB run entry that records the current Git commit, environment variables, and any parameters passed via ln.track(params={...}). The ln.Artifact.get() method fetches the input using the key supplied by the Nextflow workflow (${input_key}). After processing, ln.finish() finalizes the run and links the output artifact to the run record.
This implementation mirrors the reference code found at lines 99–124 in scientific-skills/lamindb/references/integrations.md.
Snakemake Integration
Adding Tracking to a Snakemake Rule
Snakemake supports arbitrary Python code via the run: directive, making LaminDB integration straightforward. Wrap your analysis logic with the same tracking calls used in Nextflow.
rule process_data:
input:
"data/input.csv"
output:
"data/output.csv"
run:
import lamindb as ln
import pandas as pd
# Start a LaminDB run for this rule
ln.track()
# Load the input artifact registered previously
artifact = ln.Artifact.get(key="inputs/data.csv")
df = artifact.load()
# ---- Your analysis logic ----
result = df.assign(mean=df.value.mean())
result.to_csv(output[0], index=False)
# Save the output artifact back to LaminDB
ln.Artifact(output[0], key="outputs/result.csv").save()
# End the LaminDB run
ln.finish()
The artifact key (e.g., "inputs/data.csv" and "outputs/result.csv") acts as a global identifier that other rules can reference with ln.Artifact.get(). Because the run is recorded, you can later query the lineage using ln.ViewLineage() to see which rule produced which artifact.
This example follows the official Snakemake integration recipe at lines 146–174 in scientific-skills/lamindb/references/integrations.md.
Key Source Files and References
The following files in the K-Dense-AI/scientific-agent-skills repository provide complete implementation details:
scientific-skills/lamindb/references/integrations.md— Contains the full reference for Nextflow and Snakemake integration, including the code examples above.scientific-skills/lamindb/SKILL.md— Provides high-level LaminDB installation instructions (pip install lamindb) and example imports.scientific-skills/lamindb/references/setup-deployment.md— Documents installation commands and cloud storage extras (e.g.,lamindb[s3],lamindb[gcp]).scientific-skills/latchbio-integration/references/workflow-creation.md— Shows how to register Nextflow and Snakemake workflows with LatchBio usinglatch register --nextflowandlatch register --snakemake.
Summary
- Wrap every workflow step with
ln.track()andln.finish()to capture complete provenance, including git hashes and timestamps. - Use
ln.Artifact.get(key=...)to load inputs andln.Artifact(...).save()to register outputs with automatic versioning. - Both Nextflow
processscripts and Snakemakerunblocks support inline Python execution for LaminDB calls. - Artifact keys serve as global identifiers that enable cross-step data lineage tracking via
ln.ViewLineage(). - Implementation details are maintained in
scientific-skills/lamindb/references/integrations.mdwithin the K-Dense-AI/scientific-agent-skills repository.
Frequently Asked Questions
Can I integrate LaminDB with workflow managers other than Nextflow and Snakemake?
Yes. While the K-Dense-AI/scientific-agent-skills repository specifically documents Nextflow and Snakemake, LaminDB’s Python API can be embedded in any workflow manager that supports Python script execution, including Apache Airflow, Prefect, or shell-based pipelines that invoke Python scripts.
How does LaminDB track lineage between workflow steps?
LaminDB automatically links input and output artifacts to the current run record when you use ln.Artifact.get() and ln.Artifact(...).save(). You can query this lineage programmatically using ln.ViewLineage() to visualize the complete graph of which pipeline steps produced specific artifacts and what parameters were used.
What Python environment setup is required to use LaminDB in a pipeline?
The execution environment must have lamindb installed, typically via pip install lamindb as documented in scientific-skills/lamindb/references/setup-deployment.md. If your pipeline uses cloud storage backends, install the appropriate extras such as lamindb[s3] for AWS or lamindb[gcp] for Google Cloud Platform.
Do I need to pre-register input artifacts before running the workflow?
No. While you query existing artifacts with ln.Artifact.get(key=...), new outputs are registered automatically during workflow execution via the .save() method. The system handles versioning, caching, and metadata extraction without requiring manual pre-registration of files.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →