How to Integrate LaminDB with Nextflow and Snakemake Workflows

Integrate LaminDB into Nextflow and Snakemake pipelines by wrapping your analysis logic with ln.track() at the start and ln.finish() at the end, using ln.Artifact.get() to load inputs and ln.Artifact(...).save() to persist outputs with automatic lineage tracking.

The lamindb library provides a lightweight data-management layer designed specifically for computational biology pipelines. According to the K-Dense-AI/scientific-agent-skills repository, you can integrate LaminDB with Nextflow and Snakemake workflows by embedding Python tracking calls directly inside process scripts and rules. This approach captures complete provenance—including input dependencies, output artifacts, and environment parameters—without modifying your underlying workflow engine configuration.

The Five-Step Integration Pattern

Every pipeline step follows a consistent pattern documented in scientific-skills/lamindb/references/integrations.md. Implement these five stages to ensure full data lineage and provenance tracking:

  1. Initialize tracking: Call ln.track() to start a new LaminDB run, record environment details, and prepare a cache for artifacts.
  2. Load inputs: Retrieve previously stored artifacts using ln.Artifact.get(key=...) and access their data with the .load() method.
  3. Process data: Execute your analysis logic (pandas transformations, statistical modeling, etc.).
  4. Persist outputs: Save results back to LaminDB using ln.Artifact.from_dataframe(...).save() or ln.Artifact(path, key=...).save().
  5. Finalize: Close the run with ln.finish() to write provenance information including git hashes, timestamps, and parameters.

Because LaminDB stores metadata and lineage (input → output), downstream steps can query artifacts using their unique keys without re-computing intermediate results.

Nextflow Integration

Embedding LaminDB in a Nextflow Process

In Nextflow, embed LaminDB calls within the script: block of a process definition. The Python code executes inside the process environment, requiring only that lamindb is installed in the container or conda environment.

process ANALYZE {
    input:
    val input_key
    output:
    path "result.csv"
    script:
    """
    #!/usr/bin/env python
    import lamindb as ln

    # Start tracking for this process

    ln.track()

    # Load the input artifact from LaminDB

    artifact = ln.Artifact.get(key="${input_key}")
    data = artifact.load()

    # ---- Your analysis logic ----

    # Example: simple pandas transformation

    import pandas as pd
    df = pd.read_csv(data)
    result = df.groupby("category").sum().reset_index()
    result.to_csv("result.csv", index=False)

    # Register the output artifact

    ln.Artifact("result.csv", key="outputs/result.csv").save()

    # Finish tracking – writes provenance to the DB

    ln.finish()
    """
}

The ln.track() call creates a LaminDB run entry that records the current Git commit, environment variables, and any parameters passed via ln.track(params={...}). The ln.Artifact.get() method fetches the input using the key supplied by the Nextflow workflow (${input_key}). After processing, ln.finish() finalizes the run and links the output artifact to the run record.

This implementation mirrors the reference code found at lines 99–124 in scientific-skills/lamindb/references/integrations.md.

Snakemake Integration

Adding Tracking to a Snakemake Rule

Snakemake supports arbitrary Python code via the run: directive, making LaminDB integration straightforward. Wrap your analysis logic with the same tracking calls used in Nextflow.

rule process_data:
    input:
        "data/input.csv"
    output:
        "data/output.csv"
    run:
        import lamindb as ln
        import pandas as pd

        # Start a LaminDB run for this rule

        ln.track()

        # Load the input artifact registered previously

        artifact = ln.Artifact.get(key="inputs/data.csv")
        df = artifact.load()

        # ---- Your analysis logic ----

        result = df.assign(mean=df.value.mean())
        result.to_csv(output[0], index=False)

        # Save the output artifact back to LaminDB

        ln.Artifact(output[0], key="outputs/result.csv").save()

        # End the LaminDB run

        ln.finish()

The artifact key (e.g., "inputs/data.csv" and "outputs/result.csv") acts as a global identifier that other rules can reference with ln.Artifact.get(). Because the run is recorded, you can later query the lineage using ln.ViewLineage() to see which rule produced which artifact.

This example follows the official Snakemake integration recipe at lines 146–174 in scientific-skills/lamindb/references/integrations.md.

Key Source Files and References

The following files in the K-Dense-AI/scientific-agent-skills repository provide complete implementation details:

Summary

  • Wrap every workflow step with ln.track() and ln.finish() to capture complete provenance, including git hashes and timestamps.
  • Use ln.Artifact.get(key=...) to load inputs and ln.Artifact(...).save() to register outputs with automatic versioning.
  • Both Nextflow process scripts and Snakemake run blocks support inline Python execution for LaminDB calls.
  • Artifact keys serve as global identifiers that enable cross-step data lineage tracking via ln.ViewLineage().
  • Implementation details are maintained in scientific-skills/lamindb/references/integrations.md within the K-Dense-AI/scientific-agent-skills repository.

Frequently Asked Questions

Can I integrate LaminDB with workflow managers other than Nextflow and Snakemake?

Yes. While the K-Dense-AI/scientific-agent-skills repository specifically documents Nextflow and Snakemake, LaminDB’s Python API can be embedded in any workflow manager that supports Python script execution, including Apache Airflow, Prefect, or shell-based pipelines that invoke Python scripts.

How does LaminDB track lineage between workflow steps?

LaminDB automatically links input and output artifacts to the current run record when you use ln.Artifact.get() and ln.Artifact(...).save(). You can query this lineage programmatically using ln.ViewLineage() to visualize the complete graph of which pipeline steps produced specific artifacts and what parameters were used.

What Python environment setup is required to use LaminDB in a pipeline?

The execution environment must have lamindb installed, typically via pip install lamindb as documented in scientific-skills/lamindb/references/setup-deployment.md. If your pipeline uses cloud storage backends, install the appropriate extras such as lamindb[s3] for AWS or lamindb[gcp] for Google Cloud Platform.

Do I need to pre-register input artifacts before running the workflow?

No. While you query existing artifacts with ln.Artifact.get(key=...), new outputs are registered automatically during workflow execution via the .save() method. The system handles versioning, caching, and metadata extraction without requiring manual pre-registration of files.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →