# How to Integrate LaminDB with Nextflow and Snakemake Workflows

> Integrate LaminDB with Nextflow and Snakemake workflows. Seamlessly track data lineage and manage artifacts using ln.track() and ln.Artifact for robust scientific analysis pipelines.

- Repository: [K-Dense/scientific-agent-skills](https://github.com/K-Dense-AI/scientific-agent-skills)
- Tags: how-to-guide
- Published: 2026-05-14

---

**Integrate LaminDB into Nextflow and Snakemake pipelines by wrapping your analysis logic with `ln.track()` at the start and `ln.finish()` at the end, using `ln.Artifact.get()` to load inputs and `ln.Artifact(...).save()` to persist outputs with automatic lineage tracking.**

The `lamindb` library provides a lightweight data-management layer designed specifically for computational biology pipelines. According to the K-Dense-AI/scientific-agent-skills repository, you can integrate LaminDB with Nextflow and Snakemake workflows by embedding Python tracking calls directly inside process scripts and rules. This approach captures complete provenance—including input dependencies, output artifacts, and environment parameters—without modifying your underlying workflow engine configuration.

## The Five-Step Integration Pattern

Every pipeline step follows a consistent pattern documented in [`scientific-skills/lamindb/references/integrations.md`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/scientific-skills/lamindb/references/integrations.md). Implement these five stages to ensure full data lineage and provenance tracking:

1. **Initialize tracking**: Call `ln.track()` to start a new LaminDB run, record environment details, and prepare a cache for artifacts.
2. **Load inputs**: Retrieve previously stored artifacts using `ln.Artifact.get(key=...)` and access their data with the `.load()` method.
3. **Process data**: Execute your analysis logic (pandas transformations, statistical modeling, etc.).
4. **Persist outputs**: Save results back to LaminDB using `ln.Artifact.from_dataframe(...).save()` or `ln.Artifact(path, key=...).save()`.
5. **Finalize**: Close the run with `ln.finish()` to write provenance information including git hashes, timestamps, and parameters.

Because LaminDB stores **metadata** and **lineage** (input → output), downstream steps can query artifacts using their unique keys without re-computing intermediate results.

## Nextflow Integration

### Embedding LaminDB in a Nextflow Process

In Nextflow, embed LaminDB calls within the `script:` block of a process definition. The Python code executes inside the process environment, requiring only that `lamindb` is installed in the container or conda environment.

```nextflow
process ANALYZE {
    input:
    val input_key
    output:
    path "result.csv"
    script:
    """
    #!/usr/bin/env python
    import lamindb as ln

    # Start tracking for this process

    ln.track()

    # Load the input artifact from LaminDB

    artifact = ln.Artifact.get(key="${input_key}")
    data = artifact.load()

    # ---- Your analysis logic ----

    # Example: simple pandas transformation

    import pandas as pd
    df = pd.read_csv(data)
    result = df.groupby("category").sum().reset_index()
    result.to_csv("result.csv", index=False)

    # Register the output artifact

    ln.Artifact("result.csv", key="outputs/result.csv").save()

    # Finish tracking – writes provenance to the DB

    ln.finish()
    """
}

```

The `ln.track()` call creates a LaminDB run entry that records the current Git commit, environment variables, and any parameters passed via `ln.track(params={...})`. The `ln.Artifact.get()` method fetches the input using the key supplied by the Nextflow workflow (`${input_key}`). After processing, `ln.finish()` finalizes the run and links the output artifact to the run record.

This implementation mirrors the reference code found at lines 99–124 in [`scientific-skills/lamindb/references/integrations.md`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/scientific-skills/lamindb/references/integrations.md).

## Snakemake Integration

### Adding Tracking to a Snakemake Rule

Snakemake supports arbitrary Python code via the `run:` directive, making LaminDB integration straightforward. Wrap your analysis logic with the same tracking calls used in Nextflow.

```python
rule process_data:
    input:
        "data/input.csv"
    output:
        "data/output.csv"
    run:
        import lamindb as ln
        import pandas as pd

        # Start a LaminDB run for this rule

        ln.track()

        # Load the input artifact registered previously

        artifact = ln.Artifact.get(key="inputs/data.csv")
        df = artifact.load()

        # ---- Your analysis logic ----

        result = df.assign(mean=df.value.mean())
        result.to_csv(output[0], index=False)

        # Save the output artifact back to LaminDB

        ln.Artifact(output[0], key="outputs/result.csv").save()

        # End the LaminDB run

        ln.finish()

```

The artifact key (e.g., `"inputs/data.csv"` and `"outputs/result.csv"`) acts as a **global identifier** that other rules can reference with `ln.Artifact.get()`. Because the run is recorded, you can later query the lineage using `ln.ViewLineage()` to see which rule produced which artifact.

This example follows the official Snakemake integration recipe at lines 146–174 in [`scientific-skills/lamindb/references/integrations.md`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/scientific-skills/lamindb/references/integrations.md).

## Key Source Files and References

The following files in the K-Dense-AI/scientific-agent-skills repository provide complete implementation details:

- [`scientific-skills/lamindb/references/integrations.md`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/scientific-skills/lamindb/references/integrations.md) — Contains the full reference for Nextflow and Snakemake integration, including the code examples above.
- [`scientific-skills/lamindb/SKILL.md`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/scientific-skills/lamindb/SKILL.md) — Provides high-level LaminDB installation instructions (`pip install lamindb`) and example imports.
- [`scientific-skills/lamindb/references/setup-deployment.md`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/scientific-skills/lamindb/references/setup-deployment.md) — Documents installation commands and cloud storage extras (e.g., `lamindb[s3]`, `lamindb[gcp]`).
- [`scientific-skills/latchbio-integration/references/workflow-creation.md`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/scientific-skills/latchbio-integration/references/workflow-creation.md) — Shows how to register Nextflow and Snakemake workflows with LatchBio using `latch register --nextflow` and `latch register --snakemake`.

## Summary

- Wrap every workflow step with `ln.track()` and `ln.finish()` to capture complete provenance, including git hashes and timestamps.
- Use `ln.Artifact.get(key=...)` to load inputs and `ln.Artifact(...).save()` to register outputs with automatic versioning.
- Both Nextflow `process` scripts and Snakemake `run` blocks support inline Python execution for LaminDB calls.
- Artifact keys serve as global identifiers that enable cross-step data lineage tracking via `ln.ViewLineage()`.
- Implementation details are maintained in [`scientific-skills/lamindb/references/integrations.md`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/scientific-skills/lamindb/references/integrations.md) within the K-Dense-AI/scientific-agent-skills repository.

## Frequently Asked Questions

### Can I integrate LaminDB with workflow managers other than Nextflow and Snakemake?

Yes. While the K-Dense-AI/scientific-agent-skills repository specifically documents Nextflow and Snakemake, LaminDB’s Python API can be embedded in any workflow manager that supports Python script execution, including Apache Airflow, Prefect, or shell-based pipelines that invoke Python scripts.

### How does LaminDB track lineage between workflow steps?

LaminDB automatically links input and output artifacts to the current run record when you use `ln.Artifact.get()` and `ln.Artifact(...).save()`. You can query this lineage programmatically using `ln.ViewLineage()` to visualize the complete graph of which pipeline steps produced specific artifacts and what parameters were used.

### What Python environment setup is required to use LaminDB in a pipeline?

The execution environment must have `lamindb` installed, typically via `pip install lamindb` as documented in [`scientific-skills/lamindb/references/setup-deployment.md`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/scientific-skills/lamindb/references/setup-deployment.md). If your pipeline uses cloud storage backends, install the appropriate extras such as `lamindb[s3]` for AWS or `lamindb[gcp]` for Google Cloud Platform.

### Do I need to pre-register input artifacts before running the workflow?

No. While you query existing artifacts with `ln.Artifact.get(key=...)`, new outputs are registered automatically during workflow execution via the `.save()` method. The system handles versioning, caching, and metadata extraction without requiring manual pre-registration of files.