# Best Practices for Using the Marin Framework: A Complete Guide

> Master Marin best practices for reproducible pipelines. Leverage lazy artifacts, explicit versioning, and immutable data classes for efficient language model development.

- Repository: [The Marin Project/marin](https://github.com/marin-community/marin)
- Tags: best-practices
- Published: 2026-08-27

---

**The best practices for using Marin center on leveraging lazy artifacts with explicit versioning, separating build-time configuration from runtime arguments, and adopting immutable data classes to ensure reproducible, cache-efficient language model pipelines.**

The marin-community/marin repository provides a modular, lazy-execution framework designed for building large-scale language-model pipelines. Following the best practices for using Marin ensures your experiments remain declarative, cache-friendly, and reproducible across different hardware configurations. These guidelines derive directly from the framework's core architecture as implemented in [`lib/marin/src/marin/execution/lazy.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/execution/lazy.py) and the official development standards.

## Leverage Lazy Artifacts for Declarative Pipelines

Every step in a Marin experiment returns an `ArtifactStep[T]` handle that contains only an identity (`name@version`) and a pure `build_config` function. No computation occurs until the step is lowered and executed, making pipelines both declarative and cache-friendly.

### Understand the ArtifactStep Handle

According to [`docs/explanations/lazy-artifacts.md`](https://github.com/marin-community/marin/blob/main/docs/explanations/lazy-artifacts.md), the handle represents a promise of computation rather than the result itself. The framework tracks dependencies through these handles, enabling automatic caching and parallel execution. When you compose multiple steps, you build a directed acyclic graph that Marin can optimize before execution.

### Use High-Level Experiment Helpers

Instead of constructing `ArtifactStep` directly, use the high-level helpers provided in [`lib/marin/src/marin/experiment/data.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/experiment/data.py) such as `tokenized`, `train_lm`, and `hf_download`. These functions encapsulate common patterns and automatically register dependencies, reducing boilerplate and ensuring consistent fingerprint generation.

```python
from marin.experiment.data import tokenized
from marin.training.train import train_lm

# Create a lazy handle to tokenized data (no work performed yet)

tinystories = tokenized(
    name="tokenized/tinystories",
    source="roneneldan/TinyStories",
    tokenizer=marin_tokenizer,
    version="2026.07.01",  # Bump to force re-tokenization

)

# Declare a training step that depends on the tokenized data

model_ckpt = train_lm(
    name="checkpoints/my-run",
    model=llama_nano,
    datasets={tinystories: 1.0},
    batch_size=4,
    seq_len=2048,
)

```

## Implement Version-Based Identity and Drift Detection

Marin identifies artifacts using the explicit path format `{prefix}/{name}/{version}`. Changing a literal value—such as a hyperparameter—does not alter this path; instead, it updates the **fingerprint** stored alongside the artifact.

### Handle Advisory Drift

When Marin detects a fingerprint mismatch between your code and the cached artifact, it issues an **advisory drift warning** and serves the cached output. To force a rebuild, explicitly bump the version string. This design prevents accidental recomputation while maintaining a clear audit trail of what changed.

### Use Development Versions for Rapid Iteration

During active development, use a `"dev"` version (or `"mylabel-dev"`) to skip the cache completely and always recompute. As documented in [`docs/explanations/lazy-artifacts.md`](https://github.com/marin-community/marin/blob/main/docs/explanations/lazy-artifacts.md), this practice is essential for debugging pipelines when you need to verify behavior without manipulating version strings repeatedly.

```python

# Always rebuild during development

model_ckpt = train_lm(
    name="checkpoints/my-run",
    version="dev",  # Triggers fresh execution every time

    model=llama_nano,
    optimizer=AdamConfig(lr=6e-4),
    num_train_steps=100,
)

```

## Adopt External Data with Proper Provenance

When integrating pre-existing data—such as a tokenized cache produced by an external run—bring it into the graph using `ArtifactStep.adopt`. This method, defined in [`lib/marin/src/marin/execution/lazy.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/execution/lazy.py), writes a provenance record at the canonical address and enables drift checking on future re-adoptions.

```python
from marin.execution.lazy import ArtifactStep
from marin.processing.tokenize.tokenize import TokenizedCache

external = ArtifactStep.adopt(
    name="tokenized/external",
    version="2025.12.01",
    adopt_source="gs://my-bucket/tokenized/external/",
    kind=TokenizedCache,
)

# Now `external` can be used as a normal dependency in other steps

```

## Separate Build-Time and Runtime Configuration

Values that influence *how* a step runs—such as CPU/GPU/TPU resources, storage prefix, or region—must be accessed via `ctx.runtime_arg(key)` within your step's build function. According to [`lib/marin/src/marin/execution/lazy.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/execution/lazy.py), these runtime arguments are explicitly excluded from the fingerprint, ensuring that switching hardware configurations never forces downstream artifacts to rebuild.

```python
def build_config(self, ctx):
    # Runtime argument: excluded from fingerprint

    device = ctx.runtime_arg("device")  # e.g., "cuda:0" or "tpu"

    
    # Build-time argument: included in fingerprint

    learning_rate = self.config.lr

```

## Follow Strict Python Coding Standards

The Marin project enforces specific conventions documented in [`docs/dev-guide/coding-standards.md`](https://github.com/marin-community/marin/blob/main/docs/dev-guide/coding-standards.md) to maintain code quality and prevent common errors in lazy-execution contexts.

### Optimize Import Structure

Keep all imports at the top of the file and avoid mid-function imports. This prevents hidden circular dependencies and improves static analysis capabilities across the codebase.

### Prefer Immutable Configuration Objects

Use **frozen dataclasses** for configuration objects to ensure immutability. Combine this with `dataclasses.replace` for creating modified copies of configurations. Prefer top-level functions and early returns to reduce nesting and make the lazy-artifact model easier to reason about.

## Maintain Explicit Configuration Without Hidden Defaults

All critical parameters—such as learning rate, batch size, and optimizer settings—should be passed explicitly to avoid hidden defaults that can silently alter fingerprints. Centralize default values in a single module and reference them explicitly. When you need to update a configuration, use `dataclasses.replace` to create a new immutable instance rather than mutating existing objects.

## Validate with Integration Testing

The [`tests/integration_test.py`](https://github.com/marin-community/marin/blob/main/tests/integration_test.py) file contains a miniature version of the full pipeline that runs in less than 10 minutes without requiring GPU or TPU access. Regularly run this test to catch regressions in artifact handling, fingerprinting logic, and version bump behavior. This practice ensures that changes to [`lib/marin/src/marin/execution/lazy.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/execution/lazy.py) or experiment helpers do not break the core lazy-execution model.

## Summary

- **Leverage lazy artifacts** using `ArtifactStep[T]` handles and high-level helpers like `tokenized` and `train_lm` to build declarative, cache-efficient pipelines.
- **Implement explicit versioning** with `{prefix}/{name}/{version}` paths, use `"dev"` versions to bypass cache during debugging, and handle advisory drift by bumping version strings when necessary.
- **Adopt external data** using `ArtifactStep.adopt` to maintain provenance records and enable drift detection on pre-existing artifacts.
- **Separate concerns** by accessing hardware and region settings through `ctx.runtime_arg(key)` to keep runtime changes from invalidating build caches.
- **Follow coding standards** by placing imports at file scope, using frozen dataclasses for configuration, and avoiding mid-function imports.
- **Enforce explicit configuration** by passing all parameters explicitly and using `dataclasses.replace` for immutable updates.
- **Run integration tests** from [`tests/integration_test.py`](https://github.com/marin-community/marin/blob/main/tests/integration_test.py) regularly to validate artifact handling and fingerprinting logic.

## Frequently Asked Questions

### What is the benefit of using lazy artifacts in Marin?

Lazy artifacts defer computation until execution time, creating declarative pipelines that are cache-friendly and reproducible. Each `ArtifactStep[T]` handle contains only identity metadata and a pure build function, allowing the framework to skip redundant work when inputs haven't changed, as implemented in [`lib/marin/src/marin/execution/lazy.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/execution/lazy.py).

### How do I force Marin to rebuild an artifact when debugging?

Use a `"dev"` version string (such as `"dev"` or `"mylabel-dev"`) to bypass the cache completely during rapid development. Alternatively, bump the explicit version number (e.g., from `"2026.07.01"` to `"2026.07.02"`) to invalidate the cached output for that specific step, triggering a fresh build according to the versioning logic in [`docs/explanations/lazy-artifacts.md`](https://github.com/marin-community/marin/blob/main/docs/explanations/lazy-artifacts.md).

### Why should runtime arguments be accessed separately from build configuration?

Runtime arguments like GPU allocation, TPU topology, or storage region are accessed via `ctx.runtime_arg(key)` and excluded from the artifact fingerprint. This separation ensures that switching hardware configurations or cloud regions never triggers unnecessary rebuilds of downstream artifacts, maintaining cache efficiency across different execution environments.

### What coding patterns should I avoid when writing Marin pipeline code?

Avoid mid-function imports, which can create hidden circular dependencies and hamper static analysis. According to the Marin coding standards in [`docs/dev-guide/coding-standards.md`](https://github.com/marin-community/marin/blob/main/docs/dev-guide/coding-standards.md), always place imports at the top of files, prefer top-level functions with early returns, and use frozen dataclasses for all configuration objects to ensure immutability.