# Example Applications Using Marin: Tutorials and Sample Scripts for ML Pipelines

> Explore example applications using Marin to build lazy artifact pipelines, device-agnostic model training, and distributed ML workflows. Get tutorials and sample scripts now.

- Repository: [The Marin Project/marin](https://github.com/marin-community/marin)
- Tags: tutorial
- Published: 2026-08-27

---

**Marin ships with four ready-to-run example applications in `experiments/tutorials` that demonstrate lazy artifact pipelines, device-agnostic model training, and distributed ML workflows.**

The `marin-community/marin` repository includes a comprehensive set of **example applications using Marin** that serve as practical starting points for ML engineers. These self-contained scripts illustrate how to leverage Marin’s core abstractions—**ArtifactStep**, **StepRunner**, and **BuildContext**—to execute everything from simple data processing to multi-billion-parameter training runs on heterogeneous hardware.

## Where to Find Marin Example Applications

All official examples reside in the `experiments/tutorials` package at the repository root. Each script is a self-contained entry point that can be launched with `python -m` and is fully type-checked according to the project’s style guidelines. The tutorials progress from a minimal two-step pipeline to a full-scale DCLM reproduction, demonstrating Marin’s incremental complexity model.

## Hello World: Your First Lazy Artifact Pipeline

The [`hello_world.py`](https://github.com/marin-community/marin/blob/main/hello_world.py) script provides the simplest introduction to Marin’s execution model. It constructs a two-step lazy artifact pipeline that writes a JSON file of numbers, then subsequently reads that file to compute basic statistics.

In [`experiments/tutorials/hello_world.py`](https://github.com/marin-community/marin/blob/main/experiments/tutorials/hello_world.py), the pipeline defines two configuration dataclasses—**GenerateDataConfig** and **ComputeStatsConfig**—that parameterize pure functions for data generation and statistical computation:

```python

# experiments/tutorials/hello_world.py

from marin.execution.artifact import Artifact
from marin.execution.lazy import ArtifactStep, lower
from marin.execution.step_runner import StepRunner
import json, os
from dataclasses import dataclass

@dataclass(frozen=True)
class GenerateDataConfig:
    n: int
    output_path: str

@dataclass(frozen=True)
class ComputeStatsConfig:
    input_path: str
    output_path: str

def generate_data(cfg: GenerateDataConfig) -> None:
    numbers = list(range(cfg.n))
    with open(os.path.join(cfg.output_path, "numbers.json"), "w") as f:
        json.dump(numbers, f)

def compute_stats(cfg: ComputeStatsConfig) -> None:
    with open(os.path.join(cfg.input_path, "numbers.json")) as f:
        numbers = json.load(f)
    stats = {"sum": sum(numbers), "min": min(numbers), "max": max(numbers)}
    with open(os.path.join(cfg.output_path, "stats.json"), "w") as f:
        json.dump(stats, f)

_data = ArtifactStep(
    name="hello_world/data",
    version="dev",
    artifact_type=Artifact,
    run=generate_data,
    build_config=lambda ctx: GenerateDataConfig(n=100, output_path=ctx.output_path),
)

_stats = ArtifactStep(
    name="hello_world/stats",
    version="dev",
    artifact_type=Artifact,
    run=compute_stats,
    build_config=lambda ctx: ComputeStatsConfig(
        input_path=ctx.artifact_path(_data), output_path=ctx.output_path
    ),
    deps=(_data,),
)

def build() -> ArtifactStep[Artifact]:
    return _stats

if __name__ == "__main__":
    StepRunner().run([lower(build())])

```

This example demonstrates how **ArtifactStep** declares lazy computations and how **StepRunner** executes the resulting directed acyclic graph (DAG).

## Training Tiny Language Models on Any Device

The [`train_tiny_model.py`](https://github.com/marin-community/marin/blob/main/train_tiny_model.py) tutorial shows how to train a tiny Llama model across different accelerators and datasets without modifying core logic. The script uses a declarative approach to specify **ResourceConfig** for CPUs, GPUs, or TPUs, and supports datasets like TinyStories, WikiText-2, and FineWeb-Edu.

Key components in [`experiments/tutorials/train_tiny_model.py`](https://github.com/marin-community/marin/blob/main/experiments/tutorials/train_tiny_model.py) include a `DEVICES` dictionary mapping hardware names to resource configurations, and a `dataset()` factory function that returns **ArtifactStep** instances for tokenized data:

```python

# experiments/tutorials/train_tiny_model.py

import click
from fray.types import ResourceConfig
from levanter.optim.config import AdamConfig
from marin.execution.lazy import ArtifactStep
from marin.experiment.cli import build_options
from marin.experiment.data import pretokenized, tokenized
from marin.experiment.train import train_lm
from marin.training.training import LevanterCheckpoint
from experiments.llama import llama_150m, llama_nano
from experiments.marin_tokenizer import marin_tokenizer

DEVICES = {
    "cpu": (ResourceConfig.with_cpu(), 4),
    "h100x8": (ResourceConfig.with_gpu("H100", count=8, cpu=32), 256),
    # … other devices omitted for brevity …

}

RAW_SOURCES = {
    "tinystories": "roneneldan/TinyStories",
    "wikitext": "dlwh/wikitext_2_detokenized",
}

def dataset(name: str) -> ArtifactStep[TokenizedCache]:
    if name == "fineweb-edu":
        return pretokenized(
            "fineweb-edu-10M",
            repo_id="marin-community/fineweb-edu-pretokenized-10M",
            tokenizer=marin_tokenizer,
            version="2026.06.28",
        )
    return tokenized(
        name,
        tokenizer=marin_tokenizer,
        source=RAW_SOURCES[name],
        sample_count=1_000,
        version="2026.06.28",
    )

def build(*, device: str, data: str) -> ArtifactStep[LevanterCheckpoint]:
    resources, batch_size = DEVICES[device]
    model = llama_150m if data == "fineweb-edu" else llama_nano
    return train_lm(
        name=f"checkpoints/tiny-{data}-{device}",
        run_id=f"tiny-{data}-{device}",
        model=model,
        optimizer=AdamConfig(learning_rate=6e-4, weight_decay=0.1),
        datasets={dataset(data): 1.0},
        batch_size=batch_size,
        seq_len=model.max_seq_len,
        num_train_steps=100,
        resources=resources,
        tags=["llama", "tutorial", data, device],
    )

@click.command()
@click.option("--device", type=click.Choice(tuple(DEVICES)), default="cpu")
@click.option("--dataset", "data", type=click.Choice(("tinystories", "wikitext", "fineweb-edu")), default="tinystories")
@build_options
def main(device: str, data: str) -> ArtifactStep[LevanterCheckpoint]:
    return build(device=device, data=data)

if __name__ == "__main__":
    main()

```

The `build()` function returns a training step configured for specific hardware, demonstrating Marin’s ability to abstract device placement while maintaining type safety through **ArtifactStep[LevanterCheckpoint]** return types.

## Running Hyperparameter Sweeps

For experimenting across hardware and dataset combinations, [`experiments/tutorials/train_tiny_sweep.py`](https://github.com/marin-community/marin/blob/main/experiments/tutorials/train_tiny_sweep.py) extends the tiny-model example to execute a parameter sweep. This script leverages Marin’s sweep utilities to parallelize multiple training jobs across a matrix of devices and data sources, tracking each run as a distinct artifact with its own dependency graph.

## Reproducing Large-Scale Experiments

The [`exp1078_reproduce_dclm_7b1x.py`](https://github.com/marin-community/marin/blob/main/exp1078_reproduce_dclm_7b1x.py) example demonstrates Marin’s capacity for serious distributed training. Located in `experiments/tutorials/`, this script reproduces a 7-billion-parameter DCLM (Data Compute Language Model) training run on a modest cluster. It stitches together complex data pipelines, custom tokenization, and multi-node synchronization using the same **ArtifactStep** and **StepRunner** abstractions found in the simpler tutorials, proving that the framework scales from laptop prototypes to cluster-scale reproductions without changing the underlying API.

## Core Abstractions Used in the Examples

All tutorials rely on three foundational classes defined in the Marin source:

- **ArtifactStep** – Defined in [`lib/marin/src/marin/execution/lazy.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/execution/lazy.py), this class represents a lazy computation node in the pipeline DAG. It encapsulates the execution function, configuration builder, and dependency list.
- **StepRunner** – Implemented in [`lib/marin/src/marin/execution/step_runner.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/execution/step_runner.py), this class executes the directed acyclic graph of steps, handling caching, parallelization, and artifact serialization.
- **Artifact** – Located in [`lib/marin/src/marin/execution/artifact.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/execution/artifact.py), this base class represents immutable data outputs that flow between steps, enabling reproducible pipeline states.

These abstractions allow the example applications to remain agnostic about whether they run on a local CPU, a single GPU workstation, or a multi-node TPU pod.

## Summary

- **Marin provides four official tutorials** in `experiments/tutorials/`: [`hello_world.py`](https://github.com/marin-community/marin/blob/main/hello_world.py), [`train_tiny_model.py`](https://github.com/marin-community/marin/blob/main/train_tiny_model.py), [`train_tiny_sweep.py`](https://github.com/marin-community/marin/blob/main/train_tiny_sweep.py), and [`exp1078_reproduce_dclm_7b1x.py`](https://github.com/marin-community/marin/blob/main/exp1078_reproduce_dclm_7b1x.py).
- **Entry points are self-contained** and runnable via `python -m` without additional scaffolding.
- **Lazy artifact pipelines** use **ArtifactStep** declarations in [`lib/marin/src/marin/execution/lazy.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/execution/lazy.py) to build execution graphs.
- **Device abstraction** allows the same training code to target CPU, GPU, or TPU by changing **ResourceConfig** parameters.
- **Type safety** is enforced throughout, with the examples serving as reference implementations for custom pipeline development.

## Frequently Asked Questions

### How do I run the Marin hello world example?

Execute `python -m experiments.tutorials.hello_world` from the repository root. The script uses **StepRunner** to execute the DAG defined by the `_data` and `_stats` **ArtifactStep** instances, writing [`numbers.json`](https://github.com/marin-community/marin/blob/main/numbers.json) and [`stats.json`](https://github.com/marin-community/marin/blob/main/stats.json) to the configured output paths.

### What hardware accelerators are supported in the tiny model tutorial?

The [`train_tiny_model.py`](https://github.com/marin-community/marin/blob/main/train_tiny_model.py) example supports CPU, NVIDIA GPUs (via `ResourceConfig.with_gpu()`), and TPUs. The `DEVICES` dictionary maps string keys like `"cpu"` and `"h100x8"` to specific **ResourceConfig** instances and batch sizes, allowing you to target anything from a laptop to an 8-GPU H100 node.

### Can I modify the example applications for my own datasets?

Yes. You can extend the `dataset()` function in [`train_tiny_model.py`](https://github.com/marin-community/marin/blob/main/train_tiny_model.py) to return new **ArtifactStep** instances pointing to your HuggingFace datasets or local pretokenized caches. The `build()` function and **train_lm** API remain unchanged regardless of data source.

### Where is the ArtifactStep class defined in the Marin source code?

**ArtifactStep** is defined in [`lib/marin/src/marin/execution/lazy.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/execution/lazy.py) alongside the `lower()` function used to transform step declarations into runnable tasks. The class implements the lazy evaluation logic that underpins all Marin pipelines.