Example Applications Using Marin: Tutorials and Sample Scripts for ML Pipelines

Marin ships with four ready-to-run example applications in experiments/tutorials that demonstrate lazy artifact pipelines, device-agnostic model training, and distributed ML workflows.

The marin-community/marin repository includes a comprehensive set of example applications using Marin that serve as practical starting points for ML engineers. These self-contained scripts illustrate how to leverage Marin’s core abstractions—ArtifactStep, StepRunner, and BuildContext—to execute everything from simple data processing to multi-billion-parameter training runs on heterogeneous hardware.

Where to Find Marin Example Applications

All official examples reside in the experiments/tutorials package at the repository root. Each script is a self-contained entry point that can be launched with python -m and is fully type-checked according to the project’s style guidelines. The tutorials progress from a minimal two-step pipeline to a full-scale DCLM reproduction, demonstrating Marin’s incremental complexity model.

Hello World: Your First Lazy Artifact Pipeline

The hello_world.py script provides the simplest introduction to Marin’s execution model. It constructs a two-step lazy artifact pipeline that writes a JSON file of numbers, then subsequently reads that file to compute basic statistics.

In experiments/tutorials/hello_world.py, the pipeline defines two configuration dataclasses—GenerateDataConfig and ComputeStatsConfig—that parameterize pure functions for data generation and statistical computation:


# experiments/tutorials/hello_world.py

from marin.execution.artifact import Artifact
from marin.execution.lazy import ArtifactStep, lower
from marin.execution.step_runner import StepRunner
import json, os
from dataclasses import dataclass

@dataclass(frozen=True)
class GenerateDataConfig:
    n: int
    output_path: str

@dataclass(frozen=True)
class ComputeStatsConfig:
    input_path: str
    output_path: str

def generate_data(cfg: GenerateDataConfig) -> None:
    numbers = list(range(cfg.n))
    with open(os.path.join(cfg.output_path, "numbers.json"), "w") as f:
        json.dump(numbers, f)

def compute_stats(cfg: ComputeStatsConfig) -> None:
    with open(os.path.join(cfg.input_path, "numbers.json")) as f:
        numbers = json.load(f)
    stats = {"sum": sum(numbers), "min": min(numbers), "max": max(numbers)}
    with open(os.path.join(cfg.output_path, "stats.json"), "w") as f:
        json.dump(stats, f)

_data = ArtifactStep(
    name="hello_world/data",
    version="dev",
    artifact_type=Artifact,
    run=generate_data,
    build_config=lambda ctx: GenerateDataConfig(n=100, output_path=ctx.output_path),
)

_stats = ArtifactStep(
    name="hello_world/stats",
    version="dev",
    artifact_type=Artifact,
    run=compute_stats,
    build_config=lambda ctx: ComputeStatsConfig(
        input_path=ctx.artifact_path(_data), output_path=ctx.output_path
    ),
    deps=(_data,),
)

def build() -> ArtifactStep[Artifact]:
    return _stats

if __name__ == "__main__":
    StepRunner().run([lower(build())])

This example demonstrates how ArtifactStep declares lazy computations and how StepRunner executes the resulting directed acyclic graph (DAG).

Training Tiny Language Models on Any Device

The train_tiny_model.py tutorial shows how to train a tiny Llama model across different accelerators and datasets without modifying core logic. The script uses a declarative approach to specify ResourceConfig for CPUs, GPUs, or TPUs, and supports datasets like TinyStories, WikiText-2, and FineWeb-Edu.

Key components in experiments/tutorials/train_tiny_model.py include a DEVICES dictionary mapping hardware names to resource configurations, and a dataset() factory function that returns ArtifactStep instances for tokenized data:


# experiments/tutorials/train_tiny_model.py

import click
from fray.types import ResourceConfig
from levanter.optim.config import AdamConfig
from marin.execution.lazy import ArtifactStep
from marin.experiment.cli import build_options
from marin.experiment.data import pretokenized, tokenized
from marin.experiment.train import train_lm
from marin.training.training import LevanterCheckpoint
from experiments.llama import llama_150m, llama_nano
from experiments.marin_tokenizer import marin_tokenizer

DEVICES = {
    "cpu": (ResourceConfig.with_cpu(), 4),
    "h100x8": (ResourceConfig.with_gpu("H100", count=8, cpu=32), 256),
    # … other devices omitted for brevity …

}

RAW_SOURCES = {
    "tinystories": "roneneldan/TinyStories",
    "wikitext": "dlwh/wikitext_2_detokenized",
}

def dataset(name: str) -> ArtifactStep[TokenizedCache]:
    if name == "fineweb-edu":
        return pretokenized(
            "fineweb-edu-10M",
            repo_id="marin-community/fineweb-edu-pretokenized-10M",
            tokenizer=marin_tokenizer,
            version="2026.06.28",
        )
    return tokenized(
        name,
        tokenizer=marin_tokenizer,
        source=RAW_SOURCES[name],
        sample_count=1_000,
        version="2026.06.28",
    )

def build(*, device: str, data: str) -> ArtifactStep[LevanterCheckpoint]:
    resources, batch_size = DEVICES[device]
    model = llama_150m if data == "fineweb-edu" else llama_nano
    return train_lm(
        name=f"checkpoints/tiny-{data}-{device}",
        run_id=f"tiny-{data}-{device}",
        model=model,
        optimizer=AdamConfig(learning_rate=6e-4, weight_decay=0.1),
        datasets={dataset(data): 1.0},
        batch_size=batch_size,
        seq_len=model.max_seq_len,
        num_train_steps=100,
        resources=resources,
        tags=["llama", "tutorial", data, device],
    )

@click.command()
@click.option("--device", type=click.Choice(tuple(DEVICES)), default="cpu")
@click.option("--dataset", "data", type=click.Choice(("tinystories", "wikitext", "fineweb-edu")), default="tinystories")
@build_options
def main(device: str, data: str) -> ArtifactStep[LevanterCheckpoint]:
    return build(device=device, data=data)

if __name__ == "__main__":
    main()

The build() function returns a training step configured for specific hardware, demonstrating Marin’s ability to abstract device placement while maintaining type safety through ArtifactStep[LevanterCheckpoint] return types.

Running Hyperparameter Sweeps

For experimenting across hardware and dataset combinations, experiments/tutorials/train_tiny_sweep.py extends the tiny-model example to execute a parameter sweep. This script leverages Marin’s sweep utilities to parallelize multiple training jobs across a matrix of devices and data sources, tracking each run as a distinct artifact with its own dependency graph.

Reproducing Large-Scale Experiments

The exp1078_reproduce_dclm_7b1x.py example demonstrates Marin’s capacity for serious distributed training. Located in experiments/tutorials/, this script reproduces a 7-billion-parameter DCLM (Data Compute Language Model) training run on a modest cluster. It stitches together complex data pipelines, custom tokenization, and multi-node synchronization using the same ArtifactStep and StepRunner abstractions found in the simpler tutorials, proving that the framework scales from laptop prototypes to cluster-scale reproductions without changing the underlying API.

Core Abstractions Used in the Examples

All tutorials rely on three foundational classes defined in the Marin source:

These abstractions allow the example applications to remain agnostic about whether they run on a local CPU, a single GPU workstation, or a multi-node TPU pod.

Summary

  • Marin provides four official tutorials in experiments/tutorials/: hello_world.py, train_tiny_model.py, train_tiny_sweep.py, and exp1078_reproduce_dclm_7b1x.py.
  • Entry points are self-contained and runnable via python -m without additional scaffolding.
  • Lazy artifact pipelines use ArtifactStep declarations in lib/marin/src/marin/execution/lazy.py to build execution graphs.
  • Device abstraction allows the same training code to target CPU, GPU, or TPU by changing ResourceConfig parameters.
  • Type safety is enforced throughout, with the examples serving as reference implementations for custom pipeline development.

Frequently Asked Questions

How do I run the Marin hello world example?

Execute python -m experiments.tutorials.hello_world from the repository root. The script uses StepRunner to execute the DAG defined by the _data and _stats ArtifactStep instances, writing numbers.json and stats.json to the configured output paths.

What hardware accelerators are supported in the tiny model tutorial?

The train_tiny_model.py example supports CPU, NVIDIA GPUs (via ResourceConfig.with_gpu()), and TPUs. The DEVICES dictionary maps string keys like "cpu" and "h100x8" to specific ResourceConfig instances and batch sizes, allowing you to target anything from a laptop to an 8-GPU H100 node.

Can I modify the example applications for my own datasets?

Yes. You can extend the dataset() function in train_tiny_model.py to return new ArtifactStep instances pointing to your HuggingFace datasets or local pretokenized caches. The build() function and train_lm API remain unchanged regardless of data source.

Where is the ArtifactStep class defined in the Marin source code?

ArtifactStep is defined in lib/marin/src/marin/execution/lazy.py alongside the lower() function used to transform step declarations into runnable tasks. The class implements the lazy evaluation logic that underpins all Marin pipelines.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →