# Example Projects Using Marin: 4 Ready-to-Run Tutorials

> Explore example projects using Marin with four ready-to-run tutorials. Learn lazy artifact pipelines, tokenization, and model training with StepRunner and ArtifactStep.

- Repository: [The Marin Project/marin](https://github.com/marin-community/marin)
- Tags: tutorial
- Published: 2026-08-29

---

**Marin ships with four self-contained example projects in `experiments/tutorials/` that demonstrate lazy-artifact pipelines, tokenization, and model training using `StepRunner` and `ArtifactStep`.**

The `marin-community/marin` repository includes a dedicated tutorials directory containing executable scripts that walk through the framework's core abstractions. These example projects using Marin require only the base installation to run and illustrate everything from simple data pipelines to full-scale language model training.

## Hello World: Your First Lazy-Artifact Pipeline

The **Hello World** example provides the simplest introduction to Marin's execution model. Located at [`experiments/tutorials/hello_world.py`](https://github.com/marin-community/marin/blob/main/experiments/tutorials/hello_world.py), this script constructs a two-step pipeline that generates a list of numbers and computes summary statistics, demonstrating how `ArtifactStep` declarations automatically handle dependencies and caching.

The script defines two functions—`generate_data` and `compute_stats`—and wraps them in `ArtifactStep` objects. The `_data` step uses `build_config` to specify output paths via `ctx.output_path`, while the `_stats` step accesses the previous artifact's location using `ctx.artifact_path(_data)`. Dependencies are explicitly declared via the `deps` parameter, ensuring Marin executes `_data` before `_stats`.

```python

# experiments/tutorials/hello_world.py (excerpt)

from marin.execution.artifact import Artifact
from marin.execution.lazy import ArtifactStep, lower
from marin.execution.step_runner import StepRunner
from rigging.filesystem.factory import open_url

def generate_data(config):
    numbers = list(range(config.n))
    with open_url(os.path.join(config.output_path, "numbers.json"), "w") as f:
        json.dump(numbers, f)

def compute_stats(config):
    with open_url(os.path.join(config.input_path, "numbers.json")) as f:
        numbers = json.load(f)
    stats = {"sum": sum(numbers), "min": min(numbers), "max": max(numbers)}
    with open_url(os.path.join(config.output_path, "stats.json"), "w") as f:
        json.dump(stats, f)

_data = ArtifactStep(
    name="hello_world/data",
    version="dev",
    artifact_type=Artifact,
    run=generate_data,
    build_config=lambda ctx: GenerateDataConfig(n=100, output_path=ctx.output_path),
)

_stats = ArtifactStep(
    name="hello_world/stats",
    version="dev",
    artifact_type=Artifact,
    run=compute_stats,
    build_config=lambda ctx: ComputeStatsConfig(
        input_path=ctx.artifact_path(_data), output_path=ctx.output_path
    ),
    deps=(_data,),
)

if __name__ == "__main__":
    StepRunner().run([lower(_stats)])

```

Running this file executes the dependency graph via `StepRunner().run([lower(_stats)])`, where `lower()` converts the lazy step definition into an executable specification.

## Tiny Model Training: End-to-End LLM Workflow

The **Tiny Model Training** tutorial at [`experiments/tutorials/train_tiny_model.py`](https://github.com/marin-community/marin/blob/main/experiments/tutorials/train_tiny_model.py) demonstrates a complete machine learning pipeline using the **TinyStories** dataset. This example projects using Marin shows how to wire together tokenization and training steps without manual file management.

The script utilizes `tokenized()` to preprocess the HuggingFace `roneneldan/TinyStories` dataset with a custom `marin_tokenizer`, producing a lazy artifact named `"tokenized/tinystories"`. It then passes this artifact to `train_lm()`, which configures a `llama_nano` model with `AdamConfig` optimizers and executes a 100-step training run.

```python

# experiments/tutorials/train_tiny_model.py (excerpt)

from marin.execution.lazy import lower
from marin.execution.step_runner import StepRunner
from marin.experiment.data import tokenized
from marin.experiment.train import train_lm
from experiments.marin_tokenizer import marin_tokenizer
from experiments.llama import llama_nano
from fray.cluster import ResourceConfig
from levanter.optim import AdamConfig

tinystories_tokenized = tokenized(
    name="tokenized/tinystories",
    source="roneneldan/TinyStories",
    tokenizer=marin_tokenizer,
    sample_count=1000,
)

nano_tinystories_model = train_lm(
    name="checkpoints/marin-nano-tinystories",
    version="v1",
    model=llama_nano,
    optimizer=AdamConfig(learning_rate=6e-4, weight_decay=0.1),
    datasets={tinystories_tokenized: 1.0},
    batch_size=4,
    seq_len=2048,
    num_train_steps=100,
    resources=ResourceConfig.with_cpu(),
)

if __name__ == "__main__":
    StepRunner().run([lower(nano_tinystories_model)])

```

Notice how `datasets` accepts a dictionary mapping the tokenized artifact to a sampling weight (1.0), allowing Marin to automatically resolve the data dependency before training begins.

## Hyperparameter Sweeps and Scale Experiments

Beyond basic workflows, the tutorials include advanced examples for systematic experimentation and reproduction of published results.

### Running Hyperparameter Sweeps

The [`experiments/tutorials/train_tiny_sweep.py`](https://github.com/marin-community/marin/blob/main/experiments/tutorials/train_tiny_sweep.py) file extends the tiny model training example to execute multiple runs with varying configurations. This script demonstrates how to parameterize `train_lm()` calls—modifying learning rates or architectural hyperparameters—and execute them via `StepRunner` as a batch of independent experiments.

### Reproducing DCLM at Scale

For researchers reproducing published benchmarks, [`experiments/tutorials/exp1078_reproduce_dclm_7b1x.py`](https://github.com/marin-community/marin/blob/main/experiments/tutorials/exp1078_reproduce_dclm_7b1x.py) provides a complete script that replicates the DCLM 1B/1x experiment on modest hardware. This example showcases **scaling-suite usage**, including dataset mixing strategies, checkpoint handling, and multi-step artifact chains typical in production LLM training.

## How to Run the Examples

All example projects using Marin are executable Python scripts requiring only the base Marin installation. According to the repository's installation documentation, once Marin is installed, you can run any tutorial directly:

```bash

# Run the Hello World example

python experiments/tutorials/hello_world.py

# Train the tiny model

python experiments/tutorials/train_tiny_model.py

# Execute hyperparameter sweep

python experiments/tutorials/train_tiny_sweep.py

```

The `StepRunner` automatically handles directory creation, artifact caching, and dependency resolution. Each script is self-contained and writes outputs to configurable paths via the context objects (`ctx.output_path`).

## Summary

- **Four official examples** live in `experiments/tutorials/` covering lazy pipelines, training, sweeps, and reproduction scripts.
- **Hello World** ([`hello_world.py`](https://github.com/marin-community/marin/blob/main/hello_world.py)) teaches `ArtifactStep`, `lower()`, and `StepRunner` basics with a two-step data pipeline.
- **Tiny Model Training** ([`train_tiny_model.py`](https://github.com/marin-community/marin/blob/main/train_tiny_model.py)) demonstrates end-to-end LLM workflows using `tokenized()` and `train_lm()`.
- **Advanced tutorials** include hyperparameter sweeps ([`train_tiny_sweep.py`](https://github.com/marin-community/marin/blob/main/train_tiny_sweep.py)) and full-scale DCLM reproduction ([`exp1078_reproduce_dclm_7b1x.py`](https://github.com/marin-community/marin/blob/main/exp1078_reproduce_dclm_7b1x.py)).
- All examples use **lazy evaluation** via `lower()` and explicit **dependency declarations** through `deps` parameters or artifact references.

## Frequently Asked Questions

### Where are the Marin example projects located?

All example projects reside in the `experiments/tutorials/` directory at the repository root. The primary entry points are [`hello_world.py`](https://github.com/marin-community/marin/blob/main/hello_world.py), [`train_tiny_model.py`](https://github.com/marin-community/marin/blob/main/train_tiny_model.py), [`train_tiny_sweep.py`](https://github.com/marin-community/marin/blob/main/train_tiny_sweep.py), and [`exp1078_reproduce_dclm_7b1x.py`](https://github.com/marin-community/marin/blob/main/exp1078_reproduce_dclm_7b1x.py), each demonstrating different aspects of the framework's lazy-artifact system.

### What dependencies are required to run the Marin tutorials?

The examples require only the base Marin installation as documented in the repository's installation guide. Specific examples like the Tiny Model Training may pull additional dependencies (such as `levanter` for optimization configs or `fray` for resource management) automatically when Marin is installed with recommended extras.

### How do the Marin examples handle artifact dependencies?

Dependencies are declared explicitly using the `deps` parameter in `ArtifactStep` constructors or implicitly by passing tokenized artifacts to training functions. Marin uses `ctx.artifact_path()` to resolve these dependencies at runtime, ensuring that upstream steps complete before downstream steps execute.

### Can I modify the example configurations for my own datasets?

Yes. Each example uses configuration objects (like `GenerateDataConfig` or `AdamConfig`) that accept custom parameters. You can replace the TinyStories dataset reference in [`train_tiny_model.py`](https://github.com/marin-community/marin/blob/main/train_tiny_model.py) with any HuggingFace dataset identifier, or adjust the `sample_count`, `batch_size`, and `num_train_steps` parameters to suit your computational constraints.