Example Projects Using Marin: 4 Ready-to-Run Tutorials
Marin ships with four self-contained example projects in experiments/tutorials/ that demonstrate lazy-artifact pipelines, tokenization, and model training using StepRunner and ArtifactStep.
The marin-community/marin repository includes a dedicated tutorials directory containing executable scripts that walk through the framework's core abstractions. These example projects using Marin require only the base installation to run and illustrate everything from simple data pipelines to full-scale language model training.
Hello World: Your First Lazy-Artifact Pipeline
The Hello World example provides the simplest introduction to Marin's execution model. Located at experiments/tutorials/hello_world.py, this script constructs a two-step pipeline that generates a list of numbers and computes summary statistics, demonstrating how ArtifactStep declarations automatically handle dependencies and caching.
The script defines two functions—generate_data and compute_stats—and wraps them in ArtifactStep objects. The _data step uses build_config to specify output paths via ctx.output_path, while the _stats step accesses the previous artifact's location using ctx.artifact_path(_data). Dependencies are explicitly declared via the deps parameter, ensuring Marin executes _data before _stats.
# experiments/tutorials/hello_world.py (excerpt)
from marin.execution.artifact import Artifact
from marin.execution.lazy import ArtifactStep, lower
from marin.execution.step_runner import StepRunner
from rigging.filesystem.factory import open_url
def generate_data(config):
numbers = list(range(config.n))
with open_url(os.path.join(config.output_path, "numbers.json"), "w") as f:
json.dump(numbers, f)
def compute_stats(config):
with open_url(os.path.join(config.input_path, "numbers.json")) as f:
numbers = json.load(f)
stats = {"sum": sum(numbers), "min": min(numbers), "max": max(numbers)}
with open_url(os.path.join(config.output_path, "stats.json"), "w") as f:
json.dump(stats, f)
_data = ArtifactStep(
name="hello_world/data",
version="dev",
artifact_type=Artifact,
run=generate_data,
build_config=lambda ctx: GenerateDataConfig(n=100, output_path=ctx.output_path),
)
_stats = ArtifactStep(
name="hello_world/stats",
version="dev",
artifact_type=Artifact,
run=compute_stats,
build_config=lambda ctx: ComputeStatsConfig(
input_path=ctx.artifact_path(_data), output_path=ctx.output_path
),
deps=(_data,),
)
if __name__ == "__main__":
StepRunner().run([lower(_stats)])
Running this file executes the dependency graph via StepRunner().run([lower(_stats)]), where lower() converts the lazy step definition into an executable specification.
Tiny Model Training: End-to-End LLM Workflow
The Tiny Model Training tutorial at experiments/tutorials/train_tiny_model.py demonstrates a complete machine learning pipeline using the TinyStories dataset. This example projects using Marin shows how to wire together tokenization and training steps without manual file management.
The script utilizes tokenized() to preprocess the HuggingFace roneneldan/TinyStories dataset with a custom marin_tokenizer, producing a lazy artifact named "tokenized/tinystories". It then passes this artifact to train_lm(), which configures a llama_nano model with AdamConfig optimizers and executes a 100-step training run.
# experiments/tutorials/train_tiny_model.py (excerpt)
from marin.execution.lazy import lower
from marin.execution.step_runner import StepRunner
from marin.experiment.data import tokenized
from marin.experiment.train import train_lm
from experiments.marin_tokenizer import marin_tokenizer
from experiments.llama import llama_nano
from fray.cluster import ResourceConfig
from levanter.optim import AdamConfig
tinystories_tokenized = tokenized(
name="tokenized/tinystories",
source="roneneldan/TinyStories",
tokenizer=marin_tokenizer,
sample_count=1000,
)
nano_tinystories_model = train_lm(
name="checkpoints/marin-nano-tinystories",
version="v1",
model=llama_nano,
optimizer=AdamConfig(learning_rate=6e-4, weight_decay=0.1),
datasets={tinystories_tokenized: 1.0},
batch_size=4,
seq_len=2048,
num_train_steps=100,
resources=ResourceConfig.with_cpu(),
)
if __name__ == "__main__":
StepRunner().run([lower(nano_tinystories_model)])
Notice how datasets accepts a dictionary mapping the tokenized artifact to a sampling weight (1.0), allowing Marin to automatically resolve the data dependency before training begins.
Hyperparameter Sweeps and Scale Experiments
Beyond basic workflows, the tutorials include advanced examples for systematic experimentation and reproduction of published results.
Running Hyperparameter Sweeps
The experiments/tutorials/train_tiny_sweep.py file extends the tiny model training example to execute multiple runs with varying configurations. This script demonstrates how to parameterize train_lm() calls—modifying learning rates or architectural hyperparameters—and execute them via StepRunner as a batch of independent experiments.
Reproducing DCLM at Scale
For researchers reproducing published benchmarks, experiments/tutorials/exp1078_reproduce_dclm_7b1x.py provides a complete script that replicates the DCLM 1B/1x experiment on modest hardware. This example showcases scaling-suite usage, including dataset mixing strategies, checkpoint handling, and multi-step artifact chains typical in production LLM training.
How to Run the Examples
All example projects using Marin are executable Python scripts requiring only the base Marin installation. According to the repository's installation documentation, once Marin is installed, you can run any tutorial directly:
# Run the Hello World example
python experiments/tutorials/hello_world.py
# Train the tiny model
python experiments/tutorials/train_tiny_model.py
# Execute hyperparameter sweep
python experiments/tutorials/train_tiny_sweep.py
The StepRunner automatically handles directory creation, artifact caching, and dependency resolution. Each script is self-contained and writes outputs to configurable paths via the context objects (ctx.output_path).
Summary
- Four official examples live in
experiments/tutorials/covering lazy pipelines, training, sweeps, and reproduction scripts. - Hello World (
hello_world.py) teachesArtifactStep,lower(), andStepRunnerbasics with a two-step data pipeline. - Tiny Model Training (
train_tiny_model.py) demonstrates end-to-end LLM workflows usingtokenized()andtrain_lm(). - Advanced tutorials include hyperparameter sweeps (
train_tiny_sweep.py) and full-scale DCLM reproduction (exp1078_reproduce_dclm_7b1x.py). - All examples use lazy evaluation via
lower()and explicit dependency declarations throughdepsparameters or artifact references.
Frequently Asked Questions
Where are the Marin example projects located?
All example projects reside in the experiments/tutorials/ directory at the repository root. The primary entry points are hello_world.py, train_tiny_model.py, train_tiny_sweep.py, and exp1078_reproduce_dclm_7b1x.py, each demonstrating different aspects of the framework's lazy-artifact system.
What dependencies are required to run the Marin tutorials?
The examples require only the base Marin installation as documented in the repository's installation guide. Specific examples like the Tiny Model Training may pull additional dependencies (such as levanter for optimization configs or fray for resource management) automatically when Marin is installed with recommended extras.
How do the Marin examples handle artifact dependencies?
Dependencies are declared explicitly using the deps parameter in ArtifactStep constructors or implicitly by passing tokenized artifacts to training functions. Marin uses ctx.artifact_path() to resolve these dependencies at runtime, ensuring that upstream steps complete before downstream steps execute.
Can I modify the example configurations for my own datasets?
Yes. Each example uses configuration objects (like GenerateDataConfig or AdamConfig) that accept custom parameters. You can replace the TinyStories dataset reference in train_tiny_model.py with any HuggingFace dataset identifier, or adjust the sample_count, batch_size, and num_train_steps parameters to suit your computational constraints.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →