# How to Set Up Marin for Local Development: Complete Installation Guide

> Easily set up Marin for local development with this complete installation guide. Follow simple steps to clone the repository, install dependencies, and run the tiny-model tutorial.

- Repository: [The Marin Project/marin](https://github.com/marin-community/marin)
- Tags: getting-started
- Published: 2026-08-27

---

**Install Python 3.12+, clone the marin-community/marin repository, create a uv virtual environment, sync dependencies with hardware-specific extras, set three environment variables, and run the tiny-model tutorial to verify.**

Marin is a modular research platform for building and evaluating large language models built around a **lazy-artifact execution engine**. Whether you are extending the **StepRunner** scheduler or adding new dataset readers, setting up a local development environment requires installing the Python stack, optionally building Rust wheels, and configuring runtime credentials. This guide walks you through the exact steps derived from the marin-community/marin source code to get from zero to running your first experiment.

## Prerequisites

Before cloning the repository, ensure your system meets the following requirements:

- **Python 3.12 or newer** — Marin relies on modern Python features; install via `sudo apt install python3.12` or use `pyenv`.
- **uv** — A fast Python package manager and resolver; install with `pip install uv`.
- **Git** — Required for cloning submodules; install with `sudo apt install git`.
- **Rust toolchain** (optional) — Only needed if you plan to build native wheels from source instead of using pre-built PyPI artifacts. Install with `curl https://sh.rustup.rs -sSf | sh` and set the specific version used by the project: `rustup toolchain install 1.91.0 && rustup default 1.91.0`.

On macOS, you will also need `cmake`, `pkg-config`, and `coreutils` to build SentencePiece tokenizers.

## Clone and Install

### Clone the Repository

Fetch the full source including submodules for **Levanter**, **Iris**, and **Zephyr**:

```bash
git clone https://github.com/marin-community/marin.git
cd marin

```

### Create a Virtual Environment

Use uv to create an isolated Python 3.12 environment:

```bash
uv venv --python 3.12
source .venv/bin/activate

```

On Windows, activate with `.venv\Scripts\activate` instead.

### Install Packages and Dependencies

Marin uses a monorepo structure with multiple `lib/*` packages. Install everything in editable mode with the appropriate hardware extra:

- **CPU-only** (default):

  ```bash
  uv sync --all-packages
  ```

- **GPU with CUDA 13**:

  ```bash
  uv sync --all-packages --extra=gpu
  ```

- **TPU**:

  ```bash
  uv sync --all-packages --extra=tpu
  ```

The `--all-packages` flag ensures that `lib/marin`, `lib/levanter`, `lib/iris`, `lib/zephyr`, and utility packages are all installed in development mode.

### Build Rust Crates from Source (Optional)

By default, pre-built wheels are fetched from PyPI. To compile native packages locally—which is useful when modifying the Rust source—run:

```bash
make rust-dev
uv sync --all-packages

```

This command modifies [`pyproject.toml`](https://github.com/marin-community/marin/blob/main/pyproject.toml) to point to local Cargo builds. **Do not commit** these modified files; revert to production wheels later with `make rust-user`.

## Configure Environment Variables

Marin requires three runtime variables to function correctly. Set them in your shell or a local `.env` file:

- **`WANDB_API_KEY`** — Enables Weights & Biases experiment tracking (optional but recommended).
- **`HF_TOKEN`** — Grants access to gated Hugging Face models and tokenizers.
- **`MARIN_PREFIX`** — Defines the root directory for all artifacts. This can be a local path or any `fsspec` URI such as `gs://my-bucket`.

Example configuration:

```bash
export WANDB_API_KEY=your_key_here
export HF_TOKEN=your_token_here
export MARIN_PREFIX=$HOME/marin_store

```

Source the file with `source .env` before running experiments.

## Verify Your Installation

Confirm that the lazy-execution engine works end-to-end by running the official tiny-model tutorial. This executes a small language model training run on CPU using the **StepRunner** scheduler defined in [`lib/marin/src/marin/execution/step_runner.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/execution/step_runner.py).

Disable WandB logging if desired, then execute:

```bash
wandb offline
uv run python experiments/tutorials/train_tiny_model.py \
  --device cpu \
  --dataset tinystories \
  --version dev \
  --run

```

Successful output ends with:

```text
INFO step_runner.py -- All steps complete.

```

You should also see two new directories under `${MARIN_PREFIX}`:

- `${MARIN_PREFIX}/tokenized/tinystories/<date>/` — Contains the cached tokenized dataset.
- `${MARIN_PREFIX}/checkpoints/marin-nano-tinystories/<date>/` — Contains model checkpoints.

## Example: Minimal Experiment Script

Once verified, you can write custom experiments. The following script mirrors the tutorial structure and demonstrates the **lazy-artifact** pattern with `lower()` and `StepRunner`:

```python

# my_experiment.py

from marin.execution.lazy import lower
from marin.execution.step_runner import StepRunner
from marin.experiment.data import tokenized
from marin.experiment.train import train_lm
from levanter.models.llama import LlamaConfig
from levanter.optim.config import AdamConfig
from fray.cluster import ResourceConfig

# 1. Create lazy handle for tokenized data (no I/O yet)

tinystories = tokenized(
    name="tokenized/tinystories",
    source="roneneldan/TinyStories",
    tokenizer=marin_tokenizer,
    version="2026.06.28",
    sample_count=1000,
)

# 2. Define a tiny Llama architecture

llama_nano = LlamaConfig(
    max_seq_len=2048,
    hidden_dim=128,
    intermediate_dim=512,
    num_heads=4,
    num_kv_heads=4,
    num_layers=2,
)

# 3. Assemble training step

def build():
    return train_lm(
        name="checkpoints/marin-nano-tinystories",
        version="2026.06.28",
        model=llama_nano,
        optimizer=AdamConfig(learning_rate=6e-4, weight_decay=0.1),
        datasets={tinystories: 1.0},
        batch_size=4,
        seq_len=2048,
        num_train_steps=100,
        resources=ResourceConfig.with_cpu(),
    )

# 4. Lower the DAG and run

if __name__ == "__main__":
    StepRunner().run([lower(build())])

```

Execute with:

```bash
MARIN_PREFIX=local_store uv run python my_experiment.py

```

This demonstrates the same pipeline used by the tutorial at [`docs/tutorials/first-experiment.md`](https://github.com/marin-community/marin/blob/main/docs/tutorials/first-experiment.md), constructing a **StepSpec** DAG and executing it through the **StepRunner**.

## Summary

- **Use Python 3.12+ and uv** to create an isolated environment for the marin-community/marin monorepo.
- **Sync with `--all-packages`** and add `--extra=gpu` or `--extra=tpu` for hardware-specific JAX wheels.
- **Set three environment variables**: `WANDB_API_KEY`, `HF_TOKEN`, and `MARIN_PREFIX` to enable tracking, data access, and artifact storage.
- **Verify with [`train_tiny_model.py`](https://github.com/marin-community/marin/blob/main/train_tiny_model.py)** to ensure the **StepRunner** scheduler in [`lib/marin/src/marin/execution/step_runner.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/execution/step_runner.py) executes steps correctly.
- **Build Rust crates with `make rust-dev`** only when modifying native code; remember to revert to `make rust-user` before committing.

## Frequently Asked Questions

### What is the minimum Python version required for Marin?

Marin requires **Python 3.12 or newer**. The project uses modern Python features and type hints that are not compatible with older versions. Use `pyenv` or `uv` to install Python 3.12 if your system default is older.

### Do I need to install Rust to use Marin?

No. Pre-built wheels are available on PyPI for standard development. You only need the **Rust toolchain** when modifying the native Rust crates or debugging performance-critical components. Run `make rust-dev` to switch to source builds and `make rust-user` to revert to wheels.

### How do I switch between CPU, GPU, and TPU configurations?

Use the `uv sync` command with the appropriate extra flag. For CPU-only (default), run `uv sync --all-packages`. For NVIDIA GPUs with CUDA 13, add `--extra=gpu`. For TPU pods, add `--extra=tpu`. These flags pull the correct JAX wheel and related dependencies.

### Where does Marin store artifacts and checkpoints?

Marin stores all artifacts under the path specified by the **`MARIN_PREFIX`** environment variable. This can be a local filesystem path like `$HOME/marin_store` or a cloud URI like `gs://my-bucket`. The `StepRunner` uses this prefix to cache tokenized datasets, model checkpoints, and intermediate `StepSpec` outputs.