How to Set Up Marin for Local Development: Complete Installation Guide
Install Python 3.12+, clone the marin-community/marin repository, create a uv virtual environment, sync dependencies with hardware-specific extras, set three environment variables, and run the tiny-model tutorial to verify.
Marin is a modular research platform for building and evaluating large language models built around a lazy-artifact execution engine. Whether you are extending the StepRunner scheduler or adding new dataset readers, setting up a local development environment requires installing the Python stack, optionally building Rust wheels, and configuring runtime credentials. This guide walks you through the exact steps derived from the marin-community/marin source code to get from zero to running your first experiment.
Prerequisites
Before cloning the repository, ensure your system meets the following requirements:
- Python 3.12 or newer — Marin relies on modern Python features; install via
sudo apt install python3.12or usepyenv. - uv — A fast Python package manager and resolver; install with
pip install uv. - Git — Required for cloning submodules; install with
sudo apt install git. - Rust toolchain (optional) — Only needed if you plan to build native wheels from source instead of using pre-built PyPI artifacts. Install with
curl https://sh.rustup.rs -sSf | shand set the specific version used by the project:rustup toolchain install 1.91.0 && rustup default 1.91.0.
On macOS, you will also need cmake, pkg-config, and coreutils to build SentencePiece tokenizers.
Clone and Install
Clone the Repository
Fetch the full source including submodules for Levanter, Iris, and Zephyr:
git clone https://github.com/marin-community/marin.git
cd marin
Create a Virtual Environment
Use uv to create an isolated Python 3.12 environment:
uv venv --python 3.12
source .venv/bin/activate
On Windows, activate with .venv\Scripts\activate instead.
Install Packages and Dependencies
Marin uses a monorepo structure with multiple lib/* packages. Install everything in editable mode with the appropriate hardware extra:
-
CPU-only (default):
uv sync --all-packages -
GPU with CUDA 13:
uv sync --all-packages --extra=gpu -
TPU:
uv sync --all-packages --extra=tpu
The --all-packages flag ensures that lib/marin, lib/levanter, lib/iris, lib/zephyr, and utility packages are all installed in development mode.
Build Rust Crates from Source (Optional)
By default, pre-built wheels are fetched from PyPI. To compile native packages locally—which is useful when modifying the Rust source—run:
make rust-dev
uv sync --all-packages
This command modifies pyproject.toml to point to local Cargo builds. Do not commit these modified files; revert to production wheels later with make rust-user.
Configure Environment Variables
Marin requires three runtime variables to function correctly. Set them in your shell or a local .env file:
WANDB_API_KEY— Enables Weights & Biases experiment tracking (optional but recommended).HF_TOKEN— Grants access to gated Hugging Face models and tokenizers.MARIN_PREFIX— Defines the root directory for all artifacts. This can be a local path or anyfsspecURI such asgs://my-bucket.
Example configuration:
export WANDB_API_KEY=your_key_here
export HF_TOKEN=your_token_here
export MARIN_PREFIX=$HOME/marin_store
Source the file with source .env before running experiments.
Verify Your Installation
Confirm that the lazy-execution engine works end-to-end by running the official tiny-model tutorial. This executes a small language model training run on CPU using the StepRunner scheduler defined in lib/marin/src/marin/execution/step_runner.py.
Disable WandB logging if desired, then execute:
wandb offline
uv run python experiments/tutorials/train_tiny_model.py \
--device cpu \
--dataset tinystories \
--version dev \
--run
Successful output ends with:
INFO step_runner.py -- All steps complete.
You should also see two new directories under ${MARIN_PREFIX}:
${MARIN_PREFIX}/tokenized/tinystories/<date>/— Contains the cached tokenized dataset.${MARIN_PREFIX}/checkpoints/marin-nano-tinystories/<date>/— Contains model checkpoints.
Example: Minimal Experiment Script
Once verified, you can write custom experiments. The following script mirrors the tutorial structure and demonstrates the lazy-artifact pattern with lower() and StepRunner:
# my_experiment.py
from marin.execution.lazy import lower
from marin.execution.step_runner import StepRunner
from marin.experiment.data import tokenized
from marin.experiment.train import train_lm
from levanter.models.llama import LlamaConfig
from levanter.optim.config import AdamConfig
from fray.cluster import ResourceConfig
# 1. Create lazy handle for tokenized data (no I/O yet)
tinystories = tokenized(
name="tokenized/tinystories",
source="roneneldan/TinyStories",
tokenizer=marin_tokenizer,
version="2026.06.28",
sample_count=1000,
)
# 2. Define a tiny Llama architecture
llama_nano = LlamaConfig(
max_seq_len=2048,
hidden_dim=128,
intermediate_dim=512,
num_heads=4,
num_kv_heads=4,
num_layers=2,
)
# 3. Assemble training step
def build():
return train_lm(
name="checkpoints/marin-nano-tinystories",
version="2026.06.28",
model=llama_nano,
optimizer=AdamConfig(learning_rate=6e-4, weight_decay=0.1),
datasets={tinystories: 1.0},
batch_size=4,
seq_len=2048,
num_train_steps=100,
resources=ResourceConfig.with_cpu(),
)
# 4. Lower the DAG and run
if __name__ == "__main__":
StepRunner().run([lower(build())])
Execute with:
MARIN_PREFIX=local_store uv run python my_experiment.py
This demonstrates the same pipeline used by the tutorial at docs/tutorials/first-experiment.md, constructing a StepSpec DAG and executing it through the StepRunner.
Summary
- Use Python 3.12+ and uv to create an isolated environment for the marin-community/marin monorepo.
- Sync with
--all-packagesand add--extra=gpuor--extra=tpufor hardware-specific JAX wheels. - Set three environment variables:
WANDB_API_KEY,HF_TOKEN, andMARIN_PREFIXto enable tracking, data access, and artifact storage. - Verify with
train_tiny_model.pyto ensure the StepRunner scheduler inlib/marin/src/marin/execution/step_runner.pyexecutes steps correctly. - Build Rust crates with
make rust-devonly when modifying native code; remember to revert tomake rust-userbefore committing.
Frequently Asked Questions
What is the minimum Python version required for Marin?
Marin requires Python 3.12 or newer. The project uses modern Python features and type hints that are not compatible with older versions. Use pyenv or uv to install Python 3.12 if your system default is older.
Do I need to install Rust to use Marin?
No. Pre-built wheels are available on PyPI for standard development. You only need the Rust toolchain when modifying the native Rust crates or debugging performance-critical components. Run make rust-dev to switch to source builds and make rust-user to revert to wheels.
How do I switch between CPU, GPU, and TPU configurations?
Use the uv sync command with the appropriate extra flag. For CPU-only (default), run uv sync --all-packages. For NVIDIA GPUs with CUDA 13, add --extra=gpu. For TPU pods, add --extra=tpu. These flags pull the correct JAX wheel and related dependencies.
Where does Marin store artifacts and checkpoints?
Marin stores all artifacts under the path specified by the MARIN_PREFIX environment variable. This can be a local filesystem path like $HOME/marin_store or a cloud URI like gs://my-bucket. The StepRunner uses this prefix to cache tokenized datasets, model checkpoints, and intermediate StepSpec outputs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →