# Prerequisites for Running Marin: From Local CPU to Distributed GPU Clusters

> Discover the essential prerequisites for running Marin, from local CPU setups to distributed GPU clusters. Learn about Python 3.12+, uv, Git, and hardware requirements.

- Repository: [The Marin Project/marin](https://github.com/marin-community/marin)
- Tags: getting-started
- Published: 2026-08-27

---

**To run Marin, you need Python 3.12+, the uv package manager, Git, and either CPU resources or hardware-specific JAX wheels for GPU/TPU, plus optional environment variables for Weights & Biases and Hugging Face authentication.**

Marin is a modular pipeline framework for training large language models maintained by the `marin-community` organization. Whether you are prototyping on a laptop or scheduling jobs across a GPU fleet, satisfying the correct prerequisites ensures the JAX-based runtime initializes properly. This guide covers the exact system requirements, hardware dependencies, and environment variables defined in the source repository.

## Core System Requirements

All Marin installations share a baseline set of tooling regardless of compute target.

- **Python 3.12 or newer**: The codebase leverages modern language features and strict type-checking introduced in recent Python versions, as declared in [`pyproject.toml`](https://github.com/marin-community/marin/blob/main/pyproject.toml).
- **uv**: This fast Python package manager handles dependency resolution and virtual-environment creation. The installation workflow documented in [`docs/tutorials/installation.md`](https://github.com/marin-community/marin/blob/main/docs/tutorials/installation.md) uses `uv sync` exclusively.
- **Git**: Required to clone the `marin-community/marin` repository and manage version-controlled configurations.
- **Rust toolchain (rustup)**: Optional unless you intend to build the native Rust wheels from source. The default workflow pulls pre-built wheels, but `Makefile` targets like `make rust-dev` require a local Rust compiler.

These tools enable you to fetch the repository, create an isolated Python environment, and install the mixed Python/Rust packages that constitute Marin’s core.

## Hardware-Specific Dependencies

Marin’s compute backend relies on JAX. The specific variant installed determines whether you train on CPU, GPU, or TPU.

### CPU-Only Setup

The simplest path requires no additional system packages beyond the core requirements. Running `uv sync --all-packages` pulls the default CPU-only JAX wheel, sufficient for development and small-scale experiments.

### GPU Requirements

For NVIDIA hardware, you must satisfy strict driver and toolkit versions:

- **NVIDIA driver ≥ 580** (CUDA 13 compatible)
- **CUDA Toolkit 13** (only if compiling custom kernels; the JAX wheel bundles its own runtime)

Install the GPU variant by adding the extra:

```bash
uv sync --all-packages --extra=gpu

```

This pulls the JAX-CUDA wheel containing CUDA, cuDNN, and NCCL Python packages. See [`docs/tutorials/local-gpu.md`](https://github.com/marin-community/marin/blob/main/docs/tutorials/local-gpu.md) for driver installation steps.

### TPU Setup

TPU training requires the TPU-specific JAX wheel installed via the `tpu` extra:

```bash
uv sync --all-packages --extra=tpu

```

No additional system drivers are required when running on Google Cloud TPU VMs, as the JAX wheel bundles necessary libtpu dependencies.

## Runtime Configuration and Environment Variables

Before executing training scripts, export these variables (typically in a `.env` file):

- **WANDB_API_KEY**: Authenticates with Weights & Biases for experiment tracking. Optional but recommended; referenced in [`docs/tutorials/installation.md`](https://github.com/marin-community/marin/blob/main/docs/tutorials/installation.md) under "Setup Weights and Biases".
- **HF_TOKEN**: Provides access to gated models and tokenizers on Hugging Face. Required when using restricted checkpoints.
- **MARIN_PREFIX**: Defines the storage root for checkpoints, logs, and artifacts. Accepts local paths or any `fsspec`-compatible URL (e.g., `s3://bucket/path`). This variable is parsed in [`config/marin.yaml`](https://github.com/marin-community/marin/blob/main/config/marin.yaml).

## Optional Cluster-Level Prerequisites (Iris)

When scaling beyond a single machine, Marin integrates with Iris, its cluster scheduling layer. Deploying to an Iris-managed Kubernetes cluster requires:

- A running **Iris** installation with configured **RBAC**, **NodePools**, **Kueue** for quota management, **Ingress** controllers, and an **object-store bucket** for artifact persistence.
- The **Iris CLI** (`iris`) installed via `uv install iris`.

The Iris controller validates these prerequisites when you run `iris cluster start`. Full contract details are documented in [`lib/iris/OPS.md`](https://github.com/marin-community/marin/blob/main/lib/iris/OPS.md).

## Step-by-Step Installation Guide

The following script demonstrates a complete CPU-based installation:

```bash

# Clone the repository

git clone https://github.com/marin-community/marin.git
cd marin

# Create and activate Python 3.12 environment

uv venv --python 3.12
source .venv/bin/activate  # Windows: .venv\Scripts\activate

# Install CPU dependencies

uv sync --all-packages

# Configure environment variables

export WANDB_API_KEY="your_key_here"
export HF_TOKEN="your_hf_token_here"
export MARIN_PREFIX="$HOME/marin_artifacts"

# Run a verification experiment

uv run python experiments/tutorials/train_tiny_model.py \
    --device cpu \
    --dataset tinystories \
    --version dev \
    --run

```

To switch to GPU, install the extra and change the device flag:

```bash
uv sync --all-packages --extra=gpu
uv run python experiments/tutorials/train_tiny_model.py \
    --device h100x8 \
    --dataset wikitext \
    --version dev \
    --run

```

For distributed execution on an Iris cluster:

```bash
uv run iris --cluster=marin job run \
    --device h100x8 \
    --script experiments/tutorials/train_tiny_model.py \
    --args "--dataset wikitext --version dev --run"

```

## Summary

- **Python 3.12+**, **uv**, and **Git** are mandatory for all installations.
- **GPU** training requires NVIDIA driver ≥580 and the `gpu` extra; **TPU** requires the `tpu` extra.
- **Environment variables** `WANDB_API_KEY`, `HF_TOKEN`, and `MARIN_PREFIX` configure external integrations and artifact storage.
- **Rust** is only needed for building native wheels from source, not for standard usage.
- **Iris clusters** demand Kubernetes infrastructure including Kueue, RBAC, and object storage for multi-node scheduling.

## Frequently Asked Questions

### Do I need a GPU to run Marin?

No. Marin runs on CPU-only systems using the default JAX wheel, suitable for development and small experiments. GPU is only required for large-scale training throughput.

### What Python version does Marin require?

Marin requires **Python 3.12 or newer**. The codebase uses modern typing features and syntax only available in 3.12+, enforced in the project's [`pyproject.toml`](https://github.com/marin-community/marin/blob/main/pyproject.toml).

### Is the Rust toolchain mandatory for installation?

No. Rust is only required if you plan to compile Marin’s native extensions from source using `make rust-dev` or `make rust-user`. Pre-built wheels are available for standard installations.

### How do I run Marin on a shared Kubernetes cluster?

Install the Iris CLI and ensure your cluster meets the prerequisites documented in [`lib/iris/OPS.md`](https://github.com/marin-community/marin/blob/main/lib/iris/OPS.md): RBAC policies, NodePools, Kueue for job queuing, and an object-store bucket accessible via the `MARIN_PREFIX` environment variable.