How to Set Up a Local Development Environment for Marin
Setting up a local development environment for Marin requires Python 3.12+, the uv package manager, and Git, followed by cloning the repository, creating a virtual environment, installing hardware-specific dependencies with uv sync, and configuring optional environment variables for Weights & Biases and Hugging Face.
Marin is an open-source machine learning framework maintained by the marin-community organization. Setting up a local development environment for Marin involves using modern Python tooling to manage dependencies and hardware-specific JAX builds for CPU, GPU, or TPU acceleration. This guide walks through the complete installation process based on the canonical documentation in docs/tutorials/installation.md.
Prerequisites
Before installing Marin, ensure your system meets the baseline requirements. You will need Python 3.12+, the uv package manager, and Git. macOS users may need additional build tools for SentencePiece compilation. Optionally, install the Rust toolchain if you plan to build native crates from source instead of using pre-built wheels.
Clone the Marin Repository
Start by cloning the repository from GitHub. This downloads the source code, configuration files defined in config/marin.yaml, and tutorial scripts to your local machine.
git clone https://github.com/marin-community/marin.git
cd marin
Create and Activate a Virtual Environment
Marin uses uv to manage Python environments and dependencies. Create an isolated virtual environment specifying Python 3.12, then activate it.
uv venv --python 3.12
source .venv/bin/activate # On Windows: .venv\Scripts\activate
Install Marin with Hardware-Specific Dependencies
The installation command varies based on your hardware acceleration needs. Marin defines hardware-specific extras in pyproject.toml to install the correct JAX build for your system.
CPU Installation
For development on machines without dedicated accelerators, install the CPU variant:
uv sync --all-packages --extra=cpu
GPU Installation
For NVIDIA GPU support, first install CUDA 12.9 and cuDNN according to the official documentation. Then install Marin with the GPU extra:
uv sync --all-packages --extra=gpu
TPU Installation
For Google Cloud TPU environments, use the TPU-specific extra:
uv sync --all-packages --extra=tpu
The uv sync command resolves all Python dependencies listed in pyproject.toml and installs the local package in editable mode, making the process take approximately 5–10 minutes due to source builds of some dependencies.
Configure Environment Variables
Marin relies on several environment variables for optional integrations and artifact storage. Create a .env file in the project root to store these values.
Set MARIN_PREFIX to define where checkpoints, datasets, and other artifacts are written. This path supports local directories or fsspec-compatible storage like GCS buckets.
Configure Weights & Biases for experiment tracking by setting WANDB_API_KEY, and optionally WANDB_ENTITY and WANDB_PROJECT. For accessing gated models on Hugging Face, set HF_TOKEN.
Example .env file:
export MARIN_PREFIX="$HOME/marin_store"
export WANDB_API_KEY="your-wandb-key"
export HF_TOKEN="your-huggingface-token"
Source the file to apply variables:
source .env
Optional: Build Rust Crates from Source
By default, Marin pulls pre-built native Rust wheels from PyPI. To compile the Rust crates from source instead, use the Makefile targets provided in the repository root.
make rust-dev # For development builds
make rust-user # For optimized user builds
This step requires the Rust toolchain to be installed and available in your PATH.
Verify Your Installation
Confirm your local development environment works by running the First Experiment tutorial. This script trains a tiny language model on the TinyStories dataset using only CPU resources.
First, optionally disable Weights & Biases logging:
wandb offline
Then execute the verification script located at experiments/tutorials/train_tiny_model.py:
uv run python experiments/tutorials/train_tiny_model.py \
--device cpu \
--dataset tinystories \
--version dev \
--run
This command tokenizes the dataset, trains a small model, and writes checkpoints to the location defined by MARIN_PREFIX. Successful completion confirms that the setup outlined in docs/tutorials/installation.md is functioning correctly.
Summary
- Install Python 3.12+, uv, and Git before starting, with optional Rust tools for building native crates.
- Clone the marin-community/marin repository and create a virtual environment with
uv venv --python 3.12. - Use
uv sync --all-packages --extra={cpu|gpu|tpu}to install hardware-specific dependencies defined inpyproject.toml. - Configure
MARIN_PREFIX,WANDB_API_KEY, andHF_TOKENin a.envfile for artifact storage and optional integrations. - Run
experiments/tutorials/train_tiny_model.pyto verify the installation trains a model successfully.
Frequently Asked Questions
What Python version is required for Marin?
Marin requires Python 3.12 or newer. The installation process explicitly uses uv venv --python 3.12 to ensure compatibility with the dependency specifications in pyproject.toml.
Do I need a GPU to run Marin locally?
No. Marin supports CPU-only development using the --extra=cpu flag during installation. The verification tutorial in experiments/tutorials/train_tiny_model.py runs entirely on CPU, making it accessible for development on laptops or workstations without dedicated accelerators.
How do I switch between CPU and GPU installations?
Re-run the installation command with the appropriate extra flag. For GPU support, first install CUDA 12.9 and cuDNN, then execute uv sync --all-packages --extra=gpu. To revert to CPU-only, use uv sync --all-packages --extra=cpu. The uv tool automatically manages the different JAX builds specified in pyproject.toml.
Where does Marin store checkpoints and datasets?
Marin writes artifacts to the path defined by the MARIN_PREFIX environment variable. According to docs/tutorials/installation.md, this variable points to a local directory or an fsspec-compatible remote storage location such as a Google Cloud Storage bucket. All checkpoints, datasets, and experimental outputs are organized under this prefix.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →