# How to Contribute to the Marin Project: A Complete Developer’s Guide

> Learn how to contribute to the Marin project. Follow our developer's guide to set up your environment, understand architecture rules, and submit successful pull requests for marin-community/marin.

- Repository: [The Marin Project/marin](https://github.com/marin-community/marin)
- Tags: how-to-guide
- Published: 2026-08-27

---

**To contribute to the Marin project, you must set up a Python 3.12 environment with `uv`, install dependencies via `uv sync --package marin-core --group dev`, follow the layered architecture rules (lower-level libraries cannot import from higher-level ones), and submit pull requests that pass `./infra/pre-commit.py --all-files --fix` and the full test suite.**

Marin is a modular research platform for large-language-model development organized into distinct architectural layers. Whether you are fixing bugs, adding optimizers to the training library, or extending dataset processing capabilities, understanding the repository structure and development workflow ensures your contribution aligns with the project’s standards.

## Understanding Marin’s Layered Architecture

The codebase is organized into four primary layers that enforce strict dependency directionality. **Lower-level libraries cannot import from higher-level ones**, ensuring clean separation of concerns.

- **Levanter**: The JAX-based training library located in `lib/levanter/`
- **Zephyr**: Dataset processing utilities in `lib/zephyr/`
- **Iris**: Job orchestration components in `lib/iris/`
- **Marin**: The top-level pipeline that coordinates the lower layers

When you contribute to the Marin project, place new functionality in the appropriate sub-package. For example, a new optimizer belongs in `lib/levanter/optim/`, while a dataset processing tool belongs in `lib/zephyr/`.

## Setting Up Your Development Environment

Before writing code, configure your local environment to match the repository’s requirements exactly.

### Prerequisites and Installation

Marin requires **Python 3.12** and uses `uv` for dependency management. Clone the repository and initialize your environment:

```bash
git clone https://github.com/marin-community/marin.git
cd marin
uv venv --python 3.12
source .venv/bin/activate

```

Install the core package along with development dependencies:

```bash
uv sync --package marin-core --group dev

```

### Running Pre-Commit Hooks

All contributions must pass the project-wide linting and formatting checks enforced by [`./infra/pre-commit.py`](https://github.com/marin-community/marin/blob/main/./infra/pre-commit.py). Run this command before every commit:

```bash
./infra/pre-commit.py --all-files --fix

```

This script executes the same checks used in CI, ensuring consistent code style across the repository.

## The Contribution Workflow

Follow these sequential steps when preparing your contribution to the Marin project.

### Reading the Contribution Guidelines

Start with the root [`CONTRIBUTING.md`](https://github.com/marin-community/marin/blob/main/CONTRIBUTING.md) file, which redirects to the comprehensive guide at [`docs/dev-guide/contributing.md`](https://github.com/marin-community/marin/blob/main/docs/dev-guide/contributing.md). This document specifies the detailed workflow, coding standards, and pull request requirements.

### Coding Standards and Type Checking

All public APIs in Marin are strictly typed. The repository uses **pyrefly** for static type checking, and you must include type hints for any new functions or classes. When adding features to libraries like `levanter`, ensure your code follows the existing patterns in `lib/levanter/` and does not introduce circular dependencies with higher-level packages.

### Testing Your Changes

Marin maintains a comprehensive test suite in the `tests/` directory. Run unit tests for specific components:

```bash
uv run pytest tests/levanter/test_example.py

```

For full validation before submitting, execute the safe-test suite:

```bash
uv run --no-project infra/ci/run_tests.py

```

### Documentation Requirements

Every code change requires corresponding documentation updates. If you modify Markdown sources in `docs/`, verify link integrity:

```bash
uv run python infra/check_docs_source_links.py

```

Optionally build the documentation locally to catch build errors:

```bash
uv run mkdocs build --strict

```

## Working with Agent Skills

Marin’s workflow utilizes **agents** defined in `.agents/skills/` to standardize common tasks. These reusable skills encapsulate higher-level actions such as adding datasets or configuring experiments. When contributing new functionality, check whether an existing skill in `.agents/skills/` already handles your use case rather than implementing ad-hoc scripts.

## Example: Adding a New Optimizer to Levanter

Here is a complete example of contributing a minimal change to the `levanter` library. This adds a new optimizer configuration:

```python

# File: lib/levanter/optim/adamx.py

from dataclasses import dataclass
from levanter.optim import OptimizerConfig

@dataclass
class AdamXConfig(OptimizerConfig):
    """AdamX with a custom decay schedule."""
    learning_rate: float = 1e-3
    beta1: float = 0.9
    beta2: float = 0.999
    weight_decay: float = 0.01

```

After creating the file, you must:

1. Add a unit test in [`tests/levanter/optim/test_adamx.py`](https://github.com/marin-community/marin/blob/main/tests/levanter/optim/test_adamx.py)
2. Update [`docs/levanter/optimizers.md`](https://github.com/marin-community/marin/blob/main/docs/levanter/optimizers.md) with documentation for the new class
3. Run `./infra/pre-commit.py --all-files --fix` to validate formatting

## Summary

Contributing to the Marin project requires understanding its strict architectural boundaries and automated quality checks. Key takeaways include:

- **Respect the layer hierarchy**: Code in `lib/levanter/` cannot import from `lib/marin/` or `lib/iris/`
- **Use `uv` for all dependency management** and always include the `--group dev` flag
- **Run `./infra/pre-commit.py --all-files --fix`** before committing to ensure CI compliance
- **Include type hints** for all public APIs and run `pyrefly` checks
- **Write tests** in the corresponding `tests/` subdirectory for every new feature
- **Update documentation** in `docs/` and verify links with [`infra/check_docs_source_links.py`](https://github.com/marin-community/marin/blob/main/infra/check_docs_source_links.py)

## Frequently Asked Questions

### What Python version is required to contribute to the Marin project?

Marin requires **Python 3.12** exactly. The [`pyproject.toml`](https://github.com/marin-community/marin/blob/main/pyproject.toml) file specifies this version constraint, and the development workflow assumes `uv` is installed to manage the virtual environment and dependencies.

### How do I run tests before submitting a contribution?

Run targeted unit tests with `uv run pytest <path-to-test-file>`. For comprehensive validation that mimics CI, execute `uv run --no-project infra/ci/run_tests.py`, which runs the full safe-test suite across all libraries including `levanter`, `zephyr`, and `iris`.

### Can I add functionality that imports from higher-level libraries?

No. The architecture strictly prohibits lower-level libraries from importing higher-level ones. For example, code in `lib/levanter/` cannot import from `lib/marin/` or `lib/iris/`. Always place your contribution in the appropriate layer based on its abstraction level.

### Where should I place new agent skills when contributing?

Place reusable agent skills in the `.agents/skills/` directory. Before creating new skills, check if an existing skill already provides the functionality you need. Skills standardize common tasks like `add-dataset` and ensure consistency across the Marin project’s automated workflows.