How to Contribute to the Marin Project: A Complete Developer’s Guide

To contribute to the Marin project, clone the repository, set up a Python 3.12 environment with uv, install dependencies via uv sync --package marin-core --group dev, enforce code quality with ./infra/pre-commit.py --all-files --fix, develop your feature in the correct architectural layer, add comprehensive tests, verify documentation integrity, and open a pull request that references related issues using Fixes #NNNN.

The Marin project is a modular research platform for large-language-model development maintained at marin-community/marin on GitHub. Its codebase is strictly organized into hierarchical libraries—levanter, zephyr, iris, and marin—each with specific responsibilities and dependency rules. Understanding this architecture and the standardized development workflow ensures your contributions align with the project’s quality standards and merge cleanly into the main branch.

Understanding Marin’s Layered Architecture

Marin’s codebase follows a strict layered design where lower-level libraries provide foundational functionality that higher-level ones consume. Violating these boundaries—such as importing from marin inside levanter—will cause CI failures.

The Four Architectural Layers

  • levanter – The foundational JAX-based training library containing optimizers, model implementations, and distributed training logic.
  • zephyr – Dataset processing and data-loading utilities that prepare training corpora.
  • iris – Job orchestration and cluster management components.
  • marin – The top-level pipeline that orchestrates end-to-end experiments, importing from the lower layers.

Dependency Direction Rules

According to the source code in docs/dev-guide/contributing.md, the cardinal rule is that lower-level libraries cannot import from higher-level ones. For example, code in lib/levanter/ may not import from lib/marin/, but lib/marin/ can freely import from lib/levanter/, lib/zephyr/, and lib/iris/. Always place new functionality in the lowest appropriate layer to maximize reusability.

Setting Up Your Development Environment

Marin requires Python 3.12 and uses uv for fast dependency management. Do not use pip directly.

First, clone the repository and create an isolated environment:

git clone https://github.com/marin-community/marin.git
cd marin
uv venv --python 3.12
source .venv/bin/activate

Next, install the core package along with development dependencies:

uv sync --package marin-core --group dev

This command installs all necessary tools—including pytest for testing, pyrefly for static type checking, and mkdocs for documentation generation—defined in pyproject.toml.

The Contribution Workflow

The canonical contribution process is documented in CONTRIBUTING.md at the repository root, which redirects to the detailed guide at docs/dev-guide/contributing.md. Follow these steps sequentially to ensure CI passes on your first submission.

1. Run the Pre-Commit Hook Locally

Before writing code, execute the project-wide linting and formatting script:

./infra/pre-commit.py --all-files --fix

This script enforces style checks across the entire repository and is identical to the checks run in CI. Running it locally prevents "fix lint" commits.

2. Develop in the Correct Sub-Package

Add your code to the appropriate library layer. For example, to add a new optimizer to levanter, create a file like lib/levanter/optim/adamx.py:


# File: lib/levanter/optim/adamx.py

from dataclasses import dataclass
from levanter.optim import OptimizerConfig

@dataclass
class AdamXConfig(OptimizerConfig):
    """AdamX with a custom decay schedule."""
    learning_rate: float = 1e-3
    beta1: float = 0.9
    beta2: float = 0.999
    weight_decay: float = 0.01

All public APIs must include type annotations, as the repository uses pyrefly for static analysis.

3. Add Tests

Every new feature requires unit tests. Place test files mirroring the source structure under tests/:


# Example: testing the new optimizer

uv run pytest tests/levanter/optim/test_adamx.py

For comprehensive validation before pushing, run the full safe-test suite:

uv run --no-project infra/ci/run_tests.py

4. Verify Documentation Integrity

If you modified Markdown files or added new modules, validate that no internal links are broken:

uv run python infra/check_docs_source_links.py

Optionally build the documentation site locally to catch rendering errors:

uv run mkdocs build --strict

Keep documentation in the docs/ directory synchronized with code changes, especially when adding public classes or functions.

5. Submit Your Pull Request

After local validation passes, push your branch and open a pull request. Follow these conventions from the contribution guide:

  • Reference the issue number with Fixes #NNNN in the PR description if applicable.
  • Follow the writing-style skill for PR titles and bodies (see .agents/skills/ for templates).
  • Ensure the pre-commit hook passes; CI will block merging if ./infra/pre-commit.py --all-files returns errors.

Working with the Agent Skills Framework

Marin’s workflow is heavily driven by agents—automated tools defined in .agents/skills/. These encapsulate common tasks like adding datasets or configuring experiments. Before writing ad-hoc scripts in the marin layer, check if an existing skill (e.g., add-dataset) can handle your use case. Reusing skills maintains consistency across the research platform and reduces maintenance burden.

Summary

  • Marin uses a strict four-layer architecture (levanter → zephyr → iris → marin) where dependencies flow downward only.
  • Setup requires Python 3.12 and uv; install with uv sync --package marin-core --group dev.
  • Quality gates include ./infra/pre-commit.py --all-files --fix, pyrefly type checking, and comprehensive pytest coverage.
  • Documentation must stay synchronized; use infra/check_docs_source_links.py to validate links.
  • Leverage agent skills in .agents/skills/ before writing new automation scripts.

Frequently Asked Questions

What Python version is required to contribute to Marin?

Marin requires Python 3.12 exactly. The pyproject.toml specifies this version constraint, and using older versions will result in dependency resolution failures when running uv sync.

How do I run only the tests relevant to my specific changes?

For rapid iteration during development, run targeted tests using uv run pytest <path>. For example, uv run pytest tests/levanter/optim/ tests only the optimizer module. Run the full safe-test suite with uv run --no-project infra/ci/run_tests.py before final submission.

Can I add utility scripts directly to the top-level marin pipeline?

While possible, the project prefers encapsulating common actions as reusable agent skills located in .agents/skills/. If your task resembles existing operations like "add dataset" or "run evaluation," extend or reuse the relevant skill rather than creating standalone scripts in the marin package.

What should I do if the pre-commit hook reports errors?

Run ./infra/pre-commit.py --all-files --fix from the repository root. This command auto-fixes most formatting issues (import sorting, code style). If errors persist after the fix phase, they typically indicate type-checking failures from pyrefly or structural issues that must be resolved manually before committing.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →