How to Contribute to the Marin Project: A Complete Developer’s Guide
To contribute to the Marin project, clone the repository, set up a Python 3.12 environment with uv, install dependencies via uv sync --package marin-core --group dev, enforce code quality with ./infra/pre-commit.py --all-files --fix, develop your feature in the correct architectural layer, add comprehensive tests, verify documentation integrity, and open a pull request that references related issues using Fixes #NNNN.
The Marin project is a modular research platform for large-language-model development maintained at marin-community/marin on GitHub. Its codebase is strictly organized into hierarchical libraries—levanter, zephyr, iris, and marin—each with specific responsibilities and dependency rules. Understanding this architecture and the standardized development workflow ensures your contributions align with the project’s quality standards and merge cleanly into the main branch.
Understanding Marin’s Layered Architecture
Marin’s codebase follows a strict layered design where lower-level libraries provide foundational functionality that higher-level ones consume. Violating these boundaries—such as importing from marin inside levanter—will cause CI failures.
The Four Architectural Layers
levanter– The foundational JAX-based training library containing optimizers, model implementations, and distributed training logic.zephyr– Dataset processing and data-loading utilities that prepare training corpora.iris– Job orchestration and cluster management components.marin– The top-level pipeline that orchestrates end-to-end experiments, importing from the lower layers.
Dependency Direction Rules
According to the source code in docs/dev-guide/contributing.md, the cardinal rule is that lower-level libraries cannot import from higher-level ones. For example, code in lib/levanter/ may not import from lib/marin/, but lib/marin/ can freely import from lib/levanter/, lib/zephyr/, and lib/iris/. Always place new functionality in the lowest appropriate layer to maximize reusability.
Setting Up Your Development Environment
Marin requires Python 3.12 and uses uv for fast dependency management. Do not use pip directly.
First, clone the repository and create an isolated environment:
git clone https://github.com/marin-community/marin.git
cd marin
uv venv --python 3.12
source .venv/bin/activate
Next, install the core package along with development dependencies:
uv sync --package marin-core --group dev
This command installs all necessary tools—including pytest for testing, pyrefly for static type checking, and mkdocs for documentation generation—defined in pyproject.toml.
The Contribution Workflow
The canonical contribution process is documented in CONTRIBUTING.md at the repository root, which redirects to the detailed guide at docs/dev-guide/contributing.md. Follow these steps sequentially to ensure CI passes on your first submission.
1. Run the Pre-Commit Hook Locally
Before writing code, execute the project-wide linting and formatting script:
./infra/pre-commit.py --all-files --fix
This script enforces style checks across the entire repository and is identical to the checks run in CI. Running it locally prevents "fix lint" commits.
2. Develop in the Correct Sub-Package
Add your code to the appropriate library layer. For example, to add a new optimizer to levanter, create a file like lib/levanter/optim/adamx.py:
# File: lib/levanter/optim/adamx.py
from dataclasses import dataclass
from levanter.optim import OptimizerConfig
@dataclass
class AdamXConfig(OptimizerConfig):
"""AdamX with a custom decay schedule."""
learning_rate: float = 1e-3
beta1: float = 0.9
beta2: float = 0.999
weight_decay: float = 0.01
All public APIs must include type annotations, as the repository uses pyrefly for static analysis.
3. Add Tests
Every new feature requires unit tests. Place test files mirroring the source structure under tests/:
# Example: testing the new optimizer
uv run pytest tests/levanter/optim/test_adamx.py
For comprehensive validation before pushing, run the full safe-test suite:
uv run --no-project infra/ci/run_tests.py
4. Verify Documentation Integrity
If you modified Markdown files or added new modules, validate that no internal links are broken:
uv run python infra/check_docs_source_links.py
Optionally build the documentation site locally to catch rendering errors:
uv run mkdocs build --strict
Keep documentation in the docs/ directory synchronized with code changes, especially when adding public classes or functions.
5. Submit Your Pull Request
After local validation passes, push your branch and open a pull request. Follow these conventions from the contribution guide:
- Reference the issue number with
Fixes #NNNNin the PR description if applicable. - Follow the writing-style skill for PR titles and bodies (see
.agents/skills/for templates). - Ensure the pre-commit hook passes; CI will block merging if
./infra/pre-commit.py --all-filesreturns errors.
Working with the Agent Skills Framework
Marin’s workflow is heavily driven by agents—automated tools defined in .agents/skills/. These encapsulate common tasks like adding datasets or configuring experiments. Before writing ad-hoc scripts in the marin layer, check if an existing skill (e.g., add-dataset) can handle your use case. Reusing skills maintains consistency across the research platform and reduces maintenance burden.
Summary
- Marin uses a strict four-layer architecture (
levanter→zephyr→iris→marin) where dependencies flow downward only. - Setup requires Python 3.12 and
uv; install withuv sync --package marin-core --group dev. - Quality gates include
./infra/pre-commit.py --all-files --fix,pyreflytype checking, and comprehensive pytest coverage. - Documentation must stay synchronized; use
infra/check_docs_source_links.pyto validate links. - Leverage agent skills in
.agents/skills/before writing new automation scripts.
Frequently Asked Questions
What Python version is required to contribute to Marin?
Marin requires Python 3.12 exactly. The pyproject.toml specifies this version constraint, and using older versions will result in dependency resolution failures when running uv sync.
How do I run only the tests relevant to my specific changes?
For rapid iteration during development, run targeted tests using uv run pytest <path>. For example, uv run pytest tests/levanter/optim/ tests only the optimizer module. Run the full safe-test suite with uv run --no-project infra/ci/run_tests.py before final submission.
Can I add utility scripts directly to the top-level marin pipeline?
While possible, the project prefers encapsulating common actions as reusable agent skills located in .agents/skills/. If your task resembles existing operations like "add dataset" or "run evaluation," extend or reuse the relevant skill rather than creating standalone scripts in the marin package.
What should I do if the pre-commit hook reports errors?
Run ./infra/pre-commit.py --all-files --fix from the repository root. This command auto-fixes most formatting issues (import sorting, code style). If errors persist after the fix phase, they typically indicate type-checking failures from pyrefly or structural issues that must be resolved manually before committing.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →