How to Contribute to the Marin Project: A Complete Developer’s Guide

To contribute to the Marin project, you must set up a Python 3.12 environment with uv, install dependencies via uv sync --package marin-core --group dev, follow the layered architecture rules (lower-level libraries cannot import from higher-level ones), and submit pull requests that pass ./infra/pre-commit.py --all-files --fix and the full test suite.

Marin is a modular research platform for large-language-model development organized into distinct architectural layers. Whether you are fixing bugs, adding optimizers to the training library, or extending dataset processing capabilities, understanding the repository structure and development workflow ensures your contribution aligns with the project’s standards.

Understanding Marin’s Layered Architecture

The codebase is organized into four primary layers that enforce strict dependency directionality. Lower-level libraries cannot import from higher-level ones, ensuring clean separation of concerns.

  • Levanter: The JAX-based training library located in lib/levanter/
  • Zephyr: Dataset processing utilities in lib/zephyr/
  • Iris: Job orchestration components in lib/iris/
  • Marin: The top-level pipeline that coordinates the lower layers

When you contribute to the Marin project, place new functionality in the appropriate sub-package. For example, a new optimizer belongs in lib/levanter/optim/, while a dataset processing tool belongs in lib/zephyr/.

Setting Up Your Development Environment

Before writing code, configure your local environment to match the repository’s requirements exactly.

Prerequisites and Installation

Marin requires Python 3.12 and uses uv for dependency management. Clone the repository and initialize your environment:

git clone https://github.com/marin-community/marin.git
cd marin
uv venv --python 3.12
source .venv/bin/activate

Install the core package along with development dependencies:

uv sync --package marin-core --group dev

Running Pre-Commit Hooks

All contributions must pass the project-wide linting and formatting checks enforced by ./infra/pre-commit.py. Run this command before every commit:

./infra/pre-commit.py --all-files --fix

This script executes the same checks used in CI, ensuring consistent code style across the repository.

The Contribution Workflow

Follow these sequential steps when preparing your contribution to the Marin project.

Reading the Contribution Guidelines

Start with the root CONTRIBUTING.md file, which redirects to the comprehensive guide at docs/dev-guide/contributing.md. This document specifies the detailed workflow, coding standards, and pull request requirements.

Coding Standards and Type Checking

All public APIs in Marin are strictly typed. The repository uses pyrefly for static type checking, and you must include type hints for any new functions or classes. When adding features to libraries like levanter, ensure your code follows the existing patterns in lib/levanter/ and does not introduce circular dependencies with higher-level packages.

Testing Your Changes

Marin maintains a comprehensive test suite in the tests/ directory. Run unit tests for specific components:

uv run pytest tests/levanter/test_example.py

For full validation before submitting, execute the safe-test suite:

uv run --no-project infra/ci/run_tests.py

Documentation Requirements

Every code change requires corresponding documentation updates. If you modify Markdown sources in docs/, verify link integrity:

uv run python infra/check_docs_source_links.py

Optionally build the documentation locally to catch build errors:

uv run mkdocs build --strict

Working with Agent Skills

Marin’s workflow utilizes agents defined in .agents/skills/ to standardize common tasks. These reusable skills encapsulate higher-level actions such as adding datasets or configuring experiments. When contributing new functionality, check whether an existing skill in .agents/skills/ already handles your use case rather than implementing ad-hoc scripts.

Example: Adding a New Optimizer to Levanter

Here is a complete example of contributing a minimal change to the levanter library. This adds a new optimizer configuration:


# File: lib/levanter/optim/adamx.py

from dataclasses import dataclass
from levanter.optim import OptimizerConfig

@dataclass
class AdamXConfig(OptimizerConfig):
    """AdamX with a custom decay schedule."""
    learning_rate: float = 1e-3
    beta1: float = 0.9
    beta2: float = 0.999
    weight_decay: float = 0.01

After creating the file, you must:

  1. Add a unit test in tests/levanter/optim/test_adamx.py
  2. Update docs/levanter/optimizers.md with documentation for the new class
  3. Run ./infra/pre-commit.py --all-files --fix to validate formatting

Summary

Contributing to the Marin project requires understanding its strict architectural boundaries and automated quality checks. Key takeaways include:

  • Respect the layer hierarchy: Code in lib/levanter/ cannot import from lib/marin/ or lib/iris/
  • Use uv for all dependency management and always include the --group dev flag
  • Run ./infra/pre-commit.py --all-files --fix before committing to ensure CI compliance
  • Include type hints for all public APIs and run pyrefly checks
  • Write tests in the corresponding tests/ subdirectory for every new feature
  • Update documentation in docs/ and verify links with infra/check_docs_source_links.py

Frequently Asked Questions

What Python version is required to contribute to the Marin project?

Marin requires Python 3.12 exactly. The pyproject.toml file specifies this version constraint, and the development workflow assumes uv is installed to manage the virtual environment and dependencies.

How do I run tests before submitting a contribution?

Run targeted unit tests with uv run pytest <path-to-test-file>. For comprehensive validation that mimics CI, execute uv run --no-project infra/ci/run_tests.py, which runs the full safe-test suite across all libraries including levanter, zephyr, and iris.

Can I add functionality that imports from higher-level libraries?

No. The architecture strictly prohibits lower-level libraries from importing higher-level ones. For example, code in lib/levanter/ cannot import from lib/marin/ or lib/iris/. Always place your contribution in the appropriate layer based on its abstraction level.

Where should I place new agent skills when contributing?

Place reusable agent skills in the .agents/skills/ directory. Before creating new skills, check if an existing skill already provides the functionality you need. Skills standardize common tasks like add-dataset and ensure consistency across the Marin project’s automated workflows.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →