How to Contribute to the Marin Project: A Complete Developer’s Guide
To contribute to the Marin project, you must set up a Python 3.12 environment with uv, install dependencies via uv sync --package marin-core --group dev, follow the layered architecture rules (lower-level libraries cannot import from higher-level ones), and submit pull requests that pass ./infra/pre-commit.py --all-files --fix and the full test suite.
Marin is a modular research platform for large-language-model development organized into distinct architectural layers. Whether you are fixing bugs, adding optimizers to the training library, or extending dataset processing capabilities, understanding the repository structure and development workflow ensures your contribution aligns with the project’s standards.
Understanding Marin’s Layered Architecture
The codebase is organized into four primary layers that enforce strict dependency directionality. Lower-level libraries cannot import from higher-level ones, ensuring clean separation of concerns.
- Levanter: The JAX-based training library located in
lib/levanter/ - Zephyr: Dataset processing utilities in
lib/zephyr/ - Iris: Job orchestration components in
lib/iris/ - Marin: The top-level pipeline that coordinates the lower layers
When you contribute to the Marin project, place new functionality in the appropriate sub-package. For example, a new optimizer belongs in lib/levanter/optim/, while a dataset processing tool belongs in lib/zephyr/.
Setting Up Your Development Environment
Before writing code, configure your local environment to match the repository’s requirements exactly.
Prerequisites and Installation
Marin requires Python 3.12 and uses uv for dependency management. Clone the repository and initialize your environment:
git clone https://github.com/marin-community/marin.git
cd marin
uv venv --python 3.12
source .venv/bin/activate
Install the core package along with development dependencies:
uv sync --package marin-core --group dev
Running Pre-Commit Hooks
All contributions must pass the project-wide linting and formatting checks enforced by ./infra/pre-commit.py. Run this command before every commit:
./infra/pre-commit.py --all-files --fix
This script executes the same checks used in CI, ensuring consistent code style across the repository.
The Contribution Workflow
Follow these sequential steps when preparing your contribution to the Marin project.
Reading the Contribution Guidelines
Start with the root CONTRIBUTING.md file, which redirects to the comprehensive guide at docs/dev-guide/contributing.md. This document specifies the detailed workflow, coding standards, and pull request requirements.
Coding Standards and Type Checking
All public APIs in Marin are strictly typed. The repository uses pyrefly for static type checking, and you must include type hints for any new functions or classes. When adding features to libraries like levanter, ensure your code follows the existing patterns in lib/levanter/ and does not introduce circular dependencies with higher-level packages.
Testing Your Changes
Marin maintains a comprehensive test suite in the tests/ directory. Run unit tests for specific components:
uv run pytest tests/levanter/test_example.py
For full validation before submitting, execute the safe-test suite:
uv run --no-project infra/ci/run_tests.py
Documentation Requirements
Every code change requires corresponding documentation updates. If you modify Markdown sources in docs/, verify link integrity:
uv run python infra/check_docs_source_links.py
Optionally build the documentation locally to catch build errors:
uv run mkdocs build --strict
Working with Agent Skills
Marin’s workflow utilizes agents defined in .agents/skills/ to standardize common tasks. These reusable skills encapsulate higher-level actions such as adding datasets or configuring experiments. When contributing new functionality, check whether an existing skill in .agents/skills/ already handles your use case rather than implementing ad-hoc scripts.
Example: Adding a New Optimizer to Levanter
Here is a complete example of contributing a minimal change to the levanter library. This adds a new optimizer configuration:
# File: lib/levanter/optim/adamx.py
from dataclasses import dataclass
from levanter.optim import OptimizerConfig
@dataclass
class AdamXConfig(OptimizerConfig):
"""AdamX with a custom decay schedule."""
learning_rate: float = 1e-3
beta1: float = 0.9
beta2: float = 0.999
weight_decay: float = 0.01
After creating the file, you must:
- Add a unit test in
tests/levanter/optim/test_adamx.py - Update
docs/levanter/optimizers.mdwith documentation for the new class - Run
./infra/pre-commit.py --all-files --fixto validate formatting
Summary
Contributing to the Marin project requires understanding its strict architectural boundaries and automated quality checks. Key takeaways include:
- Respect the layer hierarchy: Code in
lib/levanter/cannot import fromlib/marin/orlib/iris/ - Use
uvfor all dependency management and always include the--group devflag - Run
./infra/pre-commit.py --all-files --fixbefore committing to ensure CI compliance - Include type hints for all public APIs and run
pyreflychecks - Write tests in the corresponding
tests/subdirectory for every new feature - Update documentation in
docs/and verify links withinfra/check_docs_source_links.py
Frequently Asked Questions
What Python version is required to contribute to the Marin project?
Marin requires Python 3.12 exactly. The pyproject.toml file specifies this version constraint, and the development workflow assumes uv is installed to manage the virtual environment and dependencies.
How do I run tests before submitting a contribution?
Run targeted unit tests with uv run pytest <path-to-test-file>. For comprehensive validation that mimics CI, execute uv run --no-project infra/ci/run_tests.py, which runs the full safe-test suite across all libraries including levanter, zephyr, and iris.
Can I add functionality that imports from higher-level libraries?
No. The architecture strictly prohibits lower-level libraries from importing higher-level ones. For example, code in lib/levanter/ cannot import from lib/marin/ or lib/iris/. Always place your contribution in the appropriate layer based on its abstraction level.
Where should I place new agent skills when contributing?
Place reusable agent skills in the .agents/skills/ directory. Before creating new skills, check if an existing skill already provides the functionality you need. Skills standardize common tasks like add-dataset and ensure consistency across the Marin project’s automated workflows.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →