# Best Practices for Developing with Marin: A Complete Guide to LLM Pipelines

> Master Marin development with our guide to LLM pipelines. Learn best practices for architecture, configurations, and dependencies to build robust applications.

- Repository: [The Marin Project/marin](https://github.com/marin-community/marin)
- Tags: best-practices
- Published: 2026-08-29

---

**Developing with Marin requires adopting a pipeline-first architecture where experiments are expressed as reusable steps, configurations use immutable typed data structures, and dependencies flow upward through a strictly layered library stack.**

Marin is a modular, Python-centric platform for building, training, and evaluating large language models. When developing with Marin, you work within a framework that prioritizes reproducible pipelines, explicit resource configuration, and behavior-focused testing. The following guidelines synthesize the repository’s official documentation, coding conventions in [`docs/dev-guide/coding-standards.md`](https://github.com/marin-community/marin/blob/main/docs/dev-guide/coding-standards.md), and testing policies from [`TESTING.md`](https://github.com/marin-community/marin/blob/main/TESTING.md) to help you write maintainable, performant code.

## Embrace the Pipeline-First Execution Model

Marin treats experiments as directed dependency graphs rather than linear scripts. Each logical unit—whether tokenization, training, or evaluation—must be declared as an **ArtifactStep** that defines its inputs and outputs explicitly.

### Define Steps Using ArtifactStep

Steps represent discrete, cacheable units of work. In [`experiments/tutorials/train_tiny_model.py`](https://github.com/marin-community/marin/blob/main/experiments/tutorials/train_tiny_model.py), the tutorial defines a dataset step and a training step that depend on one another. The framework uses **lazy handles** via `ArtifactStep` to defer I/O until a step actually executes, avoiding unnecessary downloads or computation.

```python
from marin.experiment.data import tokenized
from marin.processing.tokenize.tokenize import TokenizedCache
from experiments.marin_tokenizer import marin_tokenizer

def tinystories_tokenized() -> ArtifactStep[TokenizedCache]:
    """Tokenize TinyStories using the shared tokenizer."""
    return tokenized(
        name="tokenized/tinystories",
        source="roneneldan/TinyStories",
        tokenizer=marin_tokenizer,
        sample_count=1000,
    )

```

### Orchestrate with StepRunner

Use the `StepRunner` class to execute steps in topological order. The runner automatically schedules dependencies so downstream steps only run after their inputs are ready, similar to a Makefile. Convert high-level step objects into lazy handles using the `lower` helper before execution.

```python
from marin.execution.step_runner import StepRunner
from marin.execution.lazy import lower

if __name__ == "__main__":
    StepRunner().run([lower(build(device="cpu", data="tinystories"))])

```

## Configure Resources and Devices Explicitly

Marin requires explicit resource allocation through the `ResourceConfig` class. Avoid hardcoding device specifications; instead, centralize configurations in dictionaries that map environment names to resource requirements and batch sizes.

```python
from fray.types import ANY_REGION, ResourceConfig

DEVICES = {
    "cpu": (ResourceConfig.with_cpu(), 4),
    "h100x8": (
        ResourceConfig.with_gpu(
            "H100", count=8, cpu=32, disk="128G", ram="128G", regions=[ANY_REGION]
        ),
        256,
    ),
}

```

Always pair device configurations with appropriate batch sizes to prevent out-of-memory errors. This centralized approach, demonstrated in [`experiments/tutorials/train_tiny_model.py`](https://github.com/marin-community/marin/blob/main/experiments/tutorials/train_tiny_model.py), allows you to add new accelerators with a single line change.

## Maintain Strict Coding Standards

According to [`docs/dev-guide/coding-standards.md`](https://github.com/marin-community/marin/blob/main/docs/dev-guide/coding-standards.md), Marin enforces specific patterns to maintain code quality across its layered architecture.

### Use Immutable Typed Configurations

Replace raw dictionaries and string literals with **immutable dataclasses** or `StrEnum`. Mark configuration objects with `@dataclass(frozen=True)` to prevent accidental mutation and enable reliable caching.

```python
from dataclasses import dataclass
from enum import StrEnum

class OptimizerType(StrEnum):
    ADAM = "adam"
    SGD = "sgd"

@dataclass(frozen=True)
class TrainingConfig:
    learning_rate: float = 6e-4
    batch_size: int = 256

```

### Follow Naming and Import Conventions

Keep all imports at the top of the file; avoid mid-function imports except when guarding optional third-party packages. Name functions as verbs describing their return type (e.g., `train_lm` rather than `run_training`). Do not create utility modules named `*_utils.py`; use descriptive domain names instead.

**Avoid boolean flags for variant behavior.** Instead of using a `docker: bool` parameter, create distinct classes such as `NativeVllm` vs `DockerVllm` to maintain type safety and clarity.

## Respect the Layered Dependency Model

Marin’s architecture organizes code into layers where low-level libraries (`iris`, `haliax`) sit beneath higher-level ones (`levanter`, `zephyr`, `marin`). As documented in [`lib/marin/AGENTS.md`](https://github.com/marin-community/marin/blob/main/lib/marin/AGENTS.md), imports must flow only upward; never import from a higher layer into a lower one. This strict directionality prevents cyclic dependencies and maintains clean architectural boundaries.

## Write Behavior-Focused Tests

The testing philosophy in [`TESTING.md`](https://github.com/marin-community/marin/blob/main/TESTING.md) mandates that tests assert observable public behavior rather than implementation details. Valid assertions include API return values, persisted artifact states, or numeric parity against reference implementations.

**Prohibited patterns** include verifying that a method exists, checking private attribute values, or asserting that log lines contain specific phrases. Prefer fakes over mocks for internal components; reserve mocks strictly for external I/O such as `gcloud` API calls or HTTP requests.

## Set Up Local Development Workflow

Before submitting changes, run the full local quality assurance pipeline as specified in [`lib/marin/AGENTS.md`](https://github.com/marin-community/marin/blob/main/lib/marin/AGENTS.md).

Execute the pre-commit formatter to ensure consistent style:

```bash
./infra/pre-commit.py --all-files --fix

```

Verify type safety using Pyrefly:

```bash
uv run pyrefly check

```

Run the fast test suite to catch regressions:

```bash
uv run --no-project infra/ci/run_tests.py

```

## Summary

- **Pipeline-first design**: Express experiments as `ArtifactStep` objects with explicit dependencies, executed via `StepRunner`.
- **Explicit configuration**: Use `ResourceConfig` helpers and centralize device definitions; employ `@dataclass(frozen=True)` for all configs.
- **Layered architecture**: Import only upward from `iris`/`haliax` to `marin`; never reverse the direction.
- **Behavior testing**: Assert public API contracts and persisted state; avoid testing internal details or using excessive mocks.
- **Quality gates**: Run [`./infra/pre-commit.py`](https://github.com/marin-community/marin/blob/main/./infra/pre-commit.py), `pyrefly check`, and the CI test suite before committing.

## Frequently Asked Questions

### What is the pipeline-first model in Marin?

The pipeline-first model treats machine learning experiments as graphs of discrete steps (e.g., tokenization, training, evaluation) defined using `ArtifactStep`. Each step declares its dependencies explicitly, and the `StepRunner` executes them in topological order. This ensures reproducibility and allows the framework to cache intermediate results and defer I/O operations until necessary.

### How do I configure GPU resources in Marin?

Use the `ResourceConfig.with_gpu()` helper to specify hardware requirements including GPU type, count, CPU cores, disk, and RAM. Store device configurations in a centralized dictionary (as shown in [`experiments/tutorials/train_tiny_model.py`](https://github.com/marin-community/marin/blob/main/experiments/tutorials/train_tiny_model.py)) that maps environment keys to resource tuples. Always pair resource definitions with appropriate batch sizes to prevent out-of-memory errors during training.

### What testing approach does Marin recommend?

Marin requires **behavior-focused testing** that verifies public API contracts, persisted artifact states, or numerical parity against reference implementations. Tests must not depend on internal implementation details such as private attributes or method existence. Use fakes for internal dependencies and reserve mocks exclusively for external I/O like cloud services or HTTP endpoints.

### How do I avoid cyclic dependencies in Marin?

Follow the layered dependency model documented in [`lib/marin/AGENTS.md`](https://github.com/marin-community/marin/blob/main/lib/marin/AGENTS.md): import only upward through the architectural stack from low-level libraries (`iris`, `haliax`) to higher-level ones (`levanter`, `marin`). Never import from a higher layer into a lower one. This one-way dependency flow prevents cycles and maintains clean separation between the orchestration, dataset processing, and model training layers.