# How to Write and Run Tests Following Marin's Testing Policy: A Complete Guide

> Learn how to write and run tests following Marin's testing policy. Master behavior-focused integration tests and ensure reliable validation with `uv run pytest`.

- Repository: [The Marin Project/marin](https://github.com/marin-community/marin)
- Tags: how-to-guide
- Published: 2026-08-29

---

**Follow Marin's testing policy by writing behavior-focused integration tests that verify public APIs, use descriptive naming conventions like `test_<subject>_<scenario>_<expected_outcome>`, and execute them with `uv run pytest` or the CI script at [`infra/ci/run_tests.py`](https://github.com/marin-community/marin/blob/main/infra/ci/run_tests.py) to ensure reliable, implementation-agnostic validation.**

The marin-community/marin repository maintains rigorous quality standards through a comprehensive testing philosophy defined in the root [`TESTING.md`](https://github.com/marin-community/marin/blob/main/TESTING.md). Understanding how to write and run tests following Marin's testing policy ensures your code integrates seamlessly with the existing CI pipeline while avoiding brittle implementation-detail checks that break during refactoring.

## Core Testing Philosophy: Behavior Over Implementation

Marin's policy mandates that tests must serve as behavioral contracts rather than mirrors of the current implementation. This approach prevents test suites from becoming maintenance burdens when internal logic changes.

### The Fail-When-Behavior-is-Wrong Rule

According to [`TESTING.md`](https://github.com/marin-community/marin/blob/main/TESTING.md), every test must **fail only when observed external behavior is incorrect**. This means assertions must validate public API outputs, persisted state, or real side-effects rather than private attributes or internal call counts. Tests that rely on implementation details create false negatives during legitimate refactoring and violate the core principles outlined in the repository's testing guidelines.

### Rejecting "Slop" Tests

The policy explicitly forbids several categories of low-value tests:

- **Tautologies** that assert `True is True` or mirror implementation logic line-for-line
- **Private-state assertions** that peek into internal attributes prefixed with underscores
- **Incidental string checks** that validate exact error message formatting rather than error types
- **Call-count verification** that only confirms a method was invoked without checking results

Remove these "slop" tests during code review to maintain suite integrity.

## Test Structure and Naming Conventions

Consistent organization enables developers to locate and understand tests quickly across the `marin` codebase.

### File and Function Naming

All test files must follow the pattern `test_<module>.py`. Inside these files, functions must use the descriptive format:

```python
def test_<subject>_<scenario>_<expected_outcome>():
    ...

```

For example, `test_job_submission_with_invalid_input_raises_value_error` clearly communicates the test's scope without requiring code inspection. This convention appears throughout the Iris sub-project at [`lib/iris/TESTING.md`](https://github.com/marin-community/marin/blob/main/lib/iris/TESTING.md) and helps pytest discovery mechanisms identify relevant tests automatically.

### Parameterization Best Practices

When testing boundary conditions, use pytest's parameterization to avoid duplication. The following example from the testing guidelines demonstrates validating multiple invalid inputs:

```python
import pytest

@pytest.mark.parametrize(
    "input_shape,expected_error",
    [
        ((), "Empty input not allowed"),
        ((0, 10), "Zero dimension not allowed"),
    ],
)
def test_invalid_input_shapes(marin_client, input_shape, expected_error):
    with pytest.raises(ValueError, match=expected_error):
        marin_client.process_tensor(shape=input_shape)

```

This approach satisfies the policy's requirement for meaningful variations while keeping test functions focused and readable.

## Handling External Dependencies with Fakes

Marin's policy strongly prefers **fakes over mocks** for external dependencies. While mocking is restricted to true I/O boundaries (subprocesses, HTTP, GCS), in-memory fakes that provide the same interface are preferred because they retain realistic behavior without network latency or costs.

### When to Use Mocks vs Fakes

For Google Cloud Storage interactions, the repository provides `InMemoryGcsService` in [`lib/iris/testing/fake_gcs.py`](https://github.com/marin-community/marin/blob/main/lib/iris/testing/fake_gcs.py). Use this fake instead of mocking GCS client methods:

```python
from iris.testing.fake_gcs import InMemoryGcsService

def test_save_to_gcs_with_fake(gcs_fake: InMemoryGcsService, marin_client):
    # Inject the fake into the client (client accepts a GCS interface)

    marin_client.gcs = gcs_fake

    # Perform an operation that writes to GCS

    marin_client.save_artifact("model.pt", b"model bytes")

    # Verify the file exists in the fake storage

    assert gcs_fake.exists("model.pt")
    assert gcs_fake.read("model.pt") == b"model bytes"

```

This integration-style test exercises the full write-and-read cycle through public APIs while remaining hermetic and fast.

## Timeouts, Markers, and CI Integration

Reliable test execution requires strict时间管理 (time management) and proper classification of resource-intensive tests.

### Managing Test Duration

The default per-test timeout is **60 seconds**. Tests requiring longer execution must explicitly declare this using:

```python
import pytest

@pytest.mark.timeout(120)  # 2-minute timeout

def test_large_dataset_processing(marin_client):
    result = marin_client.process_dataset(dataset="large")
    assert result.success

```

Avoid `time.sleep()` in tests; instead, inject a fake clock through fixtures defined in [`conftest.py`](https://github.com/marin-community/marin/blob/main/conftest.py) to simulate time passage deterministically.

### Using Markers to Classify Tests

Heavyweight tests must use specific markers to prevent accidental execution during rapid local development cycles:

- `@pytest.mark.slow` for computationally expensive operations
- `@pytest.mark.integration` for tests requiring external services
- `@pytest.mark.docker` for container-dependent validation
- `@pytest.mark.requires_cluster` for distributed system tests (Iris-specific)

The default test run excludes these markers unless explicitly requested, keeping local feedback loops fast.

## Running Tests Locally and in Continuous Integration

Marin provides two primary methods for executing tests, depending on your development phase.

For narrow, targeted validation during active development, use:

```bash
uv run pytest <path_to_test_file>

```

For comprehensive validation before pushing, execute the CI-compatible test runner:

```bash
uv run --no-project infra/ci/run_tests.py

```

The [`infra/ci/run_tests.py`](https://github.com/marin-community/marin/blob/main/infra/ci/run_tests.py) script automatically discovers all "safe" tests affected by your current branch changes, respecting the default marker exclusions (slow, integration, docker) while ensuring you don't commit broken code. This script mirrors the validation performed in the actual CI environment.

## Summary

- **Behavior-focused testing**: Write tests that fail only when external behavior changes, avoiding private attribute checks and call-count verification as specified in [`TESTING.md`](https://github.com/marin-community/marin/blob/main/TESTING.md).
- **Integration over unit**: Prefer testing public APIs and real side-effects rather than internal implementation details.
- **Naming discipline**: Use `test_<subject>_<scenario>_<expected_outcome>` functions inside `test_<module>.py` files for discoverability.
- **Fakes preferred**: Replace external services like GCS with in-memory fakes (e.g., `InMemoryGcsService`) rather than mocking interfaces.
- **Timeout awareness**: Respect the 60-second default; use `@pytest.mark.timeout()` for longer tests and avoid `time.sleep()`.
- **Marker classification**: Tag heavy tests with `@pytest.mark.slow`, `@pytest.mark.integration`, or `@pytest.mark.docker` to maintain fast local runs.
- **Execution workflow**: Use `uv run pytest` for local development and [`infra/ci/run_tests.py`](https://github.com/marin-community/marin/blob/main/infra/ci/run_tests.py) for pre-commit validation.

## Frequently Asked Questions

### What's the difference between unit and integration tests in Marin's testing policy?

Marin's policy discourages traditional unit tests that isolate internal functions in favor of **integration-style tests** that exercise public APIs and real side-effects. According to [`TESTING.md`](https://github.com/marin-community/marin/blob/main/TESTING.md), tests should validate "submit job → observe status" workflows rather than individual helper methods. This ensures tests remain valid during refactoring and catch real behavioral regressions rather than implementation changes.

### How should I handle Google Cloud Storage dependencies in my tests?

Use the `InMemoryGcsService` fake provided in [`lib/iris/testing/fake_gcs.py`](https://github.com/marin-community/marin/blob/main/lib/iris/testing/fake_gcs.py) rather than mocking GCS client methods. Inject this fake into your client code to verify that write operations actually persist data and read operations retrieve correct bytes. This approach satisfies the "use fakes over mocks" rule and provides more realistic validation than asserting that specific methods were called.

### Why does my test pass locally but fail in CI?

Likely causes include missing markers (slow, integration, docker) that cause the test to run in environments without required dependencies, or hardcoded timeouts exceeding the 60-second default. Check [`infra/ci/run_tests.py`](https://github.com/marin-community/marin/blob/main/infra/ci/run_tests.py) to see which markers are excluded by default, and ensure resource-intensive tests use appropriate `@pytest.mark` decorators. Also verify you aren't relying on local state that doesn't exist in clean CI environments.

### Can I use mocks instead of fakes for external HTTP calls?

Only at true I/O boundaries. Marin's policy permits mocking for subprocess execution, HTTP requests, and cloud service boundaries, but strongly prefers in-memory fakes that implement the same interface when available. If you must mock, ensure you're verifying behavioral outcomes (data transformation, state changes) rather than checking that specific functions were called with exact arguments.