# Marin Integration Test: Running a Full LM Pipeline in Under 10 Minutes

> Learn how the Marin integration test runs a full LM pipeline in under 10 minutes. Discover its efficient orchestration of cached steps and minimal GPT-2 setup for rapid testing.

- Repository: [The Marin Project/marin](https://github.com/marin-community/marin)
- Tags: how-to-guide
- Published: 2026-08-28

---

**The Marin integration test executes a complete language model training pipeline in under ten minutes by orchestrating a directed acyclic graph of cached steps, using a deliberately minimal GPT-2 configuration (2 layers, 2 heads, 2 training steps), and running either locally with fixture data or remotely via Iris job submission.**

The marin-community/marin repository provides a robust framework for large-scale data processing and language model training. Its integration test ([`tests/integration_test.py`](https://github.com/marin-community/marin/blob/main/tests/integration_test.py)) validates the entire pipeline—from raw HTML ingestion to model training—without exhausting CI time budgets. By leveraging aggressive caching and a tiny model architecture, the test ensures end-to-end correctness while completing in well under the default ten-minute timeout.

## Pipeline Architecture and Step Orchestration

The integration test constructs a **directed acyclic graph (DAG)** of `StepSpec` objects that represent discrete pipeline stages. In [`tests/integration_test.py`](https://github.com/marin-community/marin/blob/main/tests/integration_test.py), the `create_steps()` function (lines 79–110) assembles six sequential transformations that convert raw HTML into a trained language model. Each step declares its output path and hash attributes, enabling the `StepRunner` class to automatically skip cached results and parallelize execution where possible.

The `StepRunner` implementation in [`lib/marin/src/marin/execution/step_runner.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/execution/step_runner.py) eagerly schedules steps as soon as their dependencies are satisfied, maintaining a dependency graph to ensure correct execution order. This caching mechanism ensures that reruns of the integration test complete in seconds rather than minutes, provided the underlying data and configuration remain unchanged.

## The Six-Stage Data-to-Model Pipeline

The `create_steps()` function defines a six-step transformation pipeline that exercises every major subsystem of the Marin framework:

1. **HTML to Markdown Conversion** – Uses `SimpleHtmlToMdConfig` with the `html_to_md` utility to extract clean text from HTML documents using the Resiliparse library.

2. **Data Normalization** – The `normalize_step` creates a canonical Parquet representation of the textual data, standardizing schemas across heterogeneous inputs.

3. **Exact-Paragraph Deduplication** – The `dedup_exact_paragraph` step removes duplicate paragraphs across the corpus to prevent training data contamination.

4. **Consolidation** – The `consolidate` step removes duplicate spans within documents, ensuring clean, non-overlapping text segments.

5. **Tokenization** – The `tokenize` step wraps normalized data in a token cache compatible with the Levanter training framework.

6. **Training** – The `run_levanter_train_lm` function trains a minimal GPT-2 model (2 layers, 2 attention heads, sequence length 64) for exactly 2 steps with a batch size of 8.

```python

# Assemble the pipeline DAG (excerpt from tests/integration_test.py)

def create_steps(prefix: str, synth_data: str, tokenizer: str) -> list[StepSpec]:
    # 1. Transform HTML → Markdown

    transform_hq_data_spec = StepSpec(
        name=os.path.join(prefix, "hq-transformed"),
        hash_attrs={"extract_method": "resiliparse"},
        fn=lambda out: html_to_md(
            SimpleHtmlToMdConfig(
                input_path=os.path.join(synth_data, "pos"),
                output_path=out,
                extract_method="resiliparse",
                config=ResiliparseConfig(),
            )
        ),
    )

    # ... normalization, dedup, consolidate, tokenize, train specs ...

    return [
        transform_hq_data_spec,
        normalize_hq_spec,
        dedup_exact_paragraph_spec,
        consolidate_spec,
        tokenize_spec,
        train_spec,
    ]

```

## Execution Modes: Local Development vs. CI

The integration test supports two execution paths, both defined in [`tests/integration_test.py`](https://github.com/marin-community/marin/blob/main/tests/integration_test.py), ensuring the same code paths run regardless of environment:

### Local Mode (Default)

In local mode, the test materializes a reduced GPT-2 tokenizer fixture ([`gpt2_tokenizer.json`](https://github.com/marin-community/marin/blob/main/gpt2_tokenizer.json)) to a temporary directory via `_materialize_local_tokenizer()`. All pipeline stages execute in-process on the same host, eliminating network I/O and S3 latency. The synthetic quick-start data from `tests/quickstart-data` provides the input corpus, allowing the test to run entirely offline.

### S3 Remote Mode

When the `MARIN_CI_S3_PREFIX` environment variable is set, the test switches to remote execution. The `_upload_tree()` function uploads the synthetic data to S3, then submits the pipeline as an Iris job via `FrayIrisClient.submit()`. The remote job downloads the data and uses the public HuggingFace GPT-2 tokenizer. After completion, the temporary S3 prefix is automatically removed to prevent storage costs.

```python

# Run the pipeline – either locally or as an Iris job (excerpt)

def main():
    parser = argparse.ArgumentParser()
    parser.add_argument("--controller-url", required=True)
    args = parser.parse_args()

    if os.environ.get("MARIN_CI_S3_PREFIX"):
        # S3 mode – upload fixtures, then submit as an Iris job

        _upload_tree(LOCAL_SYNTH_DATA, synth_data)
        tokenizer = "gpt2"
        with set_current_client(iris_client):
            handle = iris_client.submit(
                JobRequest(
                    name="marin-itest",
                    entrypoint=Entrypoint.from_callable(_run_pipeline, args=(prefix, synth_data, tokenizer)),
                    resources=ResourceConfig.with_cpu(),
                    environment=create_environment(env_vars=_s3_env_vars()),
                )
            )
            handle.wait(stream_logs=True)
    else:
        # Local mode – materialize the reduced tokenizer fixture

        tokenizer = _materialize_local_tokenizer(Path(prefix) / "gpt2-tokenizer")
        steps = create_steps("quickstart-tests", str(LOCAL_SYNTH_DATA), tokenizer)
        with set_current_client(iris_client):
            StepRunner().run(steps)

```

## Performance Optimization Strategies

Several deliberate constraints enable the sub-ten-minute execution time:

- **Minimal Model Architecture** – The GPT-2 configuration specified in [`lib/levanter/models/gpt2.py`](https://github.com/marin-community/marin/blob/main/lib/levanter/models/gpt2.py) uses only 2 layers and 2 attention heads with a vocabulary size of approximately 5,000 tokens, dramatically reducing compute requirements compared to standard 12-layer models.

- **Truncated Training Run** – The training step limits `num_train_steps` to 2, sufficient to verify the training loop logic without wasting compute on convergence.

- **Intelligent Caching** – The `StepRunner` hashes each step's inputs and configuration, skipping recomputation for unchanged stages. This is particularly effective for the normalization and deduplication steps, which process the same synthetic data across test runs.

- **Fixture-Based Tokenization** – Local mode uses a vendored tokenizer to avoid downloading multi-megabyte vocabulary files from HuggingFace, shaving precious seconds off CI startup time.

## Summary

- The Marin integration test orchestrates a six-step DAG (HTML→Markdown→Normalize→Dedup→Tokenize→Train) using `StepSpec` and `StepRunner` classes.
- It completes in under ten minutes by training a tiny GPT-2 model (2 layers, 2 heads) for only 2 steps with a batch size of 8.
- Local mode runs entirely in-process with fixture data, while S3 mode validates distributed execution via Iris job submission.
- Aggressive step-level caching ensures subsequent test runs complete in seconds rather than minutes.
- The test validates both the data processing pipeline (`html_to_md`, `normalize_step`, `dedup_exact_paragraph`) and the training infrastructure (`run_levanter_train_lm`).

## Frequently Asked Questions

### What is the Marin integration test designed to validate?

The Marin integration test validates the entire end-to-end pipeline from raw HTML ingestion through language model training. It exercises the `StepRunner` orchestration logic, the deduplication and normalization transformations, and the Levanter training loop to ensure all components integrate correctly before deployment to production clusters.

### Why does the integration test use such a small model configuration?

The test uses a minimal GPT-2 architecture (2 layers, 2 heads, sequence length 64) to verify that the training infrastructure works correctly without consuming excessive compute resources. Since the goal is integration testing rather than model quality assessment, 2 training steps are sufficient to confirm that data flows correctly through the tokenization cache into the optimizer and that checkpointing functions properly.

### How does the test handle dependencies between pipeline stages?

The test represents each stage as a `StepSpec` object declared in [`lib/marin/src/marin/execution/step_spec.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/execution/step_spec.py), which encapsulates the function to execute, output paths, and hash attributes for caching. The `StepRunner` class in [`lib/marin/src/marin/execution/step_runner.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/execution/step_runner.py) builds a dependency graph from these specs and executes steps only after their prerequisite steps complete successfully, using filesystem existence checks and content hashes to skip redundant work.

### Can I run the Marin integration test on my local machine without S3 access?

Yes. By default, the test runs in local mode, which uses the synthetic dataset in `tests/quickstart-data` and a vendored GPT-2 tokenizer fixture. You only need S3 credentials if you explicitly set the `MARIN_CI_S3_PREFIX` environment variable to test the distributed Iris job submission path.