Marin Integration Test: Running a Full LM Pipeline in Under 10 Minutes

The Marin integration test executes a complete language model training pipeline in under ten minutes by orchestrating a directed acyclic graph of cached steps, using a deliberately minimal GPT-2 configuration (2 layers, 2 heads, 2 training steps), and running either locally with fixture data or remotely via Iris job submission.

The marin-community/marin repository provides a robust framework for large-scale data processing and language model training. Its integration test (tests/integration_test.py) validates the entire pipeline—from raw HTML ingestion to model training—without exhausting CI time budgets. By leveraging aggressive caching and a tiny model architecture, the test ensures end-to-end correctness while completing in well under the default ten-minute timeout.

Pipeline Architecture and Step Orchestration

The integration test constructs a directed acyclic graph (DAG) of StepSpec objects that represent discrete pipeline stages. In tests/integration_test.py, the create_steps() function (lines 79–110) assembles six sequential transformations that convert raw HTML into a trained language model. Each step declares its output path and hash attributes, enabling the StepRunner class to automatically skip cached results and parallelize execution where possible.

The StepRunner implementation in lib/marin/src/marin/execution/step_runner.py eagerly schedules steps as soon as their dependencies are satisfied, maintaining a dependency graph to ensure correct execution order. This caching mechanism ensures that reruns of the integration test complete in seconds rather than minutes, provided the underlying data and configuration remain unchanged.

The Six-Stage Data-to-Model Pipeline

The create_steps() function defines a six-step transformation pipeline that exercises every major subsystem of the Marin framework:

  1. HTML to Markdown Conversion – Uses SimpleHtmlToMdConfig with the html_to_md utility to extract clean text from HTML documents using the Resiliparse library.

  2. Data Normalization – The normalize_step creates a canonical Parquet representation of the textual data, standardizing schemas across heterogeneous inputs.

  3. Exact-Paragraph Deduplication – The dedup_exact_paragraph step removes duplicate paragraphs across the corpus to prevent training data contamination.

  4. Consolidation – The consolidate step removes duplicate spans within documents, ensuring clean, non-overlapping text segments.

  5. Tokenization – The tokenize step wraps normalized data in a token cache compatible with the Levanter training framework.

  6. Training – The run_levanter_train_lm function trains a minimal GPT-2 model (2 layers, 2 attention heads, sequence length 64) for exactly 2 steps with a batch size of 8.


# Assemble the pipeline DAG (excerpt from tests/integration_test.py)

def create_steps(prefix: str, synth_data: str, tokenizer: str) -> list[StepSpec]:
    # 1. Transform HTML → Markdown

    transform_hq_data_spec = StepSpec(
        name=os.path.join(prefix, "hq-transformed"),
        hash_attrs={"extract_method": "resiliparse"},
        fn=lambda out: html_to_md(
            SimpleHtmlToMdConfig(
                input_path=os.path.join(synth_data, "pos"),
                output_path=out,
                extract_method="resiliparse",
                config=ResiliparseConfig(),
            )
        ),
    )

    # ... normalization, dedup, consolidate, tokenize, train specs ...

    return [
        transform_hq_data_spec,
        normalize_hq_spec,
        dedup_exact_paragraph_spec,
        consolidate_spec,
        tokenize_spec,
        train_spec,
    ]

Execution Modes: Local Development vs. CI

The integration test supports two execution paths, both defined in tests/integration_test.py, ensuring the same code paths run regardless of environment:

Local Mode (Default)

In local mode, the test materializes a reduced GPT-2 tokenizer fixture (gpt2_tokenizer.json) to a temporary directory via _materialize_local_tokenizer(). All pipeline stages execute in-process on the same host, eliminating network I/O and S3 latency. The synthetic quick-start data from tests/quickstart-data provides the input corpus, allowing the test to run entirely offline.

S3 Remote Mode

When the MARIN_CI_S3_PREFIX environment variable is set, the test switches to remote execution. The _upload_tree() function uploads the synthetic data to S3, then submits the pipeline as an Iris job via FrayIrisClient.submit(). The remote job downloads the data and uses the public HuggingFace GPT-2 tokenizer. After completion, the temporary S3 prefix is automatically removed to prevent storage costs.


# Run the pipeline – either locally or as an Iris job (excerpt)

def main():
    parser = argparse.ArgumentParser()
    parser.add_argument("--controller-url", required=True)
    args = parser.parse_args()

    if os.environ.get("MARIN_CI_S3_PREFIX"):
        # S3 mode – upload fixtures, then submit as an Iris job

        _upload_tree(LOCAL_SYNTH_DATA, synth_data)
        tokenizer = "gpt2"
        with set_current_client(iris_client):
            handle = iris_client.submit(
                JobRequest(
                    name="marin-itest",
                    entrypoint=Entrypoint.from_callable(_run_pipeline, args=(prefix, synth_data, tokenizer)),
                    resources=ResourceConfig.with_cpu(),
                    environment=create_environment(env_vars=_s3_env_vars()),
                )
            )
            handle.wait(stream_logs=True)
    else:
        # Local mode – materialize the reduced tokenizer fixture

        tokenizer = _materialize_local_tokenizer(Path(prefix) / "gpt2-tokenizer")
        steps = create_steps("quickstart-tests", str(LOCAL_SYNTH_DATA), tokenizer)
        with set_current_client(iris_client):
            StepRunner().run(steps)

Performance Optimization Strategies

Several deliberate constraints enable the sub-ten-minute execution time:

  • Minimal Model Architecture – The GPT-2 configuration specified in lib/levanter/models/gpt2.py uses only 2 layers and 2 attention heads with a vocabulary size of approximately 5,000 tokens, dramatically reducing compute requirements compared to standard 12-layer models.

  • Truncated Training Run – The training step limits num_train_steps to 2, sufficient to verify the training loop logic without wasting compute on convergence.

  • Intelligent Caching – The StepRunner hashes each step's inputs and configuration, skipping recomputation for unchanged stages. This is particularly effective for the normalization and deduplication steps, which process the same synthetic data across test runs.

  • Fixture-Based Tokenization – Local mode uses a vendored tokenizer to avoid downloading multi-megabyte vocabulary files from HuggingFace, shaving precious seconds off CI startup time.

Summary

  • The Marin integration test orchestrates a six-step DAG (HTML→Markdown→Normalize→Dedup→Tokenize→Train) using StepSpec and StepRunner classes.
  • It completes in under ten minutes by training a tiny GPT-2 model (2 layers, 2 heads) for only 2 steps with a batch size of 8.
  • Local mode runs entirely in-process with fixture data, while S3 mode validates distributed execution via Iris job submission.
  • Aggressive step-level caching ensures subsequent test runs complete in seconds rather than minutes.
  • The test validates both the data processing pipeline (html_to_md, normalize_step, dedup_exact_paragraph) and the training infrastructure (run_levanter_train_lm).

Frequently Asked Questions

What is the Marin integration test designed to validate?

The Marin integration test validates the entire end-to-end pipeline from raw HTML ingestion through language model training. It exercises the StepRunner orchestration logic, the deduplication and normalization transformations, and the Levanter training loop to ensure all components integrate correctly before deployment to production clusters.

Why does the integration test use such a small model configuration?

The test uses a minimal GPT-2 architecture (2 layers, 2 heads, sequence length 64) to verify that the training infrastructure works correctly without consuming excessive compute resources. Since the goal is integration testing rather than model quality assessment, 2 training steps are sufficient to confirm that data flows correctly through the tokenization cache into the optimizer and that checkpointing functions properly.

How does the test handle dependencies between pipeline stages?

The test represents each stage as a StepSpec object declared in lib/marin/src/marin/execution/step_spec.py, which encapsulates the function to execute, output paths, and hash attributes for caching. The StepRunner class in lib/marin/src/marin/execution/step_runner.py builds a dependency graph from these specs and executes steps only after their prerequisite steps complete successfully, using filesystem existence checks and content hashes to skip redundant work.

Can I run the Marin integration test on my local machine without S3 access?

Yes. By default, the test runs in local mode, which uses the synthetic dataset in tests/quickstart-data and a vendored GPT-2 tokenizer fixture. You only need S3 credentials if you explicitly set the MARIN_CI_S3_PREFIX environment variable to test the distributed Iris job submission path.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →