How to Integrate Marin with Other Tools: A Complete Guide to Modular Pipeline Extensions

Integrating Marin with external tools involves declaring dependencies in config/external/README.md, creating thin wrapper classes that conform to layer-specific interfaces, and wiring them into the pipeline via configuration files or entry points.

Marin is architected as a modular pipeline that stitches together specialized sub‑libraries, making it straightforward to plug in new tools or replace existing components without modifying core code. Whether you need to add a custom data reader, register a new model architecture, or benchmark against an external evaluation suite, Marin’s layered design enforces clean separation through typed JSON payloads and standardized entry points. This guide walks through the practical steps to integrate tools at every layer of the Marin stack, from data ingestion to distributed training and evaluation.

Understanding Marin’s Modular Architecture

Marin decomposes the machine learning lifecycle into discrete layers, each housed in its own sub‑library under lib/. This separation allows you to swap implementations without cascading changes:

  • Data Ingestion & Transformation (lib/zephyr): Handles dataset readers, writers, and shuffling, with built‑in fsspec integration for GCS and S3 storage (Zephyr README – fsspec integration).
  • Data Cleaning & Filtering: Uses external utilities such as fastText and trafilatura via the transform package; the pipeline is agnostic to the specific HTML‑to‑text converter or classifier used.
  • Tokenization (lib/marin/tools/get_hf_dataset_schema.py): Thin wrappers around Hugging Face tokenizers, allowing any HF tokenizer to be dropped in directly.
  • Model Training (lib/levanter): Powers the training loop, invoked from experiments/simple_train_config.py, with support for custom model registration.
  • Job Orchestration (lib/iris): Provides the distributed execution engine; jobs receive an endpoint URL that downstream components call.
  • Evaluation & Benchmarking: Integrates with Harbor via docs/harbor-integration.md, using the launcher at experiments/evaluation/cli to create isolated subprocesses that write results to Parquet files.
  • Telemetry & Logging (lib/finelog and lib/finestore): Capture metrics and artefacts via a simple API used by both training and evaluation stages.

Each layer communicates via typed JSON payloads—for example, the policy JSON sent from Marin to Harbor (lines 27‑30 of docs/harbor-integration.md)—ensuring that external tools only need to understand the schema, not the internal implementation.

Integration Patterns by Pipeline Layer

Data Ingestion with Zephyr

To add support for a new storage backend or file format, implement a Reader subclass in lib/zephyr/src/zephyr/readers/. The base interface requires only that you yield (key, bytes) pairs. Zephyr already integrates with fsspec, so custom filesystems can often leverage existing URI handlers.

Key integration points:

  • Implement the Reader interface from zephyr.readers.base.
  • Register the class in a pipeline YAML under the reader key.
  • Use url_to_fs from fsspec if implementing custom storage connectors.

Model Training with Levanter

New model architectures integrate via lib/levanter. The process involves subclassing BaseModel and referencing the implementation in your training configuration.

Key integration points:

Evaluation with Harbor

Harbor benchmarks run through experiments/evaluation/cli.py, which creates an isolated subprocess communicating over an Iris endpoint. Results are serialized to Parquet files as defined in tests/evaluation/test_harbor_runner.py.

Key integration points:

  • Add Harbor policy YAML files under experiments/evaluation/configs/harbor/.
  • Invoke the launcher using python -m experiments.evaluation.cli launch.
  • Configure the handshake parameters documented in docs/harbor-integration.md.

Step‑by‑Step Integration Workflow

Regardless of which layer you extend, the pattern remains consistent: declare → wrap → wire → test.

  1. Declare a new external dependency

    Add the package to marin.external_dependencies as documented in config/external/README.md. Marin isolates the dependency in its own uv lock, preventing contamination of the core environment.

  2. Expose a thin wrapper

    Implement a wrapper conforming to the expected interface (e.g., a Reader for Zephyr or a ModelTrainer for Levanter). Most wrappers live under lib/<subproject>/src/ and are discovered via entry‑points defined in pyproject.toml.

  3. Wire the wrapper into the pipeline

    • Data readers: Plug the class into the DatasetConfig used by Zephyr.
    • Training: Reference the registered model from a training config in experiments/simple_train_config.py.
    • Evaluation: Add a Harbor policy YAML and invoke the launcher at experiments/evaluation/cli.
  4. Validate with the integration test

    Run tests/integration_test.py to verify that your connector functions correctly across all pipeline stages.

Code Examples for Common Integrations

Running a Harbor Benchmark from Python

The following script demonstrates how to programmatically launch a Harbor evaluation using the CLI module at experiments/evaluation/cli:


# examples/harbor_demo.py

import subprocess

def launch_harbor(model: str, policy: str, limit: int = 2):
    cmd = [
        "uv", "run", "python", "-m", "experiments.evaluation.cli", "launch",
        "--model", model,
        "--harbor-config", f"experiments/evaluation/configs/harbor/{policy}.yaml",
        "--limit", str(limit),
    ]
    subprocess.check_call(cmd)

if __name__ == "__main__":
    launch_harbor("qwen3-8b", "aime-smoke")

Registering a New Levanter Model

To add a custom architecture, subclass BaseModel and import it into your training script:


# lib/levanter/src/levanter/models/custom.py

from levanter.models.base import BaseModel

class MyModel(BaseModel):
    def __init__(self, config):
        super().__init__(config)
        # model construction …

# In experiments/simple_train_config.py

from levanter.models.custom import MyModel

model = MyModel(config=my_config)
train_lm(model, dataset=..., ... )

Adding a Zephyr Reader for Custom Storage

Implement the Reader interface to support non‑standard storage backends:


# lib/zephyr/src/zephyr/readers/custom_fs.py

from zephyr.readers.base import Reader
from fsspec.core import url_to_fs

class CustomFSReader(Reader):
    def __init__(self, uri: str):
        self.fs, self.path = url_to_fs(uri)

    def __iter__(self):
        # yield (key, bytes) pairs

        for file in self.fs.ls(self.path):
            yield file, self.fs.open(file).read()

Reference the reader in a pipeline configuration:


# experiments/pipelines/custom.yaml

reader:
  class: zephyr.readers.custom_fs.CustomFSReader
  args:
    uri: "s3://my-bucket/dataset/"

Summary

  • Marin’s architecture splits functionality into isolated sub‑libraries (lib/zephyr, lib/levanter, lib/iris, etc.), allowing you to integrate tools at specific layers without affecting others.
  • All integrations follow the declare → wrap → wire → test pattern: declare dependencies in config/external/README.md, wrap external tools to conform to layer interfaces, wire them via YAML configs or entry points, and validate with tests/integration_test.py.
  • Communication between layers uses typed JSON payloads, decoupling components and enabling painless substitution of implementations.
  • Key files for integration include docs/harbor-integration.md (evaluation), lib/levanter/docs/dev/Port-Models.md (training), and lib/zephyr/README.md (data ingestion).

Frequently Asked Questions

How do I add a new external package dependency to Marin?

Add the package specification to marin.external_dependencies following the guidelines in config/external/README.md. Marin manages external dependencies in isolated uv locks, ensuring that new packages do not conflict with the core environment or other tools.

Can I use a custom tokenizer that is not from Hugging Face?

Yes. While lib/marin/tools/get_hf_dataset_schema.py provides thin wrappers around Hugging Face tokenizers, the tokenization layer is interface‑based. You can implement a compatible wrapper following the patterns in tests/test_marin_tokenizer.py and inject it into the pipeline configuration without modifying core Marin code.

What is the correct way to launch a Harbor evaluation programmatically?

Use the CLI module at experiments/evaluation/cli.py via subprocess or direct Python invocation. The launcher creates an isolated subprocess that communicates with Harbor over an Iris endpoint, then writes results to Parquet files. Refer to docs/harbor-integration.md for the exact JSON schema expected in the policy configuration.

Where should I implement a new dataset reader for an unsupported storage backend?

Implement the Reader interface in lib/zephyr/src/zephyr/readers/, inheriting from zephyr.readers.base.Reader. If the storage supports fsspec, leverage url_to_fs to minimize code. Register the reader class in your pipeline YAML under the reader key, specifying the full module path and initialization arguments.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →