# How to Integrate Marin with Other Tools: A Complete Guide to Modular Pipeline Extensions

> Learn how to integrate Marin with other tools by creating wrapper classes and wiring them into your pipeline. This guide covers modular pipeline extensions for seamless integration.

- Repository: [The Marin Project/marin](https://github.com/marin-community/marin)
- Tags: how-to-guide
- Published: 2026-08-29

---

**Integrating Marin with external tools involves declaring dependencies in [`config/external/README.md`](https://github.com/marin-community/marin/blob/main/config/external/README.md), creating thin wrapper classes that conform to layer-specific interfaces, and wiring them into the pipeline via configuration files or entry points.**

Marin is architected as a modular pipeline that stitches together specialized sub‑libraries, making it straightforward to plug in new tools or replace existing components without modifying core code. Whether you need to add a custom data reader, register a new model architecture, or benchmark against an external evaluation suite, Marin’s layered design enforces clean separation through typed JSON payloads and standardized entry points. This guide walks through the practical steps to integrate tools at every layer of the Marin stack, from data ingestion to distributed training and evaluation.

## Understanding Marin’s Modular Architecture

Marin decomposes the machine learning lifecycle into discrete layers, each housed in its own sub‑library under `lib/`. This separation allows you to swap implementations without cascading changes:

- **Data Ingestion & Transformation** (`lib/zephyr`): Handles dataset readers, writers, and shuffling, with built‑in **fsspec** integration for GCS and S3 storage ([Zephyr README – fsspec integration](https://github.com/marin-community/marin/blob/main/lib/zephyr/README.md#fsspec-integration)).
- **Data Cleaning & Filtering**: Uses external utilities such as **fastText** and **trafilatura** via the `transform` package; the pipeline is agnostic to the specific HTML‑to‑text converter or classifier used.
- **Tokenization** ([`lib/marin/tools/get_hf_dataset_schema.py`](https://github.com/marin-community/marin/blob/main/lib/marin/tools/get_hf_dataset_schema.py)): Thin wrappers around Hugging Face tokenizers, allowing any HF tokenizer to be dropped in directly.
- **Model Training** (`lib/levanter`): Powers the training loop, invoked from [`experiments/simple_train_config.py`](https://github.com/marin-community/marin/blob/main/experiments/simple_train_config.py), with support for custom model registration.
- **Job Orchestration** (`lib/iris`): Provides the distributed execution engine; jobs receive an endpoint URL that downstream components call.
- **Evaluation & Benchmarking**: Integrates with **Harbor** via [`docs/harbor-integration.md`](https://github.com/marin-community/marin/blob/main/docs/harbor-integration.md), using the launcher at `experiments/evaluation/cli` to create isolated subprocesses that write results to Parquet files.
- **Telemetry & Logging** (`lib/finelog` and `lib/finestore`): Capture metrics and artefacts via a simple API used by both training and evaluation stages.

Each layer communicates via **typed JSON payloads**—for example, the policy JSON sent from Marin to Harbor (lines 27‑30 of [`docs/harbor-integration.md`](https://github.com/marin-community/marin/blob/main/docs/harbor-integration.md))—ensuring that external tools only need to understand the schema, not the internal implementation.

## Integration Patterns by Pipeline Layer

### Data Ingestion with Zephyr

To add support for a new storage backend or file format, implement a `Reader` subclass in `lib/zephyr/src/zephyr/readers/`. The base interface requires only that you yield `(key, bytes)` pairs. Zephyr already integrates with **fsspec**, so custom filesystems can often leverage existing URI handlers.

**Key integration points:**
- Implement the `Reader` interface from `zephyr.readers.base`.
- Register the class in a pipeline YAML under the `reader` key.
- Use `url_to_fs` from fsspec if implementing custom storage connectors.

### Model Training with Levanter

New model architectures integrate via `lib/levanter`. The process involves subclassing `BaseModel` and referencing the implementation in your training configuration.

**Key integration points:**
- Inherit from `BaseModel` in `lib/levanter/src/levanter/models/`.
- Register the model following the steps in [`lib/levanter/docs/dev/Port-Models.md`](https://github.com/marin-community/marin/blob/main/lib/levanter/docs/dev/Port-Models.md).
- Import and instantiate the model in [`experiments/simple_train_config.py`](https://github.com/marin-community/marin/blob/main/experiments/simple_train_config.py) before passing it to `train_lm()`.

### Evaluation with Harbor

Harbor benchmarks run through [`experiments/evaluation/cli.py`](https://github.com/marin-community/marin/blob/main/experiments/evaluation/cli.py), which creates an isolated subprocess communicating over an Iris endpoint. Results are serialized to Parquet files as defined in [`tests/evaluation/test_harbor_runner.py`](https://github.com/marin-community/marin/blob/main/tests/evaluation/test_harbor_runner.py).

**Key integration points:**
- Add Harbor policy YAML files under `experiments/evaluation/configs/harbor/`.
- Invoke the launcher using `python -m experiments.evaluation.cli launch`.
- Configure the handshake parameters documented in [`docs/harbor-integration.md`](https://github.com/marin-community/marin/blob/main/docs/harbor-integration.md).

## Step‑by‑Step Integration Workflow

Regardless of which layer you extend, the pattern remains consistent: **declare → wrap → wire → test**.

1. **Declare a new external dependency**
   
   Add the package to `marin.external_dependencies` as documented in [`config/external/README.md`](https://github.com/marin-community/marin/blob/main/config/external/README.md). Marin isolates the dependency in its own `uv` lock, preventing contamination of the core environment.

2. **Expose a thin wrapper**
   
   Implement a wrapper conforming to the expected interface (e.g., a `Reader` for Zephyr or a `ModelTrainer` for Levanter). Most wrappers live under `lib/<subproject>/src/` and are discovered via entry‑points defined in [`pyproject.toml`](https://github.com/marin-community/marin/blob/main/pyproject.toml).

3. **Wire the wrapper into the pipeline**
   
   - **Data readers:** Plug the class into the `DatasetConfig` used by Zephyr.
   - **Training:** Reference the registered model from a training config in [`experiments/simple_train_config.py`](https://github.com/marin-community/marin/blob/main/experiments/simple_train_config.py).
   - **Evaluation:** Add a Harbor policy YAML and invoke the launcher at `experiments/evaluation/cli`.

4. **Validate with the integration test**
   
   Run [`tests/integration_test.py`](https://github.com/marin-community/marin/blob/main/tests/integration_test.py) to verify that your connector functions correctly across all pipeline stages.

## Code Examples for Common Integrations

### Running a Harbor Benchmark from Python

The following script demonstrates how to programmatically launch a Harbor evaluation using the CLI module at `experiments/evaluation/cli`:

```python

# examples/harbor_demo.py

import subprocess

def launch_harbor(model: str, policy: str, limit: int = 2):
    cmd = [
        "uv", "run", "python", "-m", "experiments.evaluation.cli", "launch",
        "--model", model,
        "--harbor-config", f"experiments/evaluation/configs/harbor/{policy}.yaml",
        "--limit", str(limit),
    ]
    subprocess.check_call(cmd)

if __name__ == "__main__":
    launch_harbor("qwen3-8b", "aime-smoke")

```

### Registering a New Levanter Model

To add a custom architecture, subclass `BaseModel` and import it into your training script:

```python

# lib/levanter/src/levanter/models/custom.py

from levanter.models.base import BaseModel

class MyModel(BaseModel):
    def __init__(self, config):
        super().__init__(config)
        # model construction …

# In experiments/simple_train_config.py

from levanter.models.custom import MyModel

model = MyModel(config=my_config)
train_lm(model, dataset=..., ... )

```

### Adding a Zephyr Reader for Custom Storage

Implement the `Reader` interface to support non‑standard storage backends:

```python

# lib/zephyr/src/zephyr/readers/custom_fs.py

from zephyr.readers.base import Reader
from fsspec.core import url_to_fs

class CustomFSReader(Reader):
    def __init__(self, uri: str):
        self.fs, self.path = url_to_fs(uri)

    def __iter__(self):
        # yield (key, bytes) pairs

        for file in self.fs.ls(self.path):
            yield file, self.fs.open(file).read()

```

Reference the reader in a pipeline configuration:

```yaml

# experiments/pipelines/custom.yaml

reader:
  class: zephyr.readers.custom_fs.CustomFSReader
  args:
    uri: "s3://my-bucket/dataset/"

```

## Summary

- Marin’s architecture splits functionality into isolated sub‑libraries (`lib/zephyr`, `lib/levanter`, `lib/iris`, etc.), allowing you to integrate tools at specific layers without affecting others.
- All integrations follow the **declare → wrap → wire → test** pattern: declare dependencies in [`config/external/README.md`](https://github.com/marin-community/marin/blob/main/config/external/README.md), wrap external tools to conform to layer interfaces, wire them via YAML configs or entry points, and validate with [`tests/integration_test.py`](https://github.com/marin-community/marin/blob/main/tests/integration_test.py).
- Communication between layers uses **typed JSON payloads**, decoupling components and enabling painless substitution of implementations.
- Key files for integration include [`docs/harbor-integration.md`](https://github.com/marin-community/marin/blob/main/docs/harbor-integration.md) (evaluation), [`lib/levanter/docs/dev/Port-Models.md`](https://github.com/marin-community/marin/blob/main/lib/levanter/docs/dev/Port-Models.md) (training), and [`lib/zephyr/README.md`](https://github.com/marin-community/marin/blob/main/lib/zephyr/README.md) (data ingestion).

## Frequently Asked Questions

### How do I add a new external package dependency to Marin?

Add the package specification to `marin.external_dependencies` following the guidelines in [`config/external/README.md`](https://github.com/marin-community/marin/blob/main/config/external/README.md). Marin manages external dependencies in isolated `uv` locks, ensuring that new packages do not conflict with the core environment or other tools.

### Can I use a custom tokenizer that is not from Hugging Face?

Yes. While [`lib/marin/tools/get_hf_dataset_schema.py`](https://github.com/marin-community/marin/blob/main/lib/marin/tools/get_hf_dataset_schema.py) provides thin wrappers around Hugging Face tokenizers, the tokenization layer is interface‑based. You can implement a compatible wrapper following the patterns in [`tests/test_marin_tokenizer.py`](https://github.com/marin-community/marin/blob/main/tests/test_marin_tokenizer.py) and inject it into the pipeline configuration without modifying core Marin code.

### What is the correct way to launch a Harbor evaluation programmatically?

Use the CLI module at [`experiments/evaluation/cli.py`](https://github.com/marin-community/marin/blob/main/experiments/evaluation/cli.py) via subprocess or direct Python invocation. The launcher creates an isolated subprocess that communicates with Harbor over an Iris endpoint, then writes results to Parquet files. Refer to [`docs/harbor-integration.md`](https://github.com/marin-community/marin/blob/main/docs/harbor-integration.md) for the exact JSON schema expected in the policy configuration.

### Where should I implement a new dataset reader for an unsupported storage backend?

Implement the `Reader` interface in `lib/zephyr/src/zephyr/readers/`, inheriting from `zephyr.readers.base.Reader`. If the storage supports fsspec, leverage `url_to_fs` to minimize code. Register the reader class in your pipeline YAML under the `reader` key, specifying the full module path and initialization arguments.