How to Integrate Marin with Other Tools: A Complete Guide to Modular Pipeline Extensions
Integrating Marin with external tools involves declaring dependencies in config/external/README.md, creating thin wrapper classes that conform to layer-specific interfaces, and wiring them into the pipeline via configuration files or entry points.
Marin is architected as a modular pipeline that stitches together specialized sub‑libraries, making it straightforward to plug in new tools or replace existing components without modifying core code. Whether you need to add a custom data reader, register a new model architecture, or benchmark against an external evaluation suite, Marin’s layered design enforces clean separation through typed JSON payloads and standardized entry points. This guide walks through the practical steps to integrate tools at every layer of the Marin stack, from data ingestion to distributed training and evaluation.
Understanding Marin’s Modular Architecture
Marin decomposes the machine learning lifecycle into discrete layers, each housed in its own sub‑library under lib/. This separation allows you to swap implementations without cascading changes:
- Data Ingestion & Transformation (
lib/zephyr): Handles dataset readers, writers, and shuffling, with built‑in fsspec integration for GCS and S3 storage (Zephyr README – fsspec integration). - Data Cleaning & Filtering: Uses external utilities such as fastText and trafilatura via the
transformpackage; the pipeline is agnostic to the specific HTML‑to‑text converter or classifier used. - Tokenization (
lib/marin/tools/get_hf_dataset_schema.py): Thin wrappers around Hugging Face tokenizers, allowing any HF tokenizer to be dropped in directly. - Model Training (
lib/levanter): Powers the training loop, invoked fromexperiments/simple_train_config.py, with support for custom model registration. - Job Orchestration (
lib/iris): Provides the distributed execution engine; jobs receive an endpoint URL that downstream components call. - Evaluation & Benchmarking: Integrates with Harbor via
docs/harbor-integration.md, using the launcher atexperiments/evaluation/clito create isolated subprocesses that write results to Parquet files. - Telemetry & Logging (
lib/finelogandlib/finestore): Capture metrics and artefacts via a simple API used by both training and evaluation stages.
Each layer communicates via typed JSON payloads—for example, the policy JSON sent from Marin to Harbor (lines 27‑30 of docs/harbor-integration.md)—ensuring that external tools only need to understand the schema, not the internal implementation.
Integration Patterns by Pipeline Layer
Data Ingestion with Zephyr
To add support for a new storage backend or file format, implement a Reader subclass in lib/zephyr/src/zephyr/readers/. The base interface requires only that you yield (key, bytes) pairs. Zephyr already integrates with fsspec, so custom filesystems can often leverage existing URI handlers.
Key integration points:
- Implement the
Readerinterface fromzephyr.readers.base. - Register the class in a pipeline YAML under the
readerkey. - Use
url_to_fsfrom fsspec if implementing custom storage connectors.
Model Training with Levanter
New model architectures integrate via lib/levanter. The process involves subclassing BaseModel and referencing the implementation in your training configuration.
Key integration points:
- Inherit from
BaseModelinlib/levanter/src/levanter/models/. - Register the model following the steps in
lib/levanter/docs/dev/Port-Models.md. - Import and instantiate the model in
experiments/simple_train_config.pybefore passing it totrain_lm().
Evaluation with Harbor
Harbor benchmarks run through experiments/evaluation/cli.py, which creates an isolated subprocess communicating over an Iris endpoint. Results are serialized to Parquet files as defined in tests/evaluation/test_harbor_runner.py.
Key integration points:
- Add Harbor policy YAML files under
experiments/evaluation/configs/harbor/. - Invoke the launcher using
python -m experiments.evaluation.cli launch. - Configure the handshake parameters documented in
docs/harbor-integration.md.
Step‑by‑Step Integration Workflow
Regardless of which layer you extend, the pattern remains consistent: declare → wrap → wire → test.
-
Declare a new external dependency
Add the package to
marin.external_dependenciesas documented inconfig/external/README.md. Marin isolates the dependency in its ownuvlock, preventing contamination of the core environment. -
Expose a thin wrapper
Implement a wrapper conforming to the expected interface (e.g., a
Readerfor Zephyr or aModelTrainerfor Levanter). Most wrappers live underlib/<subproject>/src/and are discovered via entry‑points defined inpyproject.toml. -
Wire the wrapper into the pipeline
- Data readers: Plug the class into the
DatasetConfigused by Zephyr. - Training: Reference the registered model from a training config in
experiments/simple_train_config.py. - Evaluation: Add a Harbor policy YAML and invoke the launcher at
experiments/evaluation/cli.
- Data readers: Plug the class into the
-
Validate with the integration test
Run
tests/integration_test.pyto verify that your connector functions correctly across all pipeline stages.
Code Examples for Common Integrations
Running a Harbor Benchmark from Python
The following script demonstrates how to programmatically launch a Harbor evaluation using the CLI module at experiments/evaluation/cli:
# examples/harbor_demo.py
import subprocess
def launch_harbor(model: str, policy: str, limit: int = 2):
cmd = [
"uv", "run", "python", "-m", "experiments.evaluation.cli", "launch",
"--model", model,
"--harbor-config", f"experiments/evaluation/configs/harbor/{policy}.yaml",
"--limit", str(limit),
]
subprocess.check_call(cmd)
if __name__ == "__main__":
launch_harbor("qwen3-8b", "aime-smoke")
Registering a New Levanter Model
To add a custom architecture, subclass BaseModel and import it into your training script:
# lib/levanter/src/levanter/models/custom.py
from levanter.models.base import BaseModel
class MyModel(BaseModel):
def __init__(self, config):
super().__init__(config)
# model construction …
# In experiments/simple_train_config.py
from levanter.models.custom import MyModel
model = MyModel(config=my_config)
train_lm(model, dataset=..., ... )
Adding a Zephyr Reader for Custom Storage
Implement the Reader interface to support non‑standard storage backends:
# lib/zephyr/src/zephyr/readers/custom_fs.py
from zephyr.readers.base import Reader
from fsspec.core import url_to_fs
class CustomFSReader(Reader):
def __init__(self, uri: str):
self.fs, self.path = url_to_fs(uri)
def __iter__(self):
# yield (key, bytes) pairs
for file in self.fs.ls(self.path):
yield file, self.fs.open(file).read()
Reference the reader in a pipeline configuration:
# experiments/pipelines/custom.yaml
reader:
class: zephyr.readers.custom_fs.CustomFSReader
args:
uri: "s3://my-bucket/dataset/"
Summary
- Marin’s architecture splits functionality into isolated sub‑libraries (
lib/zephyr,lib/levanter,lib/iris, etc.), allowing you to integrate tools at specific layers without affecting others. - All integrations follow the declare → wrap → wire → test pattern: declare dependencies in
config/external/README.md, wrap external tools to conform to layer interfaces, wire them via YAML configs or entry points, and validate withtests/integration_test.py. - Communication between layers uses typed JSON payloads, decoupling components and enabling painless substitution of implementations.
- Key files for integration include
docs/harbor-integration.md(evaluation),lib/levanter/docs/dev/Port-Models.md(training), andlib/zephyr/README.md(data ingestion).
Frequently Asked Questions
How do I add a new external package dependency to Marin?
Add the package specification to marin.external_dependencies following the guidelines in config/external/README.md. Marin manages external dependencies in isolated uv locks, ensuring that new packages do not conflict with the core environment or other tools.
Can I use a custom tokenizer that is not from Hugging Face?
Yes. While lib/marin/tools/get_hf_dataset_schema.py provides thin wrappers around Hugging Face tokenizers, the tokenization layer is interface‑based. You can implement a compatible wrapper following the patterns in tests/test_marin_tokenizer.py and inject it into the pipeline configuration without modifying core Marin code.
What is the correct way to launch a Harbor evaluation programmatically?
Use the CLI module at experiments/evaluation/cli.py via subprocess or direct Python invocation. The launcher creates an isolated subprocess that communicates with Harbor over an Iris endpoint, then writes results to Parquet files. Refer to docs/harbor-integration.md for the exact JSON schema expected in the policy configuration.
Where should I implement a new dataset reader for an unsupported storage backend?
Implement the Reader interface in lib/zephyr/src/zephyr/readers/, inheriting from zephyr.readers.base.Reader. If the storage supports fsspec, leverage url_to_fs to minimize code. Register the reader class in your pipeline YAML under the reader key, specifying the full module path and initialization arguments.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →