How to Integrate Marin with Other Services: A Step-by-Step Guide
You integrate Marin with external services by registering them as external dependencies in config/external/, implementing adapter classes that conform to Marin protocols like InferenceEngine or IngestionManifest, and injecting them into experiments via the SweepConfig API.
The Marin framework from the marin-community/marin repository is built as a modular pipeline system that separates third-party tools from core logic through thin adapter layers. This architecture enables you to connect arbitrary services—whether data labeling platforms such as Harbor, inference engines like VLLM, or custom HTTP APIs—by declaring dependencies in the project's configuration and implementing standardized Python interfaces.
Understanding Marin's Three-Layer Architecture
Marin organizes integration points into three logical layers that determine where your code lives:
| Layer | Responsibility | Integration Point |
|---|---|---|
| External Dependencies | Describes third-party tools and pins their versions via lock files in config/external/. |
Add services by creating isolated packages and running config/update-external.py to regenerate lib/marin/src/marin/external_dependencies.py. |
| Core Framework | Provides generic abstractions for experiments, execution, and evaluation under lib/marin/src/marin/. |
Hook into pipelines using marin.experiment, marin.execution, or marin.transform APIs. |
| Service Adapters | Thin wrappers that translate external APIs into Marin's interfaces. | Implement protocols in lib/marin/src/marin/inference/ or lib/marin/src/marin/publish/ (see vllm_backend.py for reference). |
Step 1: Register the Service as an External Dependency
Before writing code, declare the service as an external dependency so Marin can track its version and artifacts.
- Create an isolated project that provides a Python wheel or Docker image for the service.
- Add a lock file under
config/external/<service>/(for example,config/external/yourservice/uv.lock). - Run the helper script to generate constants:
uv run config/update-external.py <service>
This updates lib/marin/src/marin/external_dependencies.py with a constant like YOURSERVICE = ExternalDependency(...), making the service available throughout the codebase.
Step 2: Implement a Service Adapter
Adapters must implement the protocol expected by the specific layer they extend. Marin provides three primary extension points for integrating external services:
- Data ingestion – Subclass
marin.datakit.ingestion_manifest.IngestionManifestto yieldmarin.datakit.DatasetRecordobjects. - Inference – Implement a class exposing
predict(prompt: str) -> strand inherit frommarin.inference.base.InferenceEngine. - Evaluation – Create a
marin.evaluation.runner.EvaluationRunnerthat loads the service's test suite and returns per-example metrics.
Example: Wrapping an HTTP Inference Service
The following adapter demonstrates how to integrate an external HTTP-based generation service by implementing the InferenceEngine protocol:
import requests
from marin.inference.base import InferenceEngine
class HttpInferenceEngine(InferenceEngine):
"""Thin wrapper around an external HTTP inference service."""
name = "http_service"
def __init__(self, url: str, api_key: str | None = None):
self.url = url
self.headers = {"Authorization": f"Bearer {api_key}"} if api_key else {}
def predict(self, prompt: str) -> str:
resp = requests.post(
self.url,
json={"prompt": prompt},
headers=self.headers,
)
resp.raise_for_status()
return resp.json()["output"]
Reference the VLLM backend implementation in lib/marin/src/marin/inference/vllm_backend.py for a complete template of an inference adapter that handles batching and error retry logic.
Example: Adding a Custom Data Source
To integrate a service that provides training data, implement the ingestion protocol:
from marin.datakit.ingestion_manifest import IngestionManifest, DatasetRecord
class MyCSVSource(IngestionManifest):
def __init__(self, path: str):
self.path = path
def iter_records(self):
import csv
with open(self.path, newline="") as f:
for row in csv.DictReader(f):
yield DatasetRecord(
id=row["id"],
text=row["text"],
metadata={"source": "my_csv"},
)
Step 3: Wire the Adapter into an Experiment
Use the high-level experiment API to compose your new service into a training or evaluation run. The SweepConfig class accepts your adapter via the inference_engine parameter:
from marin.experiment import train_lm
from marin.experiment.sweep import SweepConfig
from lib.marin.src.marin.inference.http_service import HttpInferenceEngine
sweep_cfg = SweepConfig(
name="my_experiment",
trainer=..., # Existing trainer config
evaluator=..., # Existing evaluator config
inference_engine=HttpInferenceEngine(
endpoint="https://api.yourservice.com",
api_key="${MY_SERVICE_KEY}"
)
)
train_lm(sweep_cfg)
The train_lm function in lib/marin/src/marin/experiment/train.py orchestrates the pipeline, ensuring your adapter is invoked during the inference phase of the experiment.
Step 4: Deploy the Service with Infrastructure as Code (Optional)
Marin includes a Pulumi stack in infra/pulumi/ for provisioning cloud resources. You can reference the generated dependency constant to deploy your service:
resource "google_cloud_run_service" "yourservice" {
name = "yourservice"
location = var.region
template {
spec {
containers {
image = external_dependencies.YOURSERVICE.docker_image
}
}
}
}
This ensures the deployed infrastructure uses the exact Docker image version pinned in your external_dependencies.py.
Step 5: Verify with Integration Tests
Validate that your adapter correctly implements the required protocols by running Marin's built-in integration test:
PYTHONPATH=. python tests/integration_test.py --prefix var
This test exercises the full pipeline across all three layers—dependency resolution, adapter execution, and experiment orchestration—giving confidence that the new service can be used in production pipelines.
Summary
- Register external services by adding lock files to
config/external/<service>/and runningconfig/update-external.pyto regeneratelib/marin/src/marin/external_dependencies.py. - Implement adapters by subclassing
InferenceEngine,IngestionManifest, orEvaluationRunnerto translate external APIs into Marin protocols. - Integrate adapters into experiments by passing them to
SweepConfigand invokingtrain_lm. - Deploy services using the Pulumi configurations in
infra/pulumi/which reference your pinned dependency constants. - Test integrations end-to-end using
tests/integration_test.pyto ensure compatibility across the entire stack.
Frequently Asked Questions
What protocol must I implement to add a custom inference service to Marin?
You must implement the InferenceEngine protocol defined in lib/marin/src/marin/inference/base.py. At minimum, your class must provide a predict(self, prompt: str) -> str method and a name attribute. For async or batched services, reference lib/marin/src/marin/inference/vllm_backend.py for implementation patterns handling concurrent requests.
How do I update the external_dependencies.py file after modifying a service?
Run uv run config/update-external.py <service_name> from the repository root. This script reads the lock files and metadata in config/external/<service_name>/ and regenerates the constants in lib/marin/src/marin/external_dependencies.py, ensuring version consistency across the project.
Can I integrate services that only provide a Docker image without a Python SDK?
Yes. When registering the service in config/external/, specify the Docker image URI in the service's configuration. The generated ExternalDependency constant will expose the image reference, which you can use in your Pulumi deployment scripts or pass to containerized adapter implementations that communicate via HTTP or gRPC.
Where should I store API keys for external services integrated with Marin?
Store sensitive credentials in environment variables or your deployment platform's secret manager (such as Google Secret Manager or AWS Secrets Manager), then reference them in your adapter's __init__ method using placeholders like "${MY_SERVICE_KEY}". Never commit API keys to the repository; instead, inject them at runtime through the experiment configuration or infrastructure templates.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →