# Development Roadmap for Marin: Scalable ML Pipeline Framework Goals and Milestones

> Explore the Marin development roadmap focusing on scalable ML pipelines. Discover our goals for multi-node JAX integration, unified observability, and expanded data connectors. Contribute to the future of Marin.

- Repository: [The Marin Project/marin](https://github.com/marin-community/marin)
- Tags: getting-started
- Published: 2026-08-27

---

**The Marin development roadmap prioritizes three strategic pillars: scalable distributed execution through multi-node JAX Levanter integration, a production-grade unified observability dashboard, and an expanding ecosystem of data connectors for third-party tools like MLflow and Weights & Biases.**

Marin is an open-source, extensible pipeline framework maintained by `marin-community/marin` that orchestrates data processing, model training, and inference across diverse compute backends. The project roadmap is actively managed through GitHub issues and project boards, with development driven by community proposals and the core team's vision for cloud-native machine-learning infrastructure.

## Core Development Pillars

The roadmap is organized around three overarching goals that guide feature prioritization and architectural decisions.

### Scalable Distributed Execution

The primary focus is expanding native support for large-scale training and inference workloads. This includes implementing robust multi-node orchestration capabilities, automatic model sharding strategies, and deeper integration with cloud-native batch-processing services. According to the architecture overview in [`README.md`](https://github.com/marin-community/marin/blob/main/README.md), the team is specifically working to enable seamless **FSDP (Fully Sharded Data Parallel)** and sharding across clusters to improve training throughput for large language models using JAX Levanter.

### User-Facing Observability and UI

Marin aims to deliver a production-grade dashboard for visualizing pipeline graphs, resource utilization, and experiment metadata. This initiative builds upon the existing `marin.inference.dashboard` package located at [`lib/marin/src/marin/inference/dashboard/README.md`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/inference/dashboard/README.md). The planned unified dashboard will consolidate inference monitoring into a full-stack UI that displays pipeline topology, real-time logs, and performance metrics through a RESTful API.

### Extensible Data Connectors and Ecosystem Integration

The framework is expanding its catalog of built-in data-source adapters to simplify ingestion from public datasets and private cloud storage. Current priorities include first-class support for StackExchange dumps, Common Crawl, Google Cloud Storage, and Pub/Sub streams. Additionally, the roadmap emphasizes simplifying integration with third-party MLOps tools such as MLflow, Weights & Biases, and Ray Serve.

## Key Roadmap Milestones and Source Files

Specific technical initiatives are documented across the repository's README files and issue tracker. The following milestones represent active development targets:

- **Multi-node JAX/Levanter Scaling**: Enable seamless FSDP and tensor sharding across compute clusters, documented in the main [`README.md`](https://github.com/marin-community/marin/blob/main/README.md) architecture section.
- **Unified Dashboard UI**: Consolidate the existing inference dashboard into a comprehensive interface showing pipeline topology and metrics, as detailed in [`lib/marin/src/marin/inference/dashboard/README.md`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/inference/dashboard/README.md).
- **Expanded Data-Kit Connectors**: Add first-class adapters for StackExchange dumps and Common Crawl, following the pattern established in [`lib/marin/src/marin/datakit/download/stackexchange/README.md`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/datakit/download/stackexchange/README.md).
- **Improved Job Orchestration (Iris)**: Refine the Iris job runner to support auto-scaling on GCP and CoreWeave, with richer failure recovery mechanisms and a higher-level API for pipeline developers, documented in [`lib/iris/README.md`](https://github.com/marin-community/marin/blob/main/lib/iris/README.md).
- **CI and Testing Enhancements**: Strengthen continuous integration with stricter type-checking, slope-test rejection policies, and performance benchmarks to prevent regressions, as defined in [`TESTING.md`](https://github.com/marin-community/marin/blob/main/TESTING.md).
- **Documentation and Community Tooling**: Produce step-by-step guides, example notebooks, and auto-generated API documentation to lower contributor barriers.

## Implementation Examples

The following code snippets demonstrate how developers interact with components targeted for roadmap enhancements.

### Submitting Distributed Jobs with Iris

The Iris job runner, central to the scaling roadmap, manages pipeline execution on cloud clusters:

```python
from marin.processing import ClassificationPipeline
from iris.runner import IrisJobRunner

# Define a pipeline that loads data, runs a classifier, and writes results

pipeline = ClassificationPipeline(
    input_path="gs://my-bucket/data",
    model_name="large-bert",
    output_path="gs://my-bucket/results",
)

# Run the pipeline on a managed Iris cluster

runner = IrisJobRunner(cluster="my-gcp-cluster")
job = runner.submit(pipeline)
print(f"Submitted job {job.id}, monitor at {job.dashboard_url}")

```

### Querying the Unified Dashboard API

The upcoming dashboard will expose REST endpoints for pipeline monitoring:

```python
import requests

resp = requests.get(
    "https://dashboard.marin.ai/api/v1/pipelines",
    headers={"Authorization": "Bearer <TOKEN>"}
)
for pipeline in resp.json():
    print(pipeline["name"], pipeline["status"], pipeline["last_updated"])

```

### Using Expanded DataKit Connectors

The StackExchange downloader exemplifies the connector pattern planned for broader dataset support:

```python
from marin.datakit.download.stackexchange import StackExchangeDownloader

downloader = StackExchangeDownloader(
    site="stackoverflow",
    start_date="2024-01-01",
    end_date="2024-01-31",
    destination="/tmp/so_jan2024",
)
downloader.run()
print("Download complete")

```

## Summary

- **Marin's development roadmap** centers on distributed scalability, observability tooling, and ecosystem integration.
- **Multi-node training** improvements target JAX Levanter with FSDP support for large language models.
- **Dashboard evolution** will transform the existing `marin.inference.dashboard` into a unified, production-grade monitoring interface.
- **Data connectivity** is expanding through standardized connectors for StackExchange, cloud storage, and streaming sources.
- **Iris job orchestration** is gaining auto-scaling capabilities and richer failure recovery for cloud-native deployments.
- **Quality assurance** is being reinforced through enhanced CI pipelines and stricter testing policies defined in [`TESTING.md`](https://github.com/marin-community/marin/blob/main/TESTING.md).

## Frequently Asked Questions

### What is the current focus of the Marin development roadmap?

The current focus involves three parallel tracks: implementing multi-node distributed training support via JAX Levanter and FSDP, building a unified observability dashboard that consolidates the existing `marin/inference/dashboard` functionality, and expanding the DataKit connector ecosystem to support additional public datasets and cloud storage backends.

### How does Marin plan to improve distributed training capabilities?

According to the repository's architecture documentation, Marin is enhancing its Iris job runner to support auto-scaling clusters on GCP and CoreWeave while integrating tighter with JAX Levanter for automatic sharding. These improvements enable seamless FSDP across multiple nodes, significantly improving throughput for large-scale model training workloads.

### What new data sources will Marin support?

The roadmap includes first-class adapters for StackExchange data dumps, Common Crawl, Google Cloud Storage, and Pub/Sub streams. These connectors follow the implementation pattern established in [`lib/marin/src/marin/datakit/download/stackexchange/README.md`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/datakit/download/stackexchange/README.md), providing standardized interfaces for both public datasets and private cloud storage backends.

### How can I contribute to the Marin roadmap?

Contributors can participate by commenting on open GitHub issues, submitting new feature proposals, or contributing code to the prioritized milestones. The project encourages community input through GitHub issues and project boards, with specific contribution guidelines documented in the main [`README.md`](https://github.com/marin-community/marin/blob/main/README.md) and testing requirements detailed in [`TESTING.md`](https://github.com/marin-community/marin/blob/main/TESTING.md).