# Marin Development Roadmap: LoRA, MoE, and Infrastructure Automation Plans

> Explore the Marin development roadmap focusing on LoRA, MoE, and infrastructure automation for efficient model fine-tuning and cost management. See active implementation details.

- Repository: [The Marin Project/marin](https://github.com/marin-community/marin)
- Tags: roadmap
- Published: 2026-08-29

---

**The marin development roadmap centers on three strategic pillars: parameter-efficient fine-tuning with LoRA, mixture-of-expert model support, and automated GCP cost management, with active implementation tracked through TODO markers in the Levanter and Zephyr source trees.**

Marin is a modular, high-performance ML pipeline built on top of the **Levanter** training library and the **Zephyr** dataset-processing framework. According to the marin development roadmap, the project is advancing scalable distributed training, high-throughput data ingestion, and DevOps automation for TPU clusters, evidenced by explicit development annotations throughout [`lib/levanter/src/levanter/trainer.py`](https://github.com/marin-community/marin/blob/main/lib/levanter/src/levanter/trainer.py) and related modules.

## Scalable Training and Model Architecture Enhancements

The core training infrastructure in [`lib/levanter/src/levanter/trainer.py`](https://github.com/marin-community/marin/blob/main/lib/levanter/src/levanter/trainer.py) already supports distributed JAX/TPU execution, but the roadmap outlines significant architectural expansions for emerging training paradigms.

### Parameter-Efficient Fine-Tuning with LoRA

The marin development roadmap prioritizes **LoRA** (Low-Rank Adaptation) implementation for parameter-efficient fine-tuning. The foundational adapter code lives in [`lib/levanter/src/levanter/adaptor/lora.py`](https://github.com/marin-community/marin/blob/main/lib/levanter/src/levanter/adaptor/lora.py), providing the scaffolding for injecting trainable low-rank matrices into frozen pretrained models.

The implementation targets support for rank-based adaptation across transformer attention layers, enabling researchers to fine-tune large language models with minimal compute overhead. This adapter system integrates seamlessly with the existing trainer loop while preserving the high-throughput characteristics essential for TPU execution.

### Mixture-of-Experts and Advanced Model Support

Support for **mixture-of-expert (MoE)** models represents another critical milestone. Current experimental implementations in [`experiments/june_tpu_67b_a2b/moe/model.py`](https://github.com/marin-community/marin/blob/main/experiments/june_tpu_67b_a2b/moe/model.py) demonstrate flexible routing mechanisms for expert parallelism.

The roadmap plans to expand MoE support to handle mismatched tokenizers and shard-wise streaming checkpoints, as tracked in [`levanter/main/train_lm.py`](https://github.com/marin-community/marin/blob/main/levanter/main/train_lm.py). Additional enhancements include flash-attention refinements and sliding-window attention support for architectures like Qwen, with active development indicated by TODO markers in [`levanter/src/levanter/models/qwen.py`](https://github.com/marin-community/marin/blob/main/levanter/src/levanter/models/qwen.py).

## Dataset Pipeline Improvements in Zephyr

The **Zephyr** dataset framework ([`lib/zephyr/src/zephyr/dataset.py`](https://github.com/marin-community/marin/blob/main/lib/zephyr/src/zephyr/dataset.py)) provides the foundation for Marin’s data ingestion, with the roadmap targeting high-impact enhancements for big-data workloads.

### Fast-Seek and Epoch Management

The marin development roadmap introduces **fast-seek and random-access** capabilities for large JSONL and Parquet sources, addressing current limitations in sequential data access. Planned improvements to **epoch reshuffling** and aggressive caching will enable more efficient training iterations over massive datasets.

These enhancements target the dataset abstraction layer, with specific TODO annotations indicating future support for deterministic shuffling across distributed training runs. The roadmap also emphasizes expanded structured-text converters for web tables, code snippets, and bio-chemical data formats, validated through the comprehensive test suite in `tests/transform/`.

## Infrastructure and DevOps Automation

Marin ships with a Pulumi-based infrastructure layer in `infra/` that manages TPU CI, X-Prof profiling, and evaluation dashboards. The roadmap significantly expands these operational capabilities for production deployments.

### Cost Management and Containerization

Automatic **GCP billing alerts** and **cost-reporting dashboards** are scheduled for rollout, as indicated by configuration files like [`ops-cost-report.yaml`](https://github.com/marin-community/marin/blob/main/ops-cost-report.yaml). The infrastructure team is developing **native Docker images** for TPU-ready training containers, with build specifications available in `docker/marin/Dockerfile.tpu-ci`.

CI/CD improvements focus on faster test selection and integrated lint-review bots, implemented in [`infra/ci/run_tests.py`](https://github.com/marin-community/marin/blob/main/infra/ci/run_tests.py). These automation efforts tighten the DevOps loop while preventing resource misconfiguration that leads to unexpected compute charges.

## Practical Code Examples from the Roadmap

The following snippets illustrate upcoming features currently in development:

### Launching a Configurable Training Sweep

```python

# experiments/tutorials/train_tiny_sweep.py

from levanter.trainer import Trainer
from levanter.config import Config

cfg = Config(
    model="gpt2_small",
    optimizer="adamw",
    num_train_steps=5000,
    batch_size=32,
    learning_rate=5e-5,
)
Trainer(cfg).run()

```

This example from [`experiments/tutorials/train_tiny_sweep.py`](https://github.com/marin-community/marin/blob/main/experiments/tutorials/train_tiny_sweep.py) demonstrates the new configuration system that supports hyperparameter sweeps across distributed TPU pods.

### Dataset Access with Future Epoch Shuffling

```python
from zephyr.dataset import Dataset

ds = Dataset.from_parquet("gs://my-bucket/train.parquet")

# Future version: ds = ds.shuffle_each_epoch(seed=42)

for batch in ds.batch(128):
    # feed batch to model

    pass

```

This pattern from [`lib/zephyr/src/zephyr/dataset.py`](https://github.com/marin-community/marin/blob/main/lib/zephyr/src/zephyr/dataset.py) shows the current API with commented roadmap features for deterministic epoch reshuffling.

### Applying LoRA Adapters

```python
from levanter.adaptor.lora import LoRAAdapter
from levanter.models.gpt2_hyena import GPT2Hyena

base = GPT2Hyena.from_pretrained("gpt2")
lora = LoRAAdapter(base, rank=8, modules_to_save=["c_attn"])
model = lora.apply()

```

This implementation from [`lib/levanter/src/levanter/adaptor/lora.py`](https://github.com/marin-community/marin/blob/main/lib/levanter/src/levanter/adaptor/lora.py) illustrates the upcoming LoRA integration for efficient model customization.

## Summary

- The marin development roadmap emphasizes **LoRA-based fine-tuning** and **mixture-of-expert architectures** to expand model capabilities while maintaining TPU efficiency.
- **Zephyr dataset improvements** target fast-seek random access, epoch reshuffling, and enhanced caching for massive-scale data pipelines.
- **Infrastructure automation** focuses on GCP cost management, TPU-optimized containers, and streamlined CI/CD through Pulumi-based configurations.
- Development progress is tracked through explicit TODO markers in critical files like [`lib/levanter/src/levanter/trainer.py`](https://github.com/marin-community/marin/blob/main/lib/levanter/src/levanter/trainer.py) and [`lib/zephyr/src/zephyr/dataset.py`](https://github.com/marin-community/marin/blob/main/lib/zephyr/src/zephyr/dataset.py).

## Frequently Asked Questions

### What is the current status of LoRA support in Marin?

LoRA support is currently in active development with foundational adapter code available in [`lib/levanter/src/levanter/adaptor/lora.py`](https://github.com/marin-community/marin/blob/main/lib/levanter/src/levanter/adaptor/lora.py). The implementation allows researchers to apply low-rank adaptations to frozen base models, significantly reducing memory requirements during fine-tuning. Full integration with the distributed training loop is targeted for the upcoming release cycle.

### How will dataset pipeline improvements affect large-scale training?

The planned **fast-seek** and **epoch reshuffling** capabilities in [`lib/zephyr/src/zephyr/dataset.py`](https://github.com/marin-community/marin/blob/main/lib/zephyr/src/zephyr/dataset.py) will eliminate I/O bottlenecks when processing terabyte-scale Parquet and JSONL datasets. These enhancements enable deterministic data access patterns across distributed TPU workers while maintaining the high-throughput ingestion rates required for efficient LLM training.

### What infrastructure changes are planned for cost management?

The roadmap includes automated **GCP billing alerts** and comprehensive **cost-reporting dashboards** managed through Pulumi configurations in the `infra/` directory. Additionally, standardized **Docker images** for TPU environments in `docker/marin/Dockerfile.tpu-ci` will reduce setup overhead and prevent resource misconfiguration that leads to unexpected compute charges.

### Where can I track progress on the marin development roadmap?

Development progress is publicly visible through TODO annotations scattered throughout the codebase, particularly in [`lib/levanter/src/levanter/trainer.py`](https://github.com/marin-community/marin/blob/main/lib/levanter/src/levanter/trainer.py), [`levanter/src/levanter/models/qwen.py`](https://github.com/marin-community/marin/blob/main/levanter/src/levanter/models/qwen.py), and [`lib/zephyr/src/zephyr/dataset.py`](https://github.com/marin-community/marin/blob/main/lib/zephyr/src/zephyr/dataset.py). The team also maintains experimental implementations in `experiments/` directories that preview upcoming stable features.