# How Marin Ensures Training Mixture Reproducibility: A Two‑Phase Pipeline

> Marin ensures training mixture reproducibility with a two phase pipeline. Discover how it guarantees identical logical mixtures for consistent results.

- Repository: [The Marin Project/marin](https://github.com/marin-community/marin)
- Tags: internals
- Published: 2026-09-10

---

**Marin guarantees training mixture reproducibility through a deterministic two‑phase pipeline that separates fingerprint‑time weight mapping from runtime cache resolution, ensuring identical logical mixtures even when underlying tokenized caches are rebuilt from scratch.**

Training mixture reproducibility is a foundational requirement for reproducible machine learning experiments. The Marin framework achieves this through a fingerprint‑driven architecture that decouples mixture configuration from physical data retrieval. By treating component identities and weight mappings as immutable during fingerprinting while deferring cache materialization to runtime, Marin ensures that re‑running an experiment always reconstructs the exact same training distribution.

## The Two‑Phase Deterministic Architecture

Marin’s approach hinges on a strict separation between **fingerprinting** (configuration validation) and **execution** (data materialization). This architecture prevents cache regeneration from altering the logical training mixture.

### Phase 1: Fingerprint‑Time Placeholder Components

During the fingerprint phase (`is_fingerprint=True`), Marin constructs a deterministic mixture map without accessing disk. In [`lib/marin/src/marin/experiment/data.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/experiment/data.py) at lines 91‑99, the `mixture()` function creates **placeholder components** that carry only cache identity metadata (`name@version`) and a constant format string. 

At this stage, the system:

- Parses training handles and validation handles (with weight 0) into a deterministic weight mapping (`handle‑name → weight`).
- Returns an `LmDataConfig` containing placeholder components rather than physical data paths.
- Bakes the weight configuration into the fingerprint itself, ensuring any weight change triggers a new build.

Because this phase never reads actual tokenized caches, deleting or regenerating cache files cannot alter the mixture’s logical structure.

### Phase 2: Runtime Resolution and Tokenizer Verification

When the pipeline actually executes (`is_fingerprint=False`), the `mixture()` function (lines 104‑118 in [`lib/marin/src/marin/experiment/data.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/experiment/data.py)) resolves each handle to a concrete `TokenizedCache`. This phase performs three critical validation steps:

1. **Cache Materialization**: Each handle is mapped to its physical `TokenizedCache` instance, supplying the tokenizer, format specification, and file paths.
2. **Tokenizer Homogeneity Check**: All tokenizers across components are verified as identical via `_verify_tokenizers_same()` in [`lib/marin/src/marin/processing/tokenize/data_configs.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/processing/tokenize/data_configs.py) (lines 76‑92). If tokenizers differ, the system raises a clear error preventing token ID mismatches.
3. **Component Assembly**: The final `LmDataConfig` contains concrete `DatasetComponent` objects, the verified tokenizer, and the preserved weight mapping from the fingerprint phase.

This strict verification ensures that **tokenizer consistency** is enforced at runtime, preventing subtle encoding variations that could otherwise corrupt reproducibility.

## Deterministic Shuffling and Seed Management

Beyond component assembly, training mixture reproducibility requires deterministic data ordering. Marin leverages Levanter’s `BlockShuffleConfig` (specifically the hierarchical block shuffle defined as `DEFAULT_LM_DATA_SHUFFLE` in [`lib/marin/src/marin/processing/tokenize/data_configs.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/processing/tokenize/data_configs.py), lines 9‑13).

Because the component list and weights are fixed during fingerprinting, the shuffle seed—derived from the fingerprint itself—yields the exact same block ordering on every run. This means:

- The shuffle configuration is stored in the `shuffle` field of `LmDataConfig`.
- Re‑running the same experiment produces identical batch sequences, eliminating nondeterminism in data presentation.

## Implementation Examples

The following examples demonstrate how to build reproducible mixtures using Marin’s API.

### Building a Mixture from Tokenized Caches

```python
from marin.experiment.data import mixture

# Assume tok_a and tok_b are TokenizedCache steps already defined

train_handles = {tok_a: 0.7, tok_b: 0.3}
validation_handles = (tok_a,)

def build_config(ctx):
    # Fingerprint phase: placeholder components + deterministic weight map

    lm_cfg = mixture(ctx, train=train_handles, validation=validation_handles)
    
    # Runtime phase: Marin automatically resolves caches and verifies tokenizers

    return lm_cfg

```

### Using the High‑Level Configuration Helper

```python
from marin.processing.tokenize.data_configs import lm_mixture_data_config
from marin.processing.tokenize.tokenize import TokenizeConfig

components = {
    "wiki": TokenizeConfig(...),   # produces a tokenized cache

    "books": TokenizeConfig(...),
}
weights = {"wiki": 0.6, "books": 0.4}

# Verifies tokenizers and builds LmDataConfig automatically

lm_data_cfg = lm_mixture_data_config(
    components=components,
    weights=weights,
    shuffle=True,  # hierarchical block shuffle (default)

)

```

## Summary

Marin ensures training mixture reproducibility through four core mechanisms:

- **Immutable Component Identity**: Fingerprinting captures `name@version` identifiers rather than physical paths, surviving cache deletion.
- **Deterministic Weight Mapping**: The weight map is constructed during fingerprinting and baked into the configuration hash.
- **Runtime Tokenizer Verification**: The `_verify_tokenizers_same()` function enforces identical tokenizers across all mixture components, preventing token ID drift.
- **Fingerprint‑Derived Seeding**: Shuffle seeds are derived from the fingerprint, guaranteeing identical data ordering across runs.

These mechanisms collectively ensure that experiments can be re‑executed months later with identical training mixtures, even if the underlying tokenized caches have been regenerated or moved.

## Frequently Asked Questions

### What happens if different mixture components use incompatible tokenizers?

Marin raises a runtime error during the resolution phase. The `_verify_tokenizers_same()` function in [`lib/marin/src/marin/processing/tokenize/data_configs.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/processing/tokenize/data_configs.py) (lines 76‑92) explicitly checks that all `TokenizedCache` instances share identical tokenizers (or equivalent ones via `_are_tokenizers_equivalent()`). This prevents subtle token ID mismatches that would otherwise break reproducibility.

### Does deleting tokenized caches affect training mixture reproducibility?

No. Because the fingerprint phase uses placeholder components containing only identity metadata (`name@version`) rather than physical paths, deleting and regenerating caches does not alter the logical mixture. When the pipeline re‑runs, Marin resolves the same handles to the newly materialized caches while preserving the original weight mapping and shuffle order.

### How does Marin ensure the same data shuffling across experiment restarts?

Marin derives the shuffle seed from the pipeline fingerprint itself. The default `BlockShuffleConfig` (referenced as `DEFAULT_LM_DATA_SHUFFLE` in [`lib/marin/src/marin/processing/tokenize/data_configs.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/processing/tokenize/data_configs.py)) uses this deterministic seed to generate identical block orderings. Since the component list and weights are fixed during fingerprinting, the shuffle sequence remains constant across runs.

### Can I modify mixture weights without changing the experiment fingerprint?

No. The weight mapping is constructed during the fingerprint phase (lines 91‑99 in [`lib/marin/src/marin/experiment/data.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/experiment/data.py)) and constitutes part of the configuration identity. Changing any training or validation weight produces a different fingerprint, triggering a new build. This design ensures that weight changes are always tracked and reproducible.