How Marin Ensures Training Mixture Reproducibility: A Two‑Phase Pipeline

Marin guarantees training mixture reproducibility through a deterministic two‑phase pipeline that separates fingerprint‑time weight mapping from runtime cache resolution, ensuring identical logical mixtures even when underlying tokenized caches are rebuilt from scratch.

Training mixture reproducibility is a foundational requirement for reproducible machine learning experiments. The Marin framework achieves this through a fingerprint‑driven architecture that decouples mixture configuration from physical data retrieval. By treating component identities and weight mappings as immutable during fingerprinting while deferring cache materialization to runtime, Marin ensures that re‑running an experiment always reconstructs the exact same training distribution.

The Two‑Phase Deterministic Architecture

Marin’s approach hinges on a strict separation between fingerprinting (configuration validation) and execution (data materialization). This architecture prevents cache regeneration from altering the logical training mixture.

Phase 1: Fingerprint‑Time Placeholder Components

During the fingerprint phase (is_fingerprint=True), Marin constructs a deterministic mixture map without accessing disk. In lib/marin/src/marin/experiment/data.py at lines 91‑99, the mixture() function creates placeholder components that carry only cache identity metadata (name@version) and a constant format string.

At this stage, the system:

  • Parses training handles and validation handles (with weight 0) into a deterministic weight mapping (handle‑name → weight).
  • Returns an LmDataConfig containing placeholder components rather than physical data paths.
  • Bakes the weight configuration into the fingerprint itself, ensuring any weight change triggers a new build.

Because this phase never reads actual tokenized caches, deleting or regenerating cache files cannot alter the mixture’s logical structure.

Phase 2: Runtime Resolution and Tokenizer Verification

When the pipeline actually executes (is_fingerprint=False), the mixture() function (lines 104‑118 in lib/marin/src/marin/experiment/data.py) resolves each handle to a concrete TokenizedCache. This phase performs three critical validation steps:

  1. Cache Materialization: Each handle is mapped to its physical TokenizedCache instance, supplying the tokenizer, format specification, and file paths.
  2. Tokenizer Homogeneity Check: All tokenizers across components are verified as identical via _verify_tokenizers_same() in lib/marin/src/marin/processing/tokenize/data_configs.py (lines 76‑92). If tokenizers differ, the system raises a clear error preventing token ID mismatches.
  3. Component Assembly: The final LmDataConfig contains concrete DatasetComponent objects, the verified tokenizer, and the preserved weight mapping from the fingerprint phase.

This strict verification ensures that tokenizer consistency is enforced at runtime, preventing subtle encoding variations that could otherwise corrupt reproducibility.

Deterministic Shuffling and Seed Management

Beyond component assembly, training mixture reproducibility requires deterministic data ordering. Marin leverages Levanter’s BlockShuffleConfig (specifically the hierarchical block shuffle defined as DEFAULT_LM_DATA_SHUFFLE in lib/marin/src/marin/processing/tokenize/data_configs.py, lines 9‑13).

Because the component list and weights are fixed during fingerprinting, the shuffle seed—derived from the fingerprint itself—yields the exact same block ordering on every run. This means:

  • The shuffle configuration is stored in the shuffle field of LmDataConfig.
  • Re‑running the same experiment produces identical batch sequences, eliminating nondeterminism in data presentation.

Implementation Examples

The following examples demonstrate how to build reproducible mixtures using Marin’s API.

Building a Mixture from Tokenized Caches

from marin.experiment.data import mixture

# Assume tok_a and tok_b are TokenizedCache steps already defined

train_handles = {tok_a: 0.7, tok_b: 0.3}
validation_handles = (tok_a,)

def build_config(ctx):
    # Fingerprint phase: placeholder components + deterministic weight map

    lm_cfg = mixture(ctx, train=train_handles, validation=validation_handles)
    
    # Runtime phase: Marin automatically resolves caches and verifies tokenizers

    return lm_cfg

Using the High‑Level Configuration Helper

from marin.processing.tokenize.data_configs import lm_mixture_data_config
from marin.processing.tokenize.tokenize import TokenizeConfig

components = {
    "wiki": TokenizeConfig(...),   # produces a tokenized cache

    "books": TokenizeConfig(...),
}
weights = {"wiki": 0.6, "books": 0.4}

# Verifies tokenizers and builds LmDataConfig automatically

lm_data_cfg = lm_mixture_data_config(
    components=components,
    weights=weights,
    shuffle=True,  # hierarchical block shuffle (default)

)

Summary

Marin ensures training mixture reproducibility through four core mechanisms:

  • Immutable Component Identity: Fingerprinting captures name@version identifiers rather than physical paths, surviving cache deletion.
  • Deterministic Weight Mapping: The weight map is constructed during fingerprinting and baked into the configuration hash.
  • Runtime Tokenizer Verification: The _verify_tokenizers_same() function enforces identical tokenizers across all mixture components, preventing token ID drift.
  • Fingerprint‑Derived Seeding: Shuffle seeds are derived from the fingerprint, guaranteeing identical data ordering across runs.

These mechanisms collectively ensure that experiments can be re‑executed months later with identical training mixtures, even if the underlying tokenized caches have been regenerated or moved.

Frequently Asked Questions

What happens if different mixture components use incompatible tokenizers?

Marin raises a runtime error during the resolution phase. The _verify_tokenizers_same() function in lib/marin/src/marin/processing/tokenize/data_configs.py (lines 76‑92) explicitly checks that all TokenizedCache instances share identical tokenizers (or equivalent ones via _are_tokenizers_equivalent()). This prevents subtle token ID mismatches that would otherwise break reproducibility.

Does deleting tokenized caches affect training mixture reproducibility?

No. Because the fingerprint phase uses placeholder components containing only identity metadata (name@version) rather than physical paths, deleting and regenerating caches does not alter the logical mixture. When the pipeline re‑runs, Marin resolves the same handles to the newly materialized caches while preserving the original weight mapping and shuffle order.

How does Marin ensure the same data shuffling across experiment restarts?

Marin derives the shuffle seed from the pipeline fingerprint itself. The default BlockShuffleConfig (referenced as DEFAULT_LM_DATA_SHUFFLE in lib/marin/src/marin/processing/tokenize/data_configs.py) uses this deterministic seed to generate identical block orderings. Since the component list and weights are fixed during fingerprinting, the shuffle sequence remains constant across runs.

Can I modify mixture weights without changing the experiment fingerprint?

No. The weight mapping is constructed during the fingerprint phase (lines 91‑99 in lib/marin/src/marin/experiment/data.py) and constitutes part of the configuration identity. Changing any training or validation weight produces a different fingerprint, triggering a new build. This design ensures that weight changes are always tracked and reproducible.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →