How Marin Ensures Training Mixture Reproducibility: A Two‑Phase Pipeline
Marin guarantees training mixture reproducibility through a deterministic two‑phase pipeline that separates fingerprint‑time weight mapping from runtime cache resolution, ensuring identical logical mixtures even when underlying tokenized caches are rebuilt from scratch.
Training mixture reproducibility is a foundational requirement for reproducible machine learning experiments. The Marin framework achieves this through a fingerprint‑driven architecture that decouples mixture configuration from physical data retrieval. By treating component identities and weight mappings as immutable during fingerprinting while deferring cache materialization to runtime, Marin ensures that re‑running an experiment always reconstructs the exact same training distribution.
The Two‑Phase Deterministic Architecture
Marin’s approach hinges on a strict separation between fingerprinting (configuration validation) and execution (data materialization). This architecture prevents cache regeneration from altering the logical training mixture.
Phase 1: Fingerprint‑Time Placeholder Components
During the fingerprint phase (is_fingerprint=True), Marin constructs a deterministic mixture map without accessing disk. In lib/marin/src/marin/experiment/data.py at lines 91‑99, the mixture() function creates placeholder components that carry only cache identity metadata (name@version) and a constant format string.
At this stage, the system:
- Parses training handles and validation handles (with weight 0) into a deterministic weight mapping (
handle‑name → weight). - Returns an
LmDataConfigcontaining placeholder components rather than physical data paths. - Bakes the weight configuration into the fingerprint itself, ensuring any weight change triggers a new build.
Because this phase never reads actual tokenized caches, deleting or regenerating cache files cannot alter the mixture’s logical structure.
Phase 2: Runtime Resolution and Tokenizer Verification
When the pipeline actually executes (is_fingerprint=False), the mixture() function (lines 104‑118 in lib/marin/src/marin/experiment/data.py) resolves each handle to a concrete TokenizedCache. This phase performs three critical validation steps:
- Cache Materialization: Each handle is mapped to its physical
TokenizedCacheinstance, supplying the tokenizer, format specification, and file paths. - Tokenizer Homogeneity Check: All tokenizers across components are verified as identical via
_verify_tokenizers_same()inlib/marin/src/marin/processing/tokenize/data_configs.py(lines 76‑92). If tokenizers differ, the system raises a clear error preventing token ID mismatches. - Component Assembly: The final
LmDataConfigcontains concreteDatasetComponentobjects, the verified tokenizer, and the preserved weight mapping from the fingerprint phase.
This strict verification ensures that tokenizer consistency is enforced at runtime, preventing subtle encoding variations that could otherwise corrupt reproducibility.
Deterministic Shuffling and Seed Management
Beyond component assembly, training mixture reproducibility requires deterministic data ordering. Marin leverages Levanter’s BlockShuffleConfig (specifically the hierarchical block shuffle defined as DEFAULT_LM_DATA_SHUFFLE in lib/marin/src/marin/processing/tokenize/data_configs.py, lines 9‑13).
Because the component list and weights are fixed during fingerprinting, the shuffle seed—derived from the fingerprint itself—yields the exact same block ordering on every run. This means:
- The shuffle configuration is stored in the
shufflefield ofLmDataConfig. - Re‑running the same experiment produces identical batch sequences, eliminating nondeterminism in data presentation.
Implementation Examples
The following examples demonstrate how to build reproducible mixtures using Marin’s API.
Building a Mixture from Tokenized Caches
from marin.experiment.data import mixture
# Assume tok_a and tok_b are TokenizedCache steps already defined
train_handles = {tok_a: 0.7, tok_b: 0.3}
validation_handles = (tok_a,)
def build_config(ctx):
# Fingerprint phase: placeholder components + deterministic weight map
lm_cfg = mixture(ctx, train=train_handles, validation=validation_handles)
# Runtime phase: Marin automatically resolves caches and verifies tokenizers
return lm_cfg
Using the High‑Level Configuration Helper
from marin.processing.tokenize.data_configs import lm_mixture_data_config
from marin.processing.tokenize.tokenize import TokenizeConfig
components = {
"wiki": TokenizeConfig(...), # produces a tokenized cache
"books": TokenizeConfig(...),
}
weights = {"wiki": 0.6, "books": 0.4}
# Verifies tokenizers and builds LmDataConfig automatically
lm_data_cfg = lm_mixture_data_config(
components=components,
weights=weights,
shuffle=True, # hierarchical block shuffle (default)
)
Summary
Marin ensures training mixture reproducibility through four core mechanisms:
- Immutable Component Identity: Fingerprinting captures
name@versionidentifiers rather than physical paths, surviving cache deletion. - Deterministic Weight Mapping: The weight map is constructed during fingerprinting and baked into the configuration hash.
- Runtime Tokenizer Verification: The
_verify_tokenizers_same()function enforces identical tokenizers across all mixture components, preventing token ID drift. - Fingerprint‑Derived Seeding: Shuffle seeds are derived from the fingerprint, guaranteeing identical data ordering across runs.
These mechanisms collectively ensure that experiments can be re‑executed months later with identical training mixtures, even if the underlying tokenized caches have been regenerated or moved.
Frequently Asked Questions
What happens if different mixture components use incompatible tokenizers?
Marin raises a runtime error during the resolution phase. The _verify_tokenizers_same() function in lib/marin/src/marin/processing/tokenize/data_configs.py (lines 76‑92) explicitly checks that all TokenizedCache instances share identical tokenizers (or equivalent ones via _are_tokenizers_equivalent()). This prevents subtle token ID mismatches that would otherwise break reproducibility.
Does deleting tokenized caches affect training mixture reproducibility?
No. Because the fingerprint phase uses placeholder components containing only identity metadata (name@version) rather than physical paths, deleting and regenerating caches does not alter the logical mixture. When the pipeline re‑runs, Marin resolves the same handles to the newly materialized caches while preserving the original weight mapping and shuffle order.
How does Marin ensure the same data shuffling across experiment restarts?
Marin derives the shuffle seed from the pipeline fingerprint itself. The default BlockShuffleConfig (referenced as DEFAULT_LM_DATA_SHUFFLE in lib/marin/src/marin/processing/tokenize/data_configs.py) uses this deterministic seed to generate identical block orderings. Since the component list and weights are fixed during fingerprinting, the shuffle sequence remains constant across runs.
Can I modify mixture weights without changing the experiment fingerprint?
No. The weight mapping is constructed during the fingerprint phase (lines 91‑99 in lib/marin/src/marin/experiment/data.py) and constitutes part of the configuration identity. Changing any training or validation weight produces a different fingerprint, triggering a new build. This design ensures that weight changes are always tracked and reproducible.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →