How the Marin Experiment Data Module Handles Tokenized Dataset Caching
The marin.experiment.data module leverages lazy ArtifactStep handles and cryptographic fingerprinting to automate tokenized dataset caching, enabling experiments to reuse existing caches through optional pin arguments or automatically detect valid cached artifacts based on configuration hashes.
The marin-community/marin repository eliminates manual cache management for machine learning experiments through its sophisticated marin/experiment/data module. Instead of tracking file paths or re-running expensive tokenization pipelines, researchers define dataset configurations that automatically resolve to cached artifacts. This system handles everything from raw text tokenization to pre-built HuggingFace cache downloads using a unified fingerprint-based caching strategy.
Core Caching Architecture
The foundation of tokenized dataset caching rests on two abstractions defined across the codebase. The ArtifactStep class, defined in lib/marin/src/marin/execution/lazy.py, serves as a lazy builder that encapsulates execution logic without immediately triggering computation. These steps defer materialization until a downstream process explicitly requests the artifact.
When targeting tokenized text, steps produce TokenizedCache artifacts—concrete representations of Levanter token caches stored on disk. As implemented in lib/marin/src/marin/processing/tokenize/tokenize.py at line 46, these artifacts encapsulate the tokenizer metadata, format specifications (such as TextLmDatasetFormat), and the actual shard files containing tokenized data.
Fingerprint‑Based Cache Invalidation
Cache reuse is determined by fingerprinting rather than manual versioning. When an ArtifactStep is instantiated, the framework generates a cryptographic fingerprint from the configuration object—whether a TokenizeConfig or PretokenizedCacheDownloadConfig—within the StepContext class. This logic resides in lib/marin/src/marin/execution/build_context.py around line 71.
The fingerprint becomes the artifact’s unique identifier and determines its output path. Identical configurations produce identical fingerprints, causing the system to automatically skip rebuilding and return the existing cache location. This eliminates redundant computation while ensuring that any change to tokenizer settings, source paths, or metadata tags invalidates the cache and triggers a fresh build.
Building Tokenized Datasets with tokenized()
The primary entry point for creating cached tokenized datasets is the tokenized() function in lib/marin/src/marin/experiment/data.py. This function accepts exactly one raw source through mutually exclusive parameters:
source— A HuggingFace repository ID (e.g.,allenai/c4) that triggers on-the-fly tokenizationpaths— Glob patterns resolved against the run prefix for local filesraw— Combined withglobfor downloadable raw data requiring local tokenization
Between lines 75 and 99, the function constructs a TokenizeConfig (or HfTokenizeConfig for HF sources) containing the cache_path (derived from the step's output directory), tokenizer name, format, and optional sample_count and metadata tags. The function then returns an ArtifactStep configured with artifact_type=TokenizedCache and a run function pointing to _tokenize_fn (lines 101-108).
# Tokenize a raw HuggingFace dataset with optional cache pinning
train_data = tokenized(
name="my_corpus",
tokenizer="gpt2",
source="allenai/c4",
pin="/data/caches/my_corpus@2024-01-01", # Reuse existing cache if present
)
Pinning Cache Locations with pin and override_path
For reproducible research requiring specific cache versions, the tokenized() function supports a pin argument. When provided, this value populates the override_path parameter of the ArtifactStep (see implementation at lines 48-50 and 101-108 in lib/marin/src/marin/experiment/data.py).
When override_path is set, the step immediately returns the pinned directory path instead of executing the tokenization run. This mechanism guarantees that experiments use known-good cache versions, bypassing fingerprint checks and preventing accidental cache regeneration when configurations drift. This is critical for ensuring bit-for-bit reproducibility across different compute environments.
Pre‑tokenized Datasets via pretokenized()
For large corpora already tokenized and hosted on HuggingFace, the pretokenized() function (lines 12-30 in lib/marin/src/marin/experiment/data.py) provides a fast path that bypasses local tokenization entirely. This function creates a PretokenizedCacheDownloadConfig and invokes fetch_pretokenized_cache (located in lib/marin/src/marin/processing/tokenize/download_pretokenized.py), which downloads the pre-built shards directly into ctx.output_path.
Because the tokenization work is already complete, this approach drastically reduces start-up costs for large corpora while still returning a standard TokenizedCache handle. Downstream components like mixture() treat these identically to freshly tokenized caches.
# Use a pre‑tokenized cache hosted on HuggingFace
prebuilt = pretokenized(
name="fineweb_edu",
repo_id="levanter/fineweb-edu",
tokenizer="gpt2",
revision="v1",
pin="/data/caches/fineweb_edu@v1",
)
Composing Datasets with mixture()
The mixture() function (lines 66-78 in lib/marin/src/marin/experiment/data.py) combines multiple ArtifactStep[TokenizedCache] handles into a unified training configuration. It accepts a mapping of handles to sampling weights and validates that all components share the same tokenizer, preventing incoherent training pipelines.
When executed, mixture() resolves each lazy handle—triggering tokenization or downloads only if fingerprints miss the cache—and constructs an LmDataConfig compatible with Levanter's language model trainer.
# Build a mixture of two tokenized datasets
lm_config = mixture(
ctx=step_ctx,
train={train_data: 0.7, prebuilt: 0.3},
validation=[prebuilt],
shuffle=True,
)
Summary
- Caching is driven by configuration fingerprints generated in
StepContext(build_context.pyline 71), not manual version strings. - The
pinargument forces cache reuse viaoverride_path, ensuring reproducibility across experimental runs. pretokenized()bypasses local tokenization by downloading pre-built HuggingFace caches, minimizing start-up costs.- All dataset handles are lazy (
ArtifactStep) and materialize only when downstream steps (likemixture) request them. - The
mixture()function enforces tokenizer consistency across all input datasets before constructing the final training configuration.
Frequently Asked Questions
How does fingerprinting determine cache reuse?
The framework generates a cryptographic hash from the dataset configuration—including tokenizer name, source path, and tags—in the StepContext class. This fingerprint defines the artifact's output directory. If the directory exists and matches the fingerprint, the system returns the cached path immediately; otherwise, it executes the tokenization run.
What is the difference between tokenized() and pretokenized()?
The tokenized() function processes raw text through a tokenizer locally, creating a TokenizeConfig and executing _tokenize_fn to generate shards. The pretokenized() function assumes the data is already tokenized on HuggingFace, creating a PretokenizedCacheDownloadConfig and merely downloading existing cache files without invoking a tokenizer.
How does the pin argument affect cache behavior?
When a pin path is provided to tokenized() or pretokenized(), it sets the override_path on the resulting ArtifactStep. This causes the step to return the pinned directory immediately, bypassing fingerprint checks and eliminating the tokenization or download run entirely.
Can I mix tokenized datasets with different tokenizers?
No. The mixture() function explicitly validates that all ArtifactStep[TokenizedCache] inputs share identical tokenizer configurations. Attempting to combine datasets with different tokenizers will raise an error, as this would produce incoherent token sequences during training.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →