# How the Marin Experiment Data Module Handles Tokenized Dataset Caching

> Discover how the marin experiment data module uses lazy ArtifactStep handles and cryptographic fingerprinting for automated tokenized dataset caching, allowing efficient cache reuse.

- Repository: [The Marin Project/marin](https://github.com/marin-community/marin)
- Tags: internals
- Published: 2026-08-28

---

**The `marin.experiment.data` module leverages lazy `ArtifactStep` handles and cryptographic fingerprinting to automate tokenized dataset caching, enabling experiments to reuse existing caches through optional `pin` arguments or automatically detect valid cached artifacts based on configuration hashes.**

The `marin-community/marin` repository eliminates manual cache management for machine learning experiments through its sophisticated `marin/experiment/data` module. Instead of tracking file paths or re-running expensive tokenization pipelines, researchers define dataset configurations that automatically resolve to cached artifacts. This system handles everything from raw text tokenization to pre-built HuggingFace cache downloads using a unified fingerprint-based caching strategy.

## Core Caching Architecture

The foundation of tokenized dataset caching rests on two abstractions defined across the codebase. The **`ArtifactStep`** class, defined in [`lib/marin/src/marin/execution/lazy.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/execution/lazy.py), serves as a lazy builder that encapsulates execution logic without immediately triggering computation. These steps defer materialization until a downstream process explicitly requests the artifact.

When targeting tokenized text, steps produce **`TokenizedCache`** artifacts—concrete representations of Levanter token caches stored on disk. As implemented in [`lib/marin/src/marin/processing/tokenize/tokenize.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/processing/tokenize/tokenize.py) at line 46, these artifacts encapsulate the tokenizer metadata, format specifications (such as `TextLmDatasetFormat`), and the actual shard files containing tokenized data.

## Fingerprint‑Based Cache Invalidation

Cache reuse is determined by **fingerprinting** rather than manual versioning. When an `ArtifactStep` is instantiated, the framework generates a cryptographic fingerprint from the configuration object—whether a `TokenizeConfig` or `PretokenizedCacheDownloadConfig`—within the `StepContext` class. This logic resides in [`lib/marin/src/marin/execution/build_context.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/execution/build_context.py) around line 71.

The fingerprint becomes the artifact’s unique identifier and determines its output path. Identical configurations produce identical fingerprints, causing the system to automatically skip rebuilding and return the existing cache location. This eliminates redundant computation while ensuring that any change to tokenizer settings, source paths, or metadata tags invalidates the cache and triggers a fresh build.

## Building Tokenized Datasets with `tokenized()`

The primary entry point for creating cached tokenized datasets is the `tokenized()` function in [`lib/marin/src/marin/experiment/data.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/experiment/data.py). This function accepts exactly one raw source through mutually exclusive parameters:

- **`source`** — A HuggingFace repository ID (e.g., `allenai/c4`) that triggers on-the-fly tokenization
- **`paths`** — Glob patterns resolved against the run prefix for local files
- **`raw`** — Combined with `glob` for downloadable raw data requiring local tokenization

Between lines 75 and 99, the function constructs a **`TokenizeConfig`** (or `HfTokenizeConfig` for HF sources) containing the `cache_path` (derived from the step's output directory), tokenizer name, format, and optional `sample_count` and metadata `tags`. The function then returns an `ArtifactStep` configured with `artifact_type=TokenizedCache` and a run function pointing to `_tokenize_fn` (lines 101-108).

```python

# Tokenize a raw HuggingFace dataset with optional cache pinning

train_data = tokenized(
    name="my_corpus",
    tokenizer="gpt2",
    source="allenai/c4",
    pin="/data/caches/my_corpus@2024-01-01",  # Reuse existing cache if present

)

```

## Pinning Cache Locations with `pin` and `override_path`

For reproducible research requiring specific cache versions, the `tokenized()` function supports a **`pin`** argument. When provided, this value populates the `override_path` parameter of the `ArtifactStep` (see implementation at lines 48-50 and 101-108 in [`lib/marin/src/marin/experiment/data.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/experiment/data.py)).

When `override_path` is set, the step immediately returns the pinned directory path instead of executing the tokenization run. This mechanism guarantees that experiments use known-good cache versions, bypassing fingerprint checks and preventing accidental cache regeneration when configurations drift. This is critical for ensuring bit-for-bit reproducibility across different compute environments.

## Pre‑tokenized Datasets via `pretokenized()`

For large corpora already tokenized and hosted on HuggingFace, the `pretokenized()` function (lines 12-30 in [`lib/marin/src/marin/experiment/data.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/experiment/data.py)) provides a fast path that bypasses local tokenization entirely. This function creates a `PretokenizedCacheDownloadConfig` and invokes `fetch_pretokenized_cache` (located in [`lib/marin/src/marin/processing/tokenize/download_pretokenized.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/processing/tokenize/download_pretokenized.py)), which downloads the pre-built shards directly into `ctx.output_path`.

Because the tokenization work is already complete, this approach drastically reduces start-up costs for large corpora while still returning a standard `TokenizedCache` handle. Downstream components like `mixture()` treat these identically to freshly tokenized caches.

```python

# Use a pre‑tokenized cache hosted on HuggingFace

prebuilt = pretokenized(
    name="fineweb_edu",
    repo_id="levanter/fineweb-edu",
    tokenizer="gpt2",
    revision="v1",
    pin="/data/caches/fineweb_edu@v1",
)

```

## Composing Datasets with `mixture()`

The `mixture()` function (lines 66-78 in [`lib/marin/src/marin/experiment/data.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/experiment/data.py)) combines multiple `ArtifactStep[TokenizedCache]` handles into a unified training configuration. It accepts a mapping of handles to sampling weights and validates that **all components share the same tokenizer**, preventing incoherent training pipelines.

When executed, `mixture()` resolves each lazy handle—triggering tokenization or downloads only if fingerprints miss the cache—and constructs an `LmDataConfig` compatible with Levanter's language model trainer.

```python

# Build a mixture of two tokenized datasets

lm_config = mixture(
    ctx=step_ctx,
    train={train_data: 0.7, prebuilt: 0.3},
    validation=[prebuilt],
    shuffle=True,
)

```

## Summary

- **Caching is driven by configuration fingerprints** generated in `StepContext` ([`build_context.py`](https://github.com/marin-community/marin/blob/main/build_context.py) line 71), not manual version strings.
- The **`pin`** argument forces cache reuse via `override_path`, ensuring reproducibility across experimental runs.
- **`pretokenized()`** bypasses local tokenization by downloading pre-built HuggingFace caches, minimizing start-up costs.
- All dataset handles are **lazy** (`ArtifactStep`) and materialize only when downstream steps (like `mixture`) request them.
- The `mixture()` function enforces tokenizer consistency across all input datasets before constructing the final training configuration.

## Frequently Asked Questions

### How does fingerprinting determine cache reuse?

The framework generates a cryptographic hash from the dataset configuration—including tokenizer name, source path, and tags—in the `StepContext` class. This fingerprint defines the artifact's output directory. If the directory exists and matches the fingerprint, the system returns the cached path immediately; otherwise, it executes the tokenization run.

### What is the difference between `tokenized()` and `pretokenized()`?

The `tokenized()` function processes raw text through a tokenizer locally, creating a `TokenizeConfig` and executing `_tokenize_fn` to generate shards. The `pretokenized()` function assumes the data is already tokenized on HuggingFace, creating a `PretokenizedCacheDownloadConfig` and merely downloading existing cache files without invoking a tokenizer.

### How does the `pin` argument affect cache behavior?

When a `pin` path is provided to `tokenized()` or `pretokenized()`, it sets the `override_path` on the resulting `ArtifactStep`. This causes the step to return the pinned directory immediately, bypassing fingerprint checks and eliminating the tokenization or download run entirely.

### Can I mix tokenized datasets with different tokenizers?

No. The `mixture()` function explicitly validates that all `ArtifactStep[TokenizedCache]` inputs share identical tokenizer configurations. Attempting to combine datasets with different tokenizers will raise an error, as this would produce incoherent token sequences during training.