# How to Add a New Dataset to Marin Using Agent Skills

> Learn to add a new Hugging Face dataset to Marin using the add dataset agent skill. Inspect schema, create lazy-builder modules, and validate registration easily.

- Repository: [The Marin Project/marin](https://github.com/marin-community/marin)
- Tags: how-to-guide
- Published: 2026-08-29

---

**You can add a new Hugging Face dataset to Marin by using the `add-dataset` agent skill to inspect the schema, creating a lazy-builder module in `experiments/datasets/`, and validating the registration without manually editing configuration files.**

The Marin framework (available at `marin-community/marin`) provides a declarative workflow for registering new Hugging Face datasets through dedicated agent skills. This approach eliminates the need to manipulate low-level configuration files manually, instead relying on programmatic schema inspection and Python modules that follow a standardized lazy-builder pattern.

## Step 1: Inspect the Dataset Schema

Before registering a dataset, you must understand its structure, including available splits, feature columns, and which field contains the raw text.

### Running the Schema Inspection Command

Use the provided utility script to fetch the dataset metadata without downloading the full corpus. From the repository root, run:

```bash
uv run lib/marin/tools/get_hf_dataset_schema.py <hf_dataset_id> [--config_name <cfg>] [--trust_remote_code]

```

If the script indicates that a configuration name is required, rerun the command with the `--config_name` flag specifying the appropriate subset.

### Interpreting the Schema Output

The inspection tool returns four critical pieces of information:

- **features**: A mapping of column names to their data types.
- **splits**: Available data partitions (typically `train`, `validation`, or `test`).
- **text_field_candidates**: An ordered list of fields that likely contain the primary text content, ranked by relevance.
- **sample_row**: A representative data excerpt for manual sanity-checking.

Review the `sample_row` to confirm the data matches your training requirements before proceeding.

## Step 2: Select the Text Field

Choose the appropriate text column based on the following priority:

1. A field exactly named `text`.
2. Any column containing the substring "text".
3. Any other string column that holds the primary content you want to model.

The first entry in `text_field_candidates` usually represents the best choice, but verify against the `sample_row` output to ensure it contains the expected raw text rather than metadata or labels.

## Step 3: Create the Lazy-Builder Module

New datasets are implemented as Python modules under the `experiments/datasets/` directory. These modules use the lazy-builder API to expose dataset handles without loading data into memory prematurely.

### Module Structure and Required Callables

Create a new file at `experiments/datasets/<your_dataset_name>.py` with the following structure:

```python
from marin.experiment.data import DatasetBuilder

class MyDatasetBuilder(DatasetBuilder):
    hf_dataset_id = "<hf_dataset_id>"
    config_name = "<optional_config_name>"  # Omit if not required

    text_field = "<chosen_text_field>"

def my_dataset():
    """Returns a single-corpus handle."""
    return MyDatasetBuilder().load_one()

def my_datasets():
    """Returns a mapping of split-specific handles."""
    return MyDatasetBuilder().load_all()

```

Replace `MyDatasetBuilder` with a descriptive class name and adjust the function names (`my_dataset` and `my_datasets`) to match your dataset's identifier.

### Reference Implementation

The [`experiments/datasets/nemotron.py`](https://github.com/marin-community/marin/blob/main/experiments/datasets/nemotron.py) file in the repository demonstrates the required implementation pattern. Study this example to see how the `DatasetBuilder` base class from [`lib/marin/src/marin/experiment/data/__init__.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/experiment/data/__init__.py) is properly subclassed and how the lazy-loading methods are structured.

## Step 4: Validate and Commit

After creating the module, invoke the validation step provided by the `add-dataset` skill. The skill automatically tests that:

- The specified configuration and splits exist.
- The selected text field maps correctly to the dataset's features.
- A sample row can be successfully tokenized using the builder.

Once validation passes, use the repository's standard `commit` skill (documented in [`.agents/skills/commit/SKILL.md`](https://github.com/marin-community/marin/blob/main/.agents/skills/commit/SKILL.md)) to stage and push your changes. The dataset is now a first-class ingredient available for any Marin experiment.

## Summary

- **Schema inspection** is performed using [`lib/marin/tools/get_hf_dataset_schema.py`](https://github.com/marin-community/marin/blob/main/lib/marin/tools/get_hf_dataset_schema.py), which retrieves features, splits, and candidate text fields without downloading the full dataset.
- **Text field selection** should prioritize exact matches for "text", then substring matches, guided by the `text_field_candidates` list.
- **Implementation** requires creating a module in `experiments/datasets/` that subclasses `DatasetBuilder` and exposes `<name>_dataset()` and `<name>_datasets()` functions.
- **Validation** occurs through the `add-dataset` skill, ensuring the builder correctly loads and tokenizes sample data.
- **Reference code** is available in [`experiments/datasets/nemotron.py`](https://github.com/marin-community/marin/blob/main/experiments/datasets/nemotron.py) and the skill documentation at [`.agents/skills/add-dataset/SKILL.md`](https://github.com/marin-community/marin/blob/main/.agents/skills/add-dataset/SKILL.md).

## Frequently Asked Questions

### What if the Hugging Face dataset requires a specific configuration name?

If [`get_hf_dataset_schema.py`](https://github.com/marin-community/marin/blob/main/get_hf_dataset_schema.py) indicates that a config is required, rerun the command with the `--config_name <name>` argument. The configuration name must be stored in the `config_name` class attribute of your `DatasetBuilder` subclass.

### How do I know if my dataset module is working correctly?

The `add-dataset` skill includes a validation step that programmatically invokes your builder. If the skill reports success, your module correctly implements the lazy-builder interface and the selected text field contains tokenizable content. Errors during this step usually indicate a mismatch between the `text_field` value and the actual column names returned in the schema.

### Can I register multiple splits of the same dataset?

Yes. Implement the `<name>_datasets()` function to return a mapping of split names to dataset handles using `MyDatasetBuilder().load_all()`. This approach is particularly useful for Hugging Face datasets that provide separate subsets for `train`, `validation`, and `test` within the same repository ID.

### Where is the DatasetBuilder base class defined?

The `DatasetBuilder` base class is exported from [`lib/marin/src/marin/experiment/data/__init__.py`](https://github.com/marin-community/marin/blob/main/lib/marin/src/marin/experiment/data/__init__.py). Your dataset modules should import this class via `from marin.experiment.data import DatasetBuilder` to ensure compatibility with the validation and loading systems.