How to Add a New Dataset to Marin Using Agent Skills
You can add a new Hugging Face dataset to Marin by using the add-dataset agent skill to inspect the schema, creating a lazy-builder module in experiments/datasets/, and validating the registration without manually editing configuration files.
The Marin framework (available at marin-community/marin) provides a declarative workflow for registering new Hugging Face datasets through dedicated agent skills. This approach eliminates the need to manipulate low-level configuration files manually, instead relying on programmatic schema inspection and Python modules that follow a standardized lazy-builder pattern.
Step 1: Inspect the Dataset Schema
Before registering a dataset, you must understand its structure, including available splits, feature columns, and which field contains the raw text.
Running the Schema Inspection Command
Use the provided utility script to fetch the dataset metadata without downloading the full corpus. From the repository root, run:
uv run lib/marin/tools/get_hf_dataset_schema.py <hf_dataset_id> [--config_name <cfg>] [--trust_remote_code]
If the script indicates that a configuration name is required, rerun the command with the --config_name flag specifying the appropriate subset.
Interpreting the Schema Output
The inspection tool returns four critical pieces of information:
- features: A mapping of column names to their data types.
- splits: Available data partitions (typically
train,validation, ortest). - text_field_candidates: An ordered list of fields that likely contain the primary text content, ranked by relevance.
- sample_row: A representative data excerpt for manual sanity-checking.
Review the sample_row to confirm the data matches your training requirements before proceeding.
Step 2: Select the Text Field
Choose the appropriate text column based on the following priority:
- A field exactly named
text. - Any column containing the substring "text".
- Any other string column that holds the primary content you want to model.
The first entry in text_field_candidates usually represents the best choice, but verify against the sample_row output to ensure it contains the expected raw text rather than metadata or labels.
Step 3: Create the Lazy-Builder Module
New datasets are implemented as Python modules under the experiments/datasets/ directory. These modules use the lazy-builder API to expose dataset handles without loading data into memory prematurely.
Module Structure and Required Callables
Create a new file at experiments/datasets/<your_dataset_name>.py with the following structure:
from marin.experiment.data import DatasetBuilder
class MyDatasetBuilder(DatasetBuilder):
hf_dataset_id = "<hf_dataset_id>"
config_name = "<optional_config_name>" # Omit if not required
text_field = "<chosen_text_field>"
def my_dataset():
"""Returns a single-corpus handle."""
return MyDatasetBuilder().load_one()
def my_datasets():
"""Returns a mapping of split-specific handles."""
return MyDatasetBuilder().load_all()
Replace MyDatasetBuilder with a descriptive class name and adjust the function names (my_dataset and my_datasets) to match your dataset's identifier.
Reference Implementation
The experiments/datasets/nemotron.py file in the repository demonstrates the required implementation pattern. Study this example to see how the DatasetBuilder base class from lib/marin/src/marin/experiment/data/__init__.py is properly subclassed and how the lazy-loading methods are structured.
Step 4: Validate and Commit
After creating the module, invoke the validation step provided by the add-dataset skill. The skill automatically tests that:
- The specified configuration and splits exist.
- The selected text field maps correctly to the dataset's features.
- A sample row can be successfully tokenized using the builder.
Once validation passes, use the repository's standard commit skill (documented in .agents/skills/commit/SKILL.md) to stage and push your changes. The dataset is now a first-class ingredient available for any Marin experiment.
Summary
- Schema inspection is performed using
lib/marin/tools/get_hf_dataset_schema.py, which retrieves features, splits, and candidate text fields without downloading the full dataset. - Text field selection should prioritize exact matches for "text", then substring matches, guided by the
text_field_candidateslist. - Implementation requires creating a module in
experiments/datasets/that subclassesDatasetBuilderand exposes<name>_dataset()and<name>_datasets()functions. - Validation occurs through the
add-datasetskill, ensuring the builder correctly loads and tokenizes sample data. - Reference code is available in
experiments/datasets/nemotron.pyand the skill documentation at.agents/skills/add-dataset/SKILL.md.
Frequently Asked Questions
What if the Hugging Face dataset requires a specific configuration name?
If get_hf_dataset_schema.py indicates that a config is required, rerun the command with the --config_name <name> argument. The configuration name must be stored in the config_name class attribute of your DatasetBuilder subclass.
How do I know if my dataset module is working correctly?
The add-dataset skill includes a validation step that programmatically invokes your builder. If the skill reports success, your module correctly implements the lazy-builder interface and the selected text field contains tokenizable content. Errors during this step usually indicate a mismatch between the text_field value and the actual column names returned in the schema.
Can I register multiple splits of the same dataset?
Yes. Implement the <name>_datasets() function to return a mapping of split names to dataset handles using MyDatasetBuilder().load_all(). This approach is particularly useful for Hugging Face datasets that provide separate subsets for train, validation, and test within the same repository ID.
Where is the DatasetBuilder base class defined?
The DatasetBuilder base class is exported from lib/marin/src/marin/experiment/data/__init__.py. Your dataset modules should import this class via from marin.experiment.data import DatasetBuilder to ensure compatibility with the validation and loading systems.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →