How to Specify Multi-Dataset Interleaving in Soup Configuration
Specify multi-dataset interleaving in Soup by setting data.train to a list of datasets and adding a data.interleave block with a chosen strategy (concat, under, over, or probs).
Soup's data pipeline lets you train on multiple datasets simultaneously by interleaving their rows during loading. The data.interleave field in your YAML recipe controls how rows from each source combine, whether you're working with local files or streaming remote data. This article walks through the four available strategies, their behavior, and how to configure them correctly based on the source code in MakazhanAlpamys/Soup.
Understanding the Interleave Specification
The interleaving logic centers on parse_interleave in src/soup_cli/utils/data_pipeline.py (lines 44-54). This validator transforms your YAML configuration into a frozen InterleaveSpec object, enforcing shape constraints, type checking, and probability validation. The resulting specification is passed to downstream loading code in src/soup_cli/data/loader.py, which delegates to Hugging Face Datasets' interleave_datasets or concatenate_datasets utilities.
The Four Interleaving Strategies
Soup supports four strategies that behave identically on local and streaming paths, differing only in dataset materialization:
| Strategy | Description | Stopping Behavior |
|---|---|---|
concat |
Append sources in order | Exhaust all sources sequentially |
under |
Truncate to smallest source | Stop when first source exhausts |
over |
Upsample smaller sources (cycling) | Stop when all sources exhaust |
probs |
Exact row ratios via probabilities | Stop when first source exhausts |
The probs strategy additionally requires a probs array where values sum to 1.0, with length matching data.train.
Configuration Examples
Concatenate Two Local Files
Use concat for simple sequential combination:
data:
train:
- dataset_a.jsonl
- dataset_b.jsonl
interleave: { strategy: concat }
eval_on_each_dataset: true
Probability-Based Interleaving
Control exact sampling ratios with probs:
data:
train:
- dataset_a.jsonl
- dataset_b.jsonl
interleave: { strategy: probs, probs: [0.7, 0.3] }
eval_on_each_dataset: true
The probs array [0.7, 0.3] ensures 70% of training rows come from dataset_a.jsonl and 30% from dataset_b.jsonl.
Streaming Interleaving from Remote URIs
Enable data.streaming: true for HTTP(S) sources:
data:
train:
- https://my-bucket.s3.amazonaws.com/dataset_a.jsonl
- https://my-bucket.s3.amazonaws.com/dataset_b.jsonl
streaming: true
interleave: { strategy: under }
The under strategy truncates all sources to the length of the shortest dataset.
Interleaving Hugging Face Hub Datasets
Reference hub datasets directly for eager loading:
data:
train:
- teknium/OpenHermes-2.5
- HuggingFaceTB/SmolLM2-Chat-Base
interleave: { strategy: over }
The over strategy cycles through smaller datasets until all sources exhaust.
Critical Configuration Rules
Several constraints govern valid multi-dataset interleaving configurations:
- List requirement:
data.trainmust be a YAML list. A single string path disables interleaving entirely. - Probability alignment: For
strategy: probs, theprobsarray length must exactly matchdata.trainlength. - Packing incompatibility:
training.packingandtraining.multipackmust be disabled when interleaving. - Streaming requirement: Remote URIs require
data.streaming: true. Local-only mode rejects HTTP(S) sources. - Uniform source types: Mixing hub dataset names with local files or remote URIs in the same list is forbidden—the loader cannot reconcile differing split semantics.
Source Files and Testing
The implementation spans multiple files with dedicated test coverage:
| File | Purpose |
|---|---|
src/soup_cli/utils/data_pipeline.py |
parse_interleave validator and InterleaveSpec definition |
src/soup_cli/data/loader.py |
Consumes spec and routes to dataset utilities |
docs/data.md (lines 602-610) |
User-facing documentation |
tests/test_v0420.py |
Version compatibility tests for interleaving |
tests/test_issue443_interleave_wiring.py |
Wiring validation between parser and loader |
Summary
- Multi-dataset interleaving in Soup requires
data.trainas a list plus adata.interleaveblock specifying strategy. - Four strategies (
concat,under,over,probs) cover sequential, truncated, upsampled, and probability-weighted combination. - Streaming support needs
data.streaming: truefor remote URIs, with behavior delegated to🤗 datasetsutilities. - Key constraints: disable packing/multipack, match
probslength to source count, and avoid mixing hub names with file paths.
Frequently Asked Questions
What happens if I omit data.interleave when data.train is a list?
The parser raises a validation error. When data.train contains multiple sources, data.interleave is mandatory to specify combination behavior. This check occurs in parse_interleave within src/soup_cli/utils/data_pipeline.py.
Can I use strategy: probs with more than two datasets?
Yes. The probs array simply needs one probability per dataset in data.train. For three datasets, use probs: [0.5, 0.3, 0.2] or any other distribution summing to 1.0. The validator enforces both length matching and proper normalization.
Why does interleaving disable packing and multipacking?
Packing and multipacking assume contiguous sequences from a single dataset source. Interleaved datasets mix rows from multiple sources, breaking the contiguous-block assumptions that packing algorithms rely on. The loader explicitly rejects this combination to prevent silent data corruption.
Is streaming interleaving slower than local file interleaving?
Streaming adds network latency but uses the same interleave_datasets or concatenate_datasets call structure. The strategy semantics remain identical—under still truncates to shortest, over still upsamples—only the data materialization path differs (network chunks vs. local memory maps).
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →