# How to Specify Multi-Dataset Interleaving in Soup Configuration

> Learn how to specify multi-dataset interleaving in Soup configuration. Easily combine datasets using concat under over or probs strategies for efficient data loading.

- Repository: [Alpamys Makazhan/Soup](https://github.com/MakazhanAlpamys/Soup)
- Tags: how-to-guide
- Published: 2026-09-06

---

**Specify multi-dataset interleaving in Soup by setting `data.train` to a list of datasets and adding a `data.interleave` block with a chosen strategy (`concat`, `under`, `over`, or `probs`).**

Soup's data pipeline lets you train on multiple datasets simultaneously by interleaving their rows during loading. The `data.interleave` field in your YAML recipe controls how rows from each source combine, whether you're working with local files or streaming remote data. This article walks through the four available strategies, their behavior, and how to configure them correctly based on the source code in `MakazhanAlpamys/Soup`.

## Understanding the Interleave Specification

The interleaving logic centers on `parse_interleave` in [`src/soup_cli/utils/data_pipeline.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/data_pipeline.py) (lines 44-54). This validator transforms your YAML configuration into a frozen **`InterleaveSpec`** object, enforcing shape constraints, type checking, and probability validation. The resulting specification is passed to downstream loading code in [`src/soup_cli/data/loader.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/data/loader.py), which delegates to Hugging Face Datasets' `interleave_datasets` or `concatenate_datasets` utilities.

## The Four Interleaving Strategies

Soup supports four strategies that behave identically on local and streaming paths, differing only in dataset materialization:

| Strategy | Description | Stopping Behavior |
|----------|-------------|-----------------|
| **`concat`** | Append sources in order | Exhaust all sources sequentially |
| **`under`** | Truncate to smallest source | Stop when first source exhausts |
| **`over`** | Upsample smaller sources (cycling) | Stop when all sources exhaust |
| **`probs`** | Exact row ratios via probabilities | Stop when first source exhausts |

The **`probs`** strategy additionally requires a `probs` array where values sum to 1.0, with length matching `data.train`.

## Configuration Examples

### Concatenate Two Local Files

Use `concat` for simple sequential combination:

```yaml
data:
  train:
    - dataset_a.jsonl
    - dataset_b.jsonl
  interleave: { strategy: concat }
  eval_on_each_dataset: true

```

### Probability-Based Interleaving

Control exact sampling ratios with `probs`:

```yaml
data:
  train:
    - dataset_a.jsonl
    - dataset_b.jsonl
  interleave: { strategy: probs, probs: [0.7, 0.3] }
  eval_on_each_dataset: true

```

The `probs` array `[0.7, 0.3]` ensures 70% of training rows come from `dataset_a.jsonl` and 30% from `dataset_b.jsonl`.

### Streaming Interleaving from Remote URIs

Enable `data.streaming: true` for HTTP(S) sources:

```yaml
data:
  train:
    - https://my-bucket.s3.amazonaws.com/dataset_a.jsonl
    - https://my-bucket.s3.amazonaws.com/dataset_b.jsonl
  streaming: true
  interleave: { strategy: under }

```

The **`under`** strategy truncates all sources to the length of the shortest dataset.

### Interleaving Hugging Face Hub Datasets

Reference hub datasets directly for eager loading:

```yaml
data:
  train:
    - teknium/OpenHermes-2.5
    - HuggingFaceTB/SmolLM2-Chat-Base
  interleave: { strategy: over }

```

The **`over`** strategy cycles through smaller datasets until all sources exhaust.

## Critical Configuration Rules

Several constraints govern valid multi-dataset interleaving configurations:

- **List requirement**: `data.train` must be a YAML list. A single string path disables interleaving entirely.
- **Probability alignment**: For `strategy: probs`, the `probs` array length must exactly match `data.train` length.
- **Packing incompatibility**: `training.packing` and `training.multipack` must be disabled when interleaving.
- **Streaming requirement**: Remote URIs require `data.streaming: true`. Local-only mode rejects HTTP(S) sources.
- **Uniform source types**: Mixing hub dataset names with local files or remote URIs in the same list is forbidden—the loader cannot reconcile differing split semantics.

## Source Files and Testing

The implementation spans multiple files with dedicated test coverage:

| File | Purpose |
|------|---------|
| [`src/soup_cli/utils/data_pipeline.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/data_pipeline.py) | `parse_interleave` validator and `InterleaveSpec` definition |
| [`src/soup_cli/data/loader.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/data/loader.py) | Consumes spec and routes to dataset utilities |
| [`docs/data.md`](https://github.com/MakazhanAlpamys/Soup/blob/main/docs/data.md) (lines 602-610) | User-facing documentation |
| [`tests/test_v0420.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/tests/test_v0420.py) | Version compatibility tests for interleaving |
| [`tests/test_issue443_interleave_wiring.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/tests/test_issue443_interleave_wiring.py) | Wiring validation between parser and loader |

## Summary

- **Multi-dataset interleaving** in Soup requires `data.train` as a list plus a `data.interleave` block specifying strategy.
- **Four strategies** (`concat`, `under`, `over`, `probs`) cover sequential, truncated, upsampled, and probability-weighted combination.
- **Streaming support** needs `data.streaming: true` for remote URIs, with behavior delegated to `🤗 datasets` utilities.
- **Key constraints**: disable packing/multipack, match `probs` length to source count, and avoid mixing hub names with file paths.

## Frequently Asked Questions

### What happens if I omit `data.interleave` when `data.train` is a list?

The parser raises a validation error. When `data.train` contains multiple sources, `data.interleave` is mandatory to specify combination behavior. This check occurs in `parse_interleave` within [`src/soup_cli/utils/data_pipeline.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/data_pipeline.py).

### Can I use `strategy: probs` with more than two datasets?

Yes. The `probs` array simply needs one probability per dataset in `data.train`. For three datasets, use `probs: [0.5, 0.3, 0.2]` or any other distribution summing to 1.0. The validator enforces both length matching and proper normalization.

### Why does interleaving disable packing and multipacking?

Packing and multipacking assume contiguous sequences from a single dataset source. Interleaved datasets mix rows from multiple sources, breaking the contiguous-block assumptions that packing algorithms rely on. The loader explicitly rejects this combination to prevent silent data corruption.

### Is streaming interleaving slower than local file interleaving?

Streaming adds network latency but uses the same `interleave_datasets` or `concatenate_datasets` call structure. The `strategy` semantics remain identical—`under` still truncates to shortest, `over` still upsamples—only the data materialization path differs (network chunks vs. local memory maps).