How to Specify Multi-Dataset Interleaving in Soup Configuration

Specify multi-dataset interleaving in Soup by setting data.train to a list of datasets and adding a data.interleave block with a chosen strategy (concat, under, over, or probs).

Soup's data pipeline lets you train on multiple datasets simultaneously by interleaving their rows during loading. The data.interleave field in your YAML recipe controls how rows from each source combine, whether you're working with local files or streaming remote data. This article walks through the four available strategies, their behavior, and how to configure them correctly based on the source code in MakazhanAlpamys/Soup.

Understanding the Interleave Specification

The interleaving logic centers on parse_interleave in src/soup_cli/utils/data_pipeline.py (lines 44-54). This validator transforms your YAML configuration into a frozen InterleaveSpec object, enforcing shape constraints, type checking, and probability validation. The resulting specification is passed to downstream loading code in src/soup_cli/data/loader.py, which delegates to Hugging Face Datasets' interleave_datasets or concatenate_datasets utilities.

The Four Interleaving Strategies

Soup supports four strategies that behave identically on local and streaming paths, differing only in dataset materialization:

Strategy Description Stopping Behavior
concat Append sources in order Exhaust all sources sequentially
under Truncate to smallest source Stop when first source exhausts
over Upsample smaller sources (cycling) Stop when all sources exhaust
probs Exact row ratios via probabilities Stop when first source exhausts

The probs strategy additionally requires a probs array where values sum to 1.0, with length matching data.train.

Configuration Examples

Concatenate Two Local Files

Use concat for simple sequential combination:

data:
  train:
    - dataset_a.jsonl
    - dataset_b.jsonl
  interleave: { strategy: concat }
  eval_on_each_dataset: true

Probability-Based Interleaving

Control exact sampling ratios with probs:

data:
  train:
    - dataset_a.jsonl
    - dataset_b.jsonl
  interleave: { strategy: probs, probs: [0.7, 0.3] }
  eval_on_each_dataset: true

The probs array [0.7, 0.3] ensures 70% of training rows come from dataset_a.jsonl and 30% from dataset_b.jsonl.

Streaming Interleaving from Remote URIs

Enable data.streaming: true for HTTP(S) sources:

data:
  train:
    - https://my-bucket.s3.amazonaws.com/dataset_a.jsonl
    - https://my-bucket.s3.amazonaws.com/dataset_b.jsonl
  streaming: true
  interleave: { strategy: under }

The under strategy truncates all sources to the length of the shortest dataset.

Interleaving Hugging Face Hub Datasets

Reference hub datasets directly for eager loading:

data:
  train:
    - teknium/OpenHermes-2.5
    - HuggingFaceTB/SmolLM2-Chat-Base
  interleave: { strategy: over }

The over strategy cycles through smaller datasets until all sources exhaust.

Critical Configuration Rules

Several constraints govern valid multi-dataset interleaving configurations:

  • List requirement: data.train must be a YAML list. A single string path disables interleaving entirely.
  • Probability alignment: For strategy: probs, the probs array length must exactly match data.train length.
  • Packing incompatibility: training.packing and training.multipack must be disabled when interleaving.
  • Streaming requirement: Remote URIs require data.streaming: true. Local-only mode rejects HTTP(S) sources.
  • Uniform source types: Mixing hub dataset names with local files or remote URIs in the same list is forbidden—the loader cannot reconcile differing split semantics.

Source Files and Testing

The implementation spans multiple files with dedicated test coverage:

File Purpose
src/soup_cli/utils/data_pipeline.py parse_interleave validator and InterleaveSpec definition
src/soup_cli/data/loader.py Consumes spec and routes to dataset utilities
docs/data.md (lines 602-610) User-facing documentation
tests/test_v0420.py Version compatibility tests for interleaving
tests/test_issue443_interleave_wiring.py Wiring validation between parser and loader

Summary

  • Multi-dataset interleaving in Soup requires data.train as a list plus a data.interleave block specifying strategy.
  • Four strategies (concat, under, over, probs) cover sequential, truncated, upsampled, and probability-weighted combination.
  • Streaming support needs data.streaming: true for remote URIs, with behavior delegated to 🤗 datasets utilities.
  • Key constraints: disable packing/multipack, match probs length to source count, and avoid mixing hub names with file paths.

Frequently Asked Questions

What happens if I omit data.interleave when data.train is a list?

The parser raises a validation error. When data.train contains multiple sources, data.interleave is mandatory to specify combination behavior. This check occurs in parse_interleave within src/soup_cli/utils/data_pipeline.py.

Can I use strategy: probs with more than two datasets?

Yes. The probs array simply needs one probability per dataset in data.train. For three datasets, use probs: [0.5, 0.3, 0.2] or any other distribution summing to 1.0. The validator enforces both length matching and proper normalization.

Why does interleaving disable packing and multipacking?

Packing and multipacking assume contiguous sequences from a single dataset source. Interleaved datasets mix rows from multiple sources, breaking the contiguous-block assumptions that packing algorithms rely on. The loader explicitly rejects this combination to prevent silent data corruption.

Is streaming interleaving slower than local file interleaving?

Streaming adds network latency but uses the same interleave_datasets or concatenate_datasets call structure. The strategy semantics remain identical—under still truncates to shortest, over still upsamples—only the data materialization path differs (network chunks vs. local memory maps).

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →