# What Data Formats Does Soup CLI Support for Fine-Tuning?

> Soup CLI supports 15+ data formats for fine-tuning including Alpaca, ShareGPT, LLaVA, and DPO. Easily ingest JSONL, CSV, Parquet, and text files for efficient model training.

- Repository: [Alpamys Makazhan/Soup](https://github.com/MakazhanAlpamys/Soup)
- Tags: api-reference
- Published: 2026-09-06

---

**Soup CLI auto-detects and ingests 15+ data formats for fine-tuning, ranging from Alpaca and ShareGPT to multimodal LLaVA and DPO preference pairs, accepting JSONL, CSV, Parquet, and plain text files.**

The `MakazhanAlpamys/Soup` repository provides a comprehensive command-line interface designed to streamline large language model fine-tuning. Understanding what data formats Soup CLI supports for fine-tuning is critical for dataset preparation, as the framework automatically identifies schemas during the ingestion process according to [`docs/data.md`](https://github.com/MakazhanAlpamys/Soup/blob/main/docs/data.md) (lines 998-1066).

## Supported Data Formats Overview

According to the documentation and the implementation in [`src/soup_cli/data/loader.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/data/loader.py), Soup CLI recognizes the following schema types for training:

### Instruction-Tuning Formats

**Alpaca** – Standard instruction-output pairs used for supervised fine-tuning.

```json
{"instruction":"…","input":"","output":"…"}

```

**ShareGPT** – Conversation-style logs preserving multi-turn dialogue history.

```json
{"conversations":[{"from":"human","value":"…"},{"from":"gpt","value":"…"}]}

```

**ChatML** – Role-based chat messages with explicit user and assistant roles.

```json
{"messages":[{"role":"user","content":"…"},{"role":"assistant","content":"…"}]}

```

### Preference Alignment Formats

**DPO / ORPO / SimPO / IPO** – Preference pairs for direct preference optimization and related alignment methods.

```json
{"prompt":"…","chosen":"…","rejected":"…"}

```

**KTO** – Unpaired preference data for Kahneman-Tversky optimization.

```json
{"prompt":"…","completion":"…","label":true}

```

### Multimodal and Vision-Language Formats

**LLaVA** – Vision-augmented dialogue combining images with text conversations.

```json
{"image":"photo.jpg","conversations":[{"from":"human","value":"<image>\nDescribe…"},{"from":"gpt","value":"…"}]}

```

**ShareGPT4V** – Multimodal vision logs for vision-language models.

```json
{"image":"chart.png","conversations":[{"from":"human","value":"<image>\nExplain…"},{"from":"gpt","value":"…"}]}

```

**Audio** – Speech plus dialogue for audio-language models.

```json
{"audio":"recording.wav","messages":[{"role":"user","content":"Transcribe."},{"role":"assistant","content":"Hello world."}]}

```

**Video** – Video-based dialogue for video-language understanding.

```json
{"video":"clip.mp4","messages":[{"role":"user","content":"Describe this clip."}]}

```

**Multimodal** – Mixed text, image, audio, and video within a single message structure.

```json
{"messages":[{"role":"user","content":[{"type":"text","text":"What’s in this?"},{"type":"image","url":"x.png"}]}]}

```

### Pre-training and Embedding Formats

**Plaintext** – Pre-training corpus with one document per line, either as JSON text fields or raw `.txt` files.

```json
{"text":"Raw text…"}

```

**Embedding** – Anchor-positive-negative triples for contrastive learning.

```json
{"anchor":"…","positive":"…","negative":"…"}

```

### Specialized Supervision Formats

**ASR** – Automatic speech recognition data in Whisper-style transcription format.

```json
{"audio":"clip.wav","text":"hello world"}

```

**PRM** – Process-reward modeling with stepwise supervision labels.

```json
{"prompt":"…","completions":["…","…"],"labels":[true,true]}

```

**Pre-tokenized** – Ready-made token IDs for bypassing tokenization during training.

```json
{"input_ids":[1,2,3],"labels":[-100,2,3],"attention_mask":[1,1,1]}

```

**Input/Output** – Segment-level loss control for masking specific portions of examples.

```json
{"segments":[{"text":"Q: hi","label":false},{"text":"A: hello","label":true}]}

```

## File Types and Auto-Detection

Soup CLI accepts **JSONL**, **JSON**, **CSV**, **Parquet**, and plain **.txt** files. The auto-detection logic resides in [`src/soup_cli/data/loader.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/data/loader.py), which parses file extensions and content structure to determine the appropriate schema without manual configuration when using commands like `soup train`.

## Configuration Schema

The `data.format` field is defined by Pydantic models in [`src/soup_cli/config/schema.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/config/schema.py). This schema validates format specifications when explicitly provided via the `--format` flag, ensuring type safety before data ingestion begins.

## Practical Usage Examples

The following commands demonstrate how to work with different data formats using the Soup CLI:

```bash

# Train on an Alpaca‑style JSONL file

soup train --data train.jsonl

# Use a ShareGPT‑style file and explicitly specify the format (optional)

soup data validate sharegpt_demo.jsonl --format sharegpt

# Convert a CSV file to the internal JSONL format

soup data convert my_data.csv --to alpaca --output converted.jsonl

# Load a plain‑text corpus for pre‑training

soup data ingest raw_corpus.txt --output corpus.jsonl

```

Sample fixture files illustrating each format are available in [`examples/data/README.md`](https://github.com/MakazhanAlpamys/Soup/blob/main/examples/data/README.md).

## Summary

- Soup CLI supports **15+ data formats** for fine-tuning, covering instruction tuning, preference alignment, multimodal tasks, and pre-training.
- Supported file types include **JSONL, JSON, CSV, Parquet, and .txt**.
- Format auto-detection is handled by [`src/soup_cli/data/loader.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/data/loader.py), eliminating the need for manual schema configuration in most cases.
- The Pydantic schema in [`src/soup_cli/config/schema.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/config/schema.py) provides runtime validation for explicit format declarations.

## Frequently Asked Questions

### Does Soup CLI require manual format specification for every dataset?

No. Soup CLI automatically detects data formats based on file content and extension using the logic in [`src/soup_cli/data/loader.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/data/loader.py). You only need to specify the format manually when using the `--format` flag for validation or conversion tasks.

### Can I use CSV files directly for fine-tuning, or do I need to convert them?

Soup CLI can ingest CSV files directly, but you may convert them to JSONL using `soup data convert` if you need to apply specific schema transformations or validate against formats like Alpaca before training.

### What is the difference between ShareGPT and ChatML formats?

**ShareGPT** uses a `conversations` array with `from` and `value` keys to denote speakers (human/gpt), while **ChatML** uses a `messages` array with `role` and `content` keys (user/assistant). Both represent multi-turn dialogue but use different schema structures as defined in [`docs/data.md`](https://github.com/MakazhanAlpamys/Soup/blob/main/docs/data.md).

### How does Soup CLI handle pre-tokenized datasets?

Pre-tokenized formats containing `input_ids`, `labels`, and `attention_mask` arrays are supported natively. The loader in [`src/soup_cli/data/loader.py`](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/data/loader.py) passes these directly to the model without re-tokenization, which is useful for experimental setups requiring fixed token boundaries.