What Data Formats Does Soup CLI Support for Fine-Tuning?
Soup CLI auto-detects and ingests 15+ data formats for fine-tuning, ranging from Alpaca and ShareGPT to multimodal LLaVA and DPO preference pairs, accepting JSONL, CSV, Parquet, and plain text files.
The MakazhanAlpamys/Soup repository provides a comprehensive command-line interface designed to streamline large language model fine-tuning. Understanding what data formats Soup CLI supports for fine-tuning is critical for dataset preparation, as the framework automatically identifies schemas during the ingestion process according to docs/data.md (lines 998-1066).
Supported Data Formats Overview
According to the documentation and the implementation in src/soup_cli/data/loader.py, Soup CLI recognizes the following schema types for training:
Instruction-Tuning Formats
Alpaca – Standard instruction-output pairs used for supervised fine-tuning.
{"instruction":"…","input":"","output":"…"}
ShareGPT – Conversation-style logs preserving multi-turn dialogue history.
{"conversations":[{"from":"human","value":"…"},{"from":"gpt","value":"…"}]}
ChatML – Role-based chat messages with explicit user and assistant roles.
{"messages":[{"role":"user","content":"…"},{"role":"assistant","content":"…"}]}
Preference Alignment Formats
DPO / ORPO / SimPO / IPO – Preference pairs for direct preference optimization and related alignment methods.
{"prompt":"…","chosen":"…","rejected":"…"}
KTO – Unpaired preference data for Kahneman-Tversky optimization.
{"prompt":"…","completion":"…","label":true}
Multimodal and Vision-Language Formats
LLaVA – Vision-augmented dialogue combining images with text conversations.
{"image":"photo.jpg","conversations":[{"from":"human","value":"<image>\nDescribe…"},{"from":"gpt","value":"…"}]}
ShareGPT4V – Multimodal vision logs for vision-language models.
{"image":"chart.png","conversations":[{"from":"human","value":"<image>\nExplain…"},{"from":"gpt","value":"…"}]}
Audio – Speech plus dialogue for audio-language models.
{"audio":"recording.wav","messages":[{"role":"user","content":"Transcribe."},{"role":"assistant","content":"Hello world."}]}
Video – Video-based dialogue for video-language understanding.
{"video":"clip.mp4","messages":[{"role":"user","content":"Describe this clip."}]}
Multimodal – Mixed text, image, audio, and video within a single message structure.
{"messages":[{"role":"user","content":[{"type":"text","text":"What’s in this?"},{"type":"image","url":"x.png"}]}]}
Pre-training and Embedding Formats
Plaintext – Pre-training corpus with one document per line, either as JSON text fields or raw .txt files.
{"text":"Raw text…"}
Embedding – Anchor-positive-negative triples for contrastive learning.
{"anchor":"…","positive":"…","negative":"…"}
Specialized Supervision Formats
ASR – Automatic speech recognition data in Whisper-style transcription format.
{"audio":"clip.wav","text":"hello world"}
PRM – Process-reward modeling with stepwise supervision labels.
{"prompt":"…","completions":["…","…"],"labels":[true,true]}
Pre-tokenized – Ready-made token IDs for bypassing tokenization during training.
{"input_ids":[1,2,3],"labels":[-100,2,3],"attention_mask":[1,1,1]}
Input/Output – Segment-level loss control for masking specific portions of examples.
{"segments":[{"text":"Q: hi","label":false},{"text":"A: hello","label":true}]}
File Types and Auto-Detection
Soup CLI accepts JSONL, JSON, CSV, Parquet, and plain .txt files. The auto-detection logic resides in src/soup_cli/data/loader.py, which parses file extensions and content structure to determine the appropriate schema without manual configuration when using commands like soup train.
Configuration Schema
The data.format field is defined by Pydantic models in src/soup_cli/config/schema.py. This schema validates format specifications when explicitly provided via the --format flag, ensuring type safety before data ingestion begins.
Practical Usage Examples
The following commands demonstrate how to work with different data formats using the Soup CLI:
# Train on an Alpaca‑style JSONL file
soup train --data train.jsonl
# Use a ShareGPT‑style file and explicitly specify the format (optional)
soup data validate sharegpt_demo.jsonl --format sharegpt
# Convert a CSV file to the internal JSONL format
soup data convert my_data.csv --to alpaca --output converted.jsonl
# Load a plain‑text corpus for pre‑training
soup data ingest raw_corpus.txt --output corpus.jsonl
Sample fixture files illustrating each format are available in examples/data/README.md.
Summary
- Soup CLI supports 15+ data formats for fine-tuning, covering instruction tuning, preference alignment, multimodal tasks, and pre-training.
- Supported file types include JSONL, JSON, CSV, Parquet, and .txt.
- Format auto-detection is handled by
src/soup_cli/data/loader.py, eliminating the need for manual schema configuration in most cases. - The Pydantic schema in
src/soup_cli/config/schema.pyprovides runtime validation for explicit format declarations.
Frequently Asked Questions
Does Soup CLI require manual format specification for every dataset?
No. Soup CLI automatically detects data formats based on file content and extension using the logic in src/soup_cli/data/loader.py. You only need to specify the format manually when using the --format flag for validation or conversion tasks.
Can I use CSV files directly for fine-tuning, or do I need to convert them?
Soup CLI can ingest CSV files directly, but you may convert them to JSONL using soup data convert if you need to apply specific schema transformations or validate against formats like Alpaca before training.
What is the difference between ShareGPT and ChatML formats?
ShareGPT uses a conversations array with from and value keys to denote speakers (human/gpt), while ChatML uses a messages array with role and content keys (user/assistant). Both represent multi-turn dialogue but use different schema structures as defined in docs/data.md.
How does Soup CLI handle pre-tokenized datasets?
Pre-tokenized formats containing input_ids, labels, and attention_mask arrays are supported natively. The loader in src/soup_cli/data/loader.py passes these directly to the model without re-tokenization, which is useful for experimental setups requiring fixed token boundaries.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →