How to Create Synthetic Training Data with Soup Data Forge: A Complete Guide
Soup data forge is a multi‑stage pipeline that converts raw documents into high‑quality synthetic training rows while preserving full provenance through chunking, LLM judging, and uncertainty‑based pruning.
The soup data forge command is the primary entry point for generating supervised fine‑tuning (SFT) datasets from unstructured text. As implemented in MakazhanAlpamys/Soup, it combines document chunking, active‑learning filtering, and provenance tracking into a single reproducible workflow.
Overview of the Soup Data Forge Pipeline
The pipeline consists of three tightly integrated stages defined in src/soup_cli/utils/data_forge.py:
- Document discovery and chunking
- Judging with uncertainty‑based pruning
- Provenance manifest generation
Each stage enforces strict security boundaries: path containment via os.lstat with S_ISLNK checks, atomic writes using tempfile → os.replace, and SSRF‑hardening for remote judge providers.
Stage 1: Document Chunking
The forge command discovers all supported text files (.txt, .md, .json, .jsonl) one level deep under the --docs directory. Each file is split into chunks using the same tokenizer that will later be used for training.
Key implementation details from src/soup_cli/utils/data_forge.py:
- Default chunk size: approximately 256 tokens
- CWD‑contained execution: the process refuses to follow symlinks or traverse dot‑directories
- Path‑traversal protection prevents reading files outside the designated document root
# Prepare a directory with source documents
mkdir -p my_docs
cp *.md *.txt my_docs/
# Run chunking as part of the full pipeline
soup data forge \
--docs ./my_docs/ \
--task sft \
--target-rows 1000 \
--output forge_dataset.jsonl
Stage 2: Judging and Active‑Learning Pruning
For every chunk, the pipeline invokes an LLM judge specified by --judge‑provider (ollama, anthropic, or vllm). The judge generates a synthetic response that is compared to the source chunk using a Jaccard‑style similarity metric.
Rows with similarity above --uncertainty‑threshold are pruned as "too easy." The remaining uncertain rows—those where the judge's output diverges meaningfully from the source—are retained because they provide the strongest learning signal.
This active‑learning approach reduces dataset noise and improves downstream model performance. The threshold is configurable; lower values retain more diverse, challenging examples.
# Retain only rows with judge‑source uncertainty below 0.4
soup data forge \
--docs ./my_docs/ \
--task sft \
--target-rows 1000 \
--uncertainty-threshold 0.4 \
--judge-provider ollama \
--judge-model llama3 \
--output forge_dataset.jsonl \
--provenance forge_provenance.json
Stage 3: Provenance Manifest Generation
Every retained synthetic row is written to the output JSONL (--output). Simultaneously, a separate provenance file (--provenance) creates a complete audit trail mapping:
- Row ID → originating document path
- Judge call ID and model version
- Chunk ID within the source file
- Filter score (uncertainty metric)
This manifest enables compliance verification, debugging, and full reproducibility of training datasets. According to the Soup source code, provenance files follow a strict schema documented in docs/data.md.
Complete End‑to‑End Workflow
The following example mirrors the official Synthetic Data Workflow from examples/synthetic_workflow.md:
# 1. Generate initial synthetic data with a prompt
soup data generate \
--prompt "Write concise Python instruction/response pairs." \
--provider ollama \
--model llama3.2:3b \
--topic "Python error handling" \
--count 200 \
--output ./synth_raw.jsonl
# 2. Filter for coherence
soup data filter ./synth_raw.jsonl \
--output ./synth_filtered.jsonl \
--min-coherence 0.5
# 3. Score data quality
soup data score --input ./synth_filtered.jsonl
# 4. Remove benchmark contamination
soup data decontaminate \
--input ./synth_filtered.jsonl \
--output ./synth_clean.jsonl \
--benchmarks mmlu,gsm8k
# 5. Train with the bundled recipe
soup train --config examples/synthetic_workflow.yaml --yes
Key Implementation Files
| Path | Purpose |
|---|---|
src/soup_cli/commands/data_forge.py |
CLI entry point defining soup data forge arguments and validation |
src/soup_cli/utils/data_forge.py |
Core logic: chunking, judging, uncertainty pruning, provenance writing |
examples/synthetic_workflow.yaml |
Training configuration referencing synthetic datasets |
examples/synthetic_workflow.md |
Narrative documentation for the full pipeline |
docs/data.md |
Flag reference and provenance schema specification |
Task Types and Customization
The --task parameter controls output format:
sft(default): instruction‑response pairs for supervised fine‑tuningpreference: prompt with chosen/rejected responses for RLHFtool: function‑calling examples with tool definitions
All tasks share the same chunking and judging infrastructure in src/soup_cli/utils/data_forge.py.
Inspecting and Validating Results
After generation, use built‑in utilities to verify dataset quality:
# Quick statistics: row count, token distribution
soup data inspect forge_dataset.jsonl
# Run full quality scorecard
soup data score --input forge_dataset.jsonl
Summary
- Soup data forge transforms raw documents into training data through three stages: chunking, judging with uncertainty pruning, and provenance tracking.
- The active‑learning filter removes rows that are too easy for the judge, preserving only high‑signal examples.
- Provenance manifests enable full audit trails and regulatory compliance.
- Security features include path containment, atomic writes, and SSRF‑hardening throughout
src/soup_cli/utils/data_forge.py.
Frequently Asked Questions
What file formats does soup data forge support?
soup data forge processes .txt, .md, .json, and .jsonl files. Discovery is limited to one level deep under the --docs directory, with automatic exclusion of dot‑files and symlinked paths to prevent traversal attacks.
How does the uncertainty threshold affect my dataset?
The --uncertainty-threshold (default varies by provider) sets the Jaccard similarity cutoff. Rows where judge output and source text are more similar than this threshold are discarded. Lower thresholds (e.g., 0.3) keep harder, more diverse examples; higher thresholds (e.g., 0.6) retain easier, more faithful reconstructions.
Can I use cloud LLMs as judges?
Yes. The --judge-provider flag accepts anthropic for Claude models, vllm for self‑hosted endpoints, or ollama for local inference. All remote providers implement SSRF‑hardening as enforced in src/soup_cli/utils/data_forge.py.
Is the provenance file required?
No, but strongly recommended. Omitting --provenance disables audit trailing. The provenance JSON maps every synthetic row to its source document, chunk ID, and filter score—essential for debugging, compliance, and reproducible ML pipelines.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →